Image encoding / decoding method and apparatus
By optimizing the inter prediction method, configuring a motion information prediction candidate group and selecting a merge candidate, the problem of insufficient inter prediction performance in the prior art is solved, and the efficiency of motion information encoding and image processing are improved.
Patent Information
- Application Number
- CN202411279872.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-07-24
- Filing Date
- 2019-07-01
- Publication Date
- 2025-07-18
AI Technical Summary
In the existing image encoding and decoding technologies, the performance improvement requirements of inter-frame prediction methods have not been met, especially in the configuration and encoding mode selection of motion information prediction candidate groups.
By determining the motion information encoding mode of the target block, configuring a motion information prediction candidate group, and selecting a merge candidate to perform inter prediction, the merge candidate list includes spatial merge candidates, temporal merge candidates and statistical merge candidates, and optimizing the encoding and decoding process of motion information.
It effectively reduces the number of bits of motion information, improves encoding performance, and improves the efficiency and quality of image processing.
Smart Images

Figure CN120343238A_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application for invention with the application date of July 1, 2019, application number 201980044609.X, and invention name "Image Coding / Decoding Method and Apparatus". Technical Field
[0002] The present disclosure relates to image coding and decoding technologies, and more particularly, to an image coding / decoding method and apparatus in inter-frame prediction. Background Art
[0003] With the provision of the Internet and mobile terminals and the development of information and communication technologies, the use of multimedia data has increased rapidly. Therefore, in order to perform various services or tasks through image prediction in all kinds of systems, the demand for improving the performance and efficiency of image processing systems has increased significantly, but the results from research and development that can respond to such an atmosphere are still insufficient.
[0004] Therefore, in the image coding and decoding methods and apparatuses of the conventional technologies, it is necessary to improve the performance of image processing, particularly the performance of image coding or image decoding. Summary of the Invention
[0005] Technical Problem
[0006] An object of the present disclosure is to provide an inter-frame prediction method.
[0007] In addition, the present disclosure provides a method and apparatus for configuring a candidate group of motion information prediction for inter-frame prediction.
[0008] In addition, the present disclosure provides a method and apparatus for performing inter-frame prediction according to a motion information coding mode.
[0009] Technical Solution
[0010] An image coding / decoding method and apparatus according to the present disclosure may: determine a motion information coding mode of a target block; configure a candidate group of motion information prediction according to the motion information coding mode; obtain the motion information of the target block by selecting one candidate from the candidate group of motion information prediction; and perform inter-frame prediction on the target block based on the motion information of the target block.
[0011] An inter prediction method performed by an image decoding device according to the present disclosure includes: determining a merge candidate list for a current block including merge candidates; selecting a merge candidate from the merge candidate list; and performing inter prediction based on the motion vector and reference picture of the selected merge candidate, and wherein the merge candidate list includes spatial merge candidates and temporal merge candidates, and when the number of the spatial merge candidates and the temporal merge candidates is less than the maximum number of the merge candidate list, the merge candidate list further includes statistical merge candidates derived from previously decoded blocks that belong to the current picture and are not adjacent to the current block, and when the number of the spatial merge candidates, the temporal merge candidates, and the statistical merge candidates is less than the maximum number of the merge candidate list, the merge candidate list further includes combined merge candidates derived from two pre-included merge candidates.
[0012] An inter prediction method performed by an image encoding device according to the present disclosure includes: determining a merge candidate list for a current block including merge candidates; selecting a merge candidate from the merge candidate list; and performing inter prediction based on the motion vector and reference picture of the selected merge candidate, and wherein the merge candidate list includes spatial merge candidates and temporal merge candidates, and when the number of the spatial merge candidates and the temporal merge candidates is less than the maximum number of the merge candidate list, the merge candidate list further includes statistical merge candidates derived from previously decoded blocks that belong to the current picture and are not adjacent to the current block, and when the number of the spatial merge candidates, the temporal merge candidates, and the statistical merge candidates is less than the maximum number of the merge candidate list, the merge candidate list further includes combined merge candidates derived from two pre-included merge candidates.
[0013] A non-transitory computer-readable recording medium according to the present disclosure stores a bitstream generated by an inter prediction method performed by an image encoding device, the method including: determining a merge candidate list for a current block including merge candidates; selecting a merge candidate from the merge candidate list; and performing inter prediction based on the motion vector and reference picture of the selected merge candidate, and wherein the merge candidate list includes spatial merge candidates and temporal merge candidates, and when the number of the spatial merge candidates and the temporal merge candidates is less than the maximum number of the merge candidate list, the merge candidate list further includes statistical merge candidates derived from previously decoded blocks that belong to the current picture and are not adjacent to the current block, and when the number of the spatial merge candidates, the temporal merge candidates, and the statistical merge candidates is less than the maximum number of the merge candidate list, the merge candidate list further includes combined merge candidates derived from two pre-included merge candidates.
[0014] Beneficial effects
[0015] When inter-frame prediction according to the present disclosure is used, since the candidate group for motion information prediction can be effectively configured to result in a reduction in the bits representing the motion information of the target block, the coding performance can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a conceptual diagram of an image coding and decoding system according to an embodiment of the present disclosure.
[0017] Figure 2 is a block diagram of components of an image coding device according to an embodiment of the present disclosure.
[0018] Figure 3 is a block diagram of components of an image decoding device according to an embodiment of the present disclosure.
[0019] Figure 4 is an exemplary diagram showing various partitioning shapes that can be obtained in the block partitioning unit of the present disclosure.
[0020] Figure 5 is an example for describing the genetic traits of members in a family and the number of families of people in the blood relationship.
[0021] Figure 6 is an example of various arrangements of related blocks that are in a horizontal relationship with the target block.
[0022] Figure 7 is an example of various arrangements of related blocks that are in a vertical relationship with the target block.
[0023] Figure 8 is an example of various arrangements of related blocks that are in a vertical relationship and a horizontal relationship with the target block.
[0024] Figure 9 is an exemplary diagram of block partitioning obtained according to the tree type. In this case, p to r represent examples of block partitioning for QT, BT, and TT.
[0025] Figure 10 is an exemplary diagram of block partitioning obtained through QT, BT, and TT.
[0026] Figure 11 is an exemplary diagram for confirming the correlation between blocks based on the partitioning method and partitioning settings.
[0027] Figure 12 is an exemplary diagram showing various situations in which a predicted block is obtained through inter-frame prediction.
[0028] Figure 13It is an exemplary diagram of a list of configuration reference pictures according to an embodiment of the present disclosure.
[0029] Figure 14 It is a conceptual diagram showing a non-translational motion model according to an embodiment of the present disclosure.
[0030] Figure 15 It is an exemplary diagram showing motion prediction in units of sub-blocks according to an embodiment of the present disclosure.
[0031] Figure 16 It is an exemplary diagram regarding the arrangement of blocks adjacent to a basic block spatially or temporally according to an embodiment of the present disclosure. Detailed implementation
[0032] Best mode
[0033] An image encoding / decoding method and apparatus according to the present disclosure may: determine a motion information encoding mode of a target block; configure a motion information prediction candidate group according to the motion information encoding mode; obtain the motion information of the target block by selecting one candidate from the motion information prediction candidate group; and perform inter-frame prediction on the target block based on the motion information of the target block.
[0034] Aspects of the present invention
[0035] Various changes and modifications can be made to the present invention, and the present invention can be described with reference to different exemplary embodiments. Some embodiments of different exemplary embodiments will be described and shown in the drawings. However, these embodiments are not intended to limit the present invention, but should be understood to include all modifications, equivalents, and alternatives belonging to the spirit and technical scope of the present invention. In the drawings, like reference numerals always refer to like elements.
[0036] Although terms such as first, second, etc. may be used to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the teachings of the present invention, the first element may be referred to as the second element, and similarly, the second element may also be referred to as the first element. The term "and / or" includes any combination and all combinations of a plurality of related listed items.
[0037] It should be understood that when an element is referred to as "connected to" or "coupled to" another element, the element may be directly connected or coupled to the other element or an intermediate element. In contrast, when an element is referred to as "directly connected to" or "directly coupled to" another element, there is no intermediate element.
[0038] The terms used in this document are for the purpose of describing specific embodiments only and are not intended to limit the present invention. Unless the context clearly indicates otherwise, as used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms. It should also be understood that when used in this specification, the terms "comprising" and / or "having" specify the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups of other features, integers, steps, operations, elements, components.
[0039] Unless otherwise defined, all terms used herein - including technical or scientific terms - have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Terms that are commonly used and defined in a dictionary should be interpreted as having the same meaning as in the context of the relevant art, and unless explicitly defined in this disclosure, they should not be interpreted as ideal or overly formal.
[0040] Generally, one or more color spaces can be configured according to the color format of an image. One or more pictures of a certain size or one or more pictures of different sizes can be configured according to the color format. In an example, in a YCbCr color configuration, color formats such as 4:4:4, 4:2:2, 4:2:0, monochrome (configured with only Y) etc. can be supported. In an example, for YCbCr 4:2:0, 1 luminance component (Y in this example) and 2 chrominance components (Cb / Cr in this example) can be configured, and in this case, the configuration ratio of the chrominance components to the luminance component can be an aspect ratio of 1:2. In an example, for 4:4:4, it can have the same aspect ratio. When pictures are configured with one or more color spaces as in the above examples, the pictures can be divided into each color space.
[0041] Since an image can be classified as I, P, B, etc. according to the image type (e.g., picture type, slice type, tile group type, tile type, brick type, etc.), the image type I can represent an image that is self - encoded without using a reference picture, the image type P can represent an image that is encoded by using a reference picture but only allows forward prediction, the image type B can represent an image that is encoded by using a reference picture and allows forward / backward prediction, but a combination of some of the above types (combining P and B) can be made according to the encoding settings, or image types in other configurations can be supported.
[0042] Various encoded / decoded information generated in the present disclosure may be processed explicitly or implicitly. In this regard, it can be understood that explicit processing may generate selection information indicating a selection of one candidate from among a plurality of candidate groups regarding information encoded in sequences, slices, tile groups, tiles, brick-shaped blocks, blocks, sub-blocks, etc., store the selection information in a bitstream, and parse relevant information in the decoder in the same units as in the encoder to reconstruct the relevant information into decoded information. In this case, it can be understood that implicit processing processes the encoded / decoded information in the encoder and the decoder with the same processing, rules, etc.
[0043] Figure 1 is a conceptual diagram of an image encoding and decoding system according to an embodiment of the present disclosure.
[0044] Referring to Figure 1 , the image encoding device (105) and the decoding device (100) may be user terminals such as a personal computer (PC), a notebook, a personal digital assistant (PDA), a portable multimedia player (PMP), a portable game console (PSP), a wireless communication terminal, a smart phone, or a TV, or server terminals such as an application server, a service server, etc., and the image encoding device (105) and the decoding device (100) may include various devices equipped with communication devices, for example, a communication modem for communicating with various instruments or wired and wireless communication networks, etc., memories (120, 125) for storing all kinds of programs and data for inter-frame prediction or intra-frame prediction to encode or decode an image, or processors (110, 115) for program operation and control by running programs, etc.
[0045] In addition, the image encoded in the bitstream by the image encoding device (105) may be sent to the image decoding device (100) in real time or non-real time through wired and wireless communication networks such as the Internet, a wireless local area network, a wireless Lan network, a wireless broadband network, or a mobile radio communication network, etc., or through various communication interfaces such as a cable or a universal serial bus, and decoded in the image decoding device (100). And it may be reconstructed into an image and played. In addition, the image encoded in the bitstream by the image encoding device (105) may be sent from the image encoding device (105) to the image decoding device (100) through a computer-readable recording medium.
[0046] The above-mentioned image encoding device and image decoding device may be separate devices, respectively. However, according to an embodiment, they may be integrated into an image encoding / decoding device. In this case, some configurations of the image encoding device may be implemented as substantially the same technical elements as some configurations of the image decoding device, including at least the same structure or at least performing the same function.
[0047] Therefore, in the detailed description of the following technical elements and their operation principles, etc., the repeated description of the corresponding technical elements will be omitted. In addition, since the image decoding device corresponds to a computing device that applies the image encoding method to be executed in the image encoding device to decoding, the following will mainly describe the image encoding device.
[0048] The computing device may include: a memory that stores programs or software modules for implementing the image encoding method and / or the image decoding method; and a processor that is connected to the memory to execute the programs. In this case, respectively, the image encoding device may be referred to as an encoder, and the image decoding device may be referred to as a decoder.
[0049] Figure 2 is a block diagram of components of an image encoding device according to an embodiment of the present disclosure.
[0050] Referring to Figure 2 , the image encoding apparatus (20) may include a prediction unit (200), a subtraction unit (205), a transform unit (210), a quantization unit (215), a dequantization unit (220), an inverse transform unit (225), an addition unit (230), a filter unit (235), a coded picture buffer (240), and an entropy encoding unit (245).
[0051] The prediction unit (200) may be implemented by using a prediction module, a software module, and may generate a prediction block for a block to be encoded by an intra prediction method or an inter prediction method. The prediction unit (200) may generate a prediction block by predicting a target block to be currently encoded in the image. In other words, the prediction unit (200) may generate a prediction block according to intra prediction or inter prediction by using the predicted pixel value of each pixel generated by predicting the pixel value of each pixel in the target block to be encoded in the image. In addition, the prediction unit (200) may cause the encoding unit to encode the information about the prediction mode by sending the information required to generate the prediction block, such as the information about the prediction mode such as the intra prediction mode or the inter prediction mode, to the encoding unit. In this case, the processing unit for performing the prediction and the processing unit for determining the prediction method and the specific content may be determined according to the encoding setting. For example, the prediction method, the prediction mode, etc. may be determined in the prediction unit, and the prediction may be performed in the transform unit.
[0052] An intra prediction unit may have directional prediction modes such as a horizontal mode, a vertical mode, etc. used according to a prediction direction, and non-directional prediction modes such as DC, planar, etc. using methods such as averaging of reference pixels, interpolation, etc. An intra prediction mode candidate group may be configured by a directional mode and a non-directional mode, and one of various candidates such as 35 prediction modes (33 directional prediction modes + 2 non-directional prediction modes), 67 prediction modes (65 directional prediction modes + 2 non-directional prediction modes), 131 prediction modes (129 directional prediction modes + 2 non-directional prediction modes), etc. that can be used as a candidate group.
[0053] The intra prediction unit may include a reference pixel construction unit, a reference pixel filtering unit, a reference pixel interpolation unit, a prediction mode determination unit, a prediction block generation unit, and a prediction mode encoding unit. The reference pixel construction unit may configure pixels that belong to a block adjacent to a target block and are adjacent to the target block as reference pixels for intra prediction. According to the coding settings, one nearest reference pixel line may be configured as a reference pixel, another adjacent reference pixel line may be configured as a reference pixel, or multiple reference pixel lines may be configured as reference pixels. When some of the reference pixels are unavailable, reference pixels may be generated by using available reference pixels, and when all reference pixels are unavailable, reference pixels may be generated by using a preset value (e.g., a central value within the range of reference values represented by the bit depth, etc.).
[0054] The reference pixel filtering unit of the intra prediction unit may perform filtering on the reference pixels to reduce residual degradation in the coding process. In this case, the filter may be a low-pass filter, such as a 3-tap filter [1 / 4, 1 / 2, 1 / 4], a 5-tap filter [2 / 16, 3 / 16, 6 / 16, 3 / 16, 2 / 16], etc. According to the coding information (e.g., the size or shape of the block, the prediction mode, etc.), it may be determined whether to apply filtering, the type of filtering, etc.
[0055] The reference pixel interpolation unit of the intra prediction unit may generate pixels in fractional units in the linear interpolation process of the reference pixels according to the prediction mode, and may determine the interpolation filter to be applied according to the coding information. In this case, the interpolation filter may include a 4-tap cubic filter, a 4-tap Gaussian filter, a 6-tap Wiener filter, an 8-tap Kalman filter, etc. Generally, interpolation is performed separately from the process of performing low-pass filtering, but filtering processing may be performed by integrating the filters applied to the two processes into one filter.
[0056] The prediction mode determination unit of the intra prediction unit may select at least one best prediction mode from a group of prediction mode candidates by considering the coding cost. The prediction block generation unit may generate a prediction block by using the corresponding prediction mode. In the prediction mode coding unit, the best prediction mode may be coded based on the prediction value. In this case, the prediction information may be adaptively coded according to the case where the prediction value is accurate and the case where the prediction value is inaccurate.
[0057] In the intra prediction unit, the prediction value may be referred to as the MPM (Most Probable Mode), and among all the modes belonging to the group of prediction mode candidates, some of the modes may be configured as an MPM candidate group. The MPM candidate group may include preset prediction modes (e.g., DC, planar, vertical, horizontal, diagonal modes, etc.) or the prediction modes of spatially neighboring blocks (e.g., left block, upper block, upper left block, upper right block, lower left block, etc.). In addition, the modes derived from the modes pre-included in the MPM candidate group (for the direction mode, the difference is +1, -1, etc.) may be configured as the MPM candidate group.
[0058] There may be a priority for the prediction modes for configuring the MPM candidate group. The order to be included in the MPM candidate group may be determined according to the priority. If the number of MPM candidates is filled according to the priority (determined according to the number of prediction mode candidates), the configuration of the MPM candidate group may be terminated. In this case, the priority may be determined in the order of the prediction modes of spatially neighboring blocks, preset prediction modes, and the modes derived from the prediction modes pre-included in the MPM candidate group, but other modifications are also possible.
[0059] For example, it may be included in the candidate group in the order of left block – upper block – lower left block – upper right block – upper left block, etc. in the spatially neighboring blocks, may be included in the candidate group in the order of DC mode – planar mode – vertical mode – horizontal mode, etc. in the preset prediction modes, and the modes obtained by adding +1, -1, etc. to the pre-included modes may be included in the candidate group to configure a total of 6 modes as the candidate group. Alternatively, it may be included in the candidate group according to a priority such as left – upper – DC – planar – lower left – upper right – upper left – (left + 1) – (left - 1) – (upper + 1), etc. to configure a total of 7 modes as the candidate group.
[0060] The subtraction unit (205) may generate a residual block by subtracting the prediction block from the target block. In other words, the subtraction unit (205) may calculate the difference between the pixel value of each pixel in the target block to be coded and the predicted pixel value of the corresponding pixel in the prediction block generated by the prediction unit to generate a residual signal in the form of a block, i.e., a residual block. In addition, the subtraction unit (205) may generate a residual block in a unit other than the block obtained by the block splitting unit described later.
[0061] The transform unit (210) may transform a spatial signal into a frequency signal. The signal obtained through the transform process is referred to as a transform coefficient. For example, a residual block having a residual signal received from a subtraction unit may be transformed into a transform block having transform coefficients, and the input signal is determined according to an encoding setting, and the input signal is not limited to the residual signal.
[0062] The transform unit may transform the residual block through a transform scheme such as, but not limited to, a Hadamard transform, a transform based on a discrete sine transform (DST), or a DCT-based transform. These transform schemes may be changed and modified in various ways.
[0063] At least one of the transform schemes may be supported, and at least one sub-transform scheme of each transform scheme may be supported. The sub-transform scheme may be obtained by modifying a part of the basis vectors in the transform scheme.
[0064] For example, in the case of DCT, one or more of the sub-transform schemes DCT-1 to DCT-8 may be supported, and in the case of DST, one or more of the sub-transform schemes DST-1 to DST-8 may be supported. A candidate group of transform schemes may be constructed using a part of the sub-transform schemes. For example, DCT-2, DCT-8, and DST-7 may be grouped into a candidate group for transformation.
[0065] The transformation may be performed in the horizontal direction / vertical direction. For example, a one-dimensional transformation may be performed in the horizontal direction through DCT-2, and a one-dimensional transformation may be performed in the vertical direction through DST-7. Using a two-dimensional transformation, the pixel values may be transformed from the spatial domain to the frequency domain.
[0066] A fixed transform scheme may be adopted, or the transform scheme may be adaptively selected according to the encoding setting. In the latter case, the transform scheme may be selected explicitly or implicitly. When the transform scheme is selected explicitly, information about the transform scheme or set of transform schemes applied in each of the horizontal and vertical directions may be generated, for example, at the block level. When the transform scheme is selected implicitly, the encoding setting may be defined according to the image type (I / P / B), color component, block size, block shape, block position, intra prediction mode, etc., and a predetermined transform scheme may be selected according to the encoding setting.
[0067] In addition, some transforms may be skipped according to the encoding setting. That is, one or more of the horizontal unit and the vertical unit may be explicitly or implicitly omitted.
[0068] In addition, the transform unit may send information required to generate a transform block to the encoding unit, such that the encoding unit encodes the information, includes the encoded information in a bitstream, and sends the bitstream to the decoder. Accordingly, the decoding unit of the decoder may parse the information from the bitstream for inverse transformation.
[0069] The quantization unit (215) may quantize an input signal. The signal obtained from quantization is referred to as quantization coefficients. For example, the quantization unit (215) may obtain a quantized block with quantization coefficients by quantizing a residual block having residual transform coefficients received from the transform unit, and the input signal may be determined according to encoding settings, which is not limited to residual transform coefficients.
[0070] The quantization unit may quantize the transformed residual block by quantization schemes such as, but not limited to, Dead Zone Uniform Threshold Quantization, quantization weighting matrix, etc. The above quantization schemes may be changed and modified in various ways.
[0071] Quantization may be skipped according to encoding settings. For example, quantization (and dequantization) may be skipped according to encoding settings (e.g., the quantization parameter is 0, i.e., lossless compression environment). In another example, when the quantization-based compression performance is not exerted according to the characteristics of the image, the quantization process may be omitted. Quantization may be skipped in the entire region or a partial region (M / 2×N / 2, M×N / 2, or M / 2×N) of the quantized block (M×N), and the quantization skip selection information may be set explicitly or implicitly.
[0072] The quantization unit may send information required to generate a quantized block to the encoding unit, such that the encoding unit encodes the information, includes the encoded information on the bitstream, and sends the bitstream to the decoder. Accordingly, the decoding unit of the decoder may parse the information from the bitstream for dequantization.
[0073] Although the above examples have been described assuming that the residual block is transformed and quantized by the transform unit and the quantization unit, a residual block with transform coefficients may be generated by transforming the residual signal, and it may not be quantized. The residual block may only undergo quantization without undergoing transformation. In addition, the residual block may undergo both transformation and quantization. These operations may be determined according to encoding settings.
[0074] The dequantization unit (220) dequantizes the residual block quantized by the quantization unit (215). That is, the dequantization unit (220) generates a residual block with frequency coefficients by dequantizing the sequence of quantized frequency coefficients.
[0075] The inverse transform unit (225) performs an inverse transform on the residual block dequantized by the dequantization unit (220). That is, the inverse transform unit (225) performs an inverse transform on the frequency coefficients of the dequantized residual block to generate a residual block having pixel values, that is, a reconstructed residual block. The inverse transform unit (225) may perform the inverse transform by reversely executing the transform scheme used by the transform unit (210).
[0076] The addition unit (230) reconstructs the target block by adding the predicted block predicted by the prediction unit (200) and the residual block restored by the inverse transform unit (225). The reconstructed target block is stored in the coded picture buffer (240) as a reference picture (or reference block) to be used as a reference picture when encoding the next block, another block, or another picture of the target block later.
[0077] The filter unit (235) may include one or more post-processing filters, such as a deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF). The deblocking filter may remove block distortion occurring at the boundaries between blocks in the reconstructed picture. The ALF may perform filtering based on values obtained by comparing the reconstructed image after filtering the block by the deblocking filter with the original image. The SAO may reconstruct the offset difference between the original image and the residual block to which the deblocking filter is applied at the pixel level. These post-processing filters may be applied to the reconstructed picture or block.
[0078] The coded picture buffer (240) may store the block or picture reconstructed by the filter unit (235). The reconstructed block or reconstructed picture stored in the coded picture buffer (240) may be provided to the prediction unit (200) that performs intra prediction or inter prediction.
[0079] The entropy coding unit (245) scans the generated sequence of quantized frequency coefficients in various scanning methods to generate a quantized coefficient sequence, encodes the quantized coefficient sequence by entropy coding, and outputs the entropy-coded coefficient sequence. The scanning pattern may be configured as one of various patterns such as zigzag, diagonal, and raster. In addition, encoded data including the encoded information received from each component may be generated, and the encoded data may be output as a bitstream.
[0080] Figure 3 is a block diagram showing an image decoding apparatus according to an embodiment of the present disclosure.
[0081] Refer to Figure 3, the image decoding apparatus (30) may be configured to include an entropy decoder (305), a prediction unit (310), a dequantization unit (315), an inverse transformation unit (320), an adder / subtractor unit (325), a filter (330), and a decoded picture buffer (335).
[0082] Furthermore, the prediction unit (310) may be configured to include an intra prediction module and an inter prediction module.
[0083] When receiving an image bitstream from the image encoding apparatus 20, the image bitstream may be sent to the entropy decoder 305.
[0084] The entropy decoder (305) may decode the bitstream into decoded data including quantized coefficients and decoding information to be sent to each component.
[0085] The prediction unit (310) may generate a prediction block based on the data received from the entropy decoder (305). Based on the reference images stored in the decoded picture buffer (335), a reference picture list may be generated using a default configuration scheme.
[0086] The intra prediction unit may include a reference pixel construction unit, a reference pixel filtering unit, a reference pixel interpolation unit, a prediction block generation unit, and a prediction mode decoding unit, and some may perform the same processing as the encoder, and some may perform the derivation process reversely.
[0087] The dequantization unit (315) may dequantize the quantized transform coefficients provided in the bitstream and decoded by the entropy decoder (305).
[0088] The inverse transformation unit (320) may generate a residual block by applying an inverse DCT, an inverse integer transformation, or a similar inverse transformation technique to the transform coefficients.
[0089] The dequantization unit (315) and the inverse transformation unit (320) may perform the processing of the transformation unit (210) and the quantization unit (215) of the above-mentioned image encoding apparatus 20 reversely, and may be implemented in various ways. For example, the dequantization unit (315) and the inverse transformation unit (320) may use the same processing and inverse transformation shared with the transformation unit (210) and the quantization unit (215), and may perform the transformation and quantization reversely using the information about the transformation and quantization processing received from the image encoding apparatus 20 (e.g., transformation size, transformation shape, quantization type, etc.).
[0090] The residual block that has been de-quantized and inversely transformed can be added to the prediction block obtained by the prediction unit (310) to generate a reconstructed image block. This addition can be performed by the addition unit / subtraction unit (325).
[0091] Regarding the filter (330), when needed, a de-blocking filter can be applied to eliminate the blocking artifacts in the reconstructed image block. To improve the video quality before and after the decoding process, other loop filters can be additionally used.
[0092] The reconstructed and filtered image block can be stored in the decoded picture buffer (335).
[0093] Although not shown in the drawings, the image encoding / decoding device may further include a block splitting unit.
[0094] The block splitting unit can split into blocks of various units and sizes. The basic coding unit (or the largest coding unit, coding tree unit CTU) can be a reference for the basic (or starting) unit used for prediction, transformation, quantization, etc. in the image encoding process. In this case, the basic coding unit can be composed of one luminance basic coding block (the largest coding block or CTB) and two basic chrominance coding blocks according to the color format (YCbCr in this example), and the size of each block can be determined according to the color format. The coding block (CB) can be obtained according to the partitioning process. The CB can be understood as a unit that is no longer further divided due to certain limitations and can be set as the starting unit for dividing into sub-units. In the present disclosure, the block conceptually includes various shapes such as triangles, circles, etc., and is not limited to squares.
[0095] Although the following description is given in the case of one color component, it also applies to other color components with some modifications proportionally according to the ratio of the color format (for example, in the case of YCbCr 4:2:0, the aspect ratio of the luminance component to the chrominance component is 2:1). In addition, although block partitioning according to other color components (for example, according to the block partitioning result of Y in Cb / Cr) is possible, it should be understood that independent block partitioning for each color component is also possible. In addition, although a common block partitioning configuration can be used (considering the proportionality to the aspect ratio), it is also necessary to consider and understand that separate block partitioning configurations are used according to the color components.
[0096] In a block splitting unit, a block can be represented as M×N, and the maximum and minimum values of each block can be obtained within this range. For example, if the maximum and minimum values of a block are 256×256 and 4×4 respectively, blocks of size 2m×2n (in this example, m and n are integers from 2 to 8), blocks of size 2m×2n (in this example, m and n are integers from 2 to 128), or blocks of size m×m (in this example, m and n are integers from 4 to 256) can be obtained. In this article, m and n can be equal or different, and one or more ranges supporting the blocks, such as the maximum and minimum values, can be generated.
[0097] For example, information about the maximum and minimum sizes of the blocks can be generated, and in some partitioning configurations, information about the maximum and minimum sizes of the blocks can be generated. In the former case, the information can be range information about the maximum and minimum sizes that can be produced in the image, while in the latter case, the information can be information about the maximum and minimum sizes that can be produced according to some partitioning configurations. The partitioning configuration can be defined by the image type (I / P / B), color component (YCbCr, etc.), block type (encoding / prediction / transformation / quantization), partitioning type (index or type), and partitioning scheme (quad-tree (QT), binary tree (BT), and ternary tree (TT) as tree methods, and SI2, SI3, and SI4 as type methods).
[0098] In addition, there can be constraints on the aspect ratio (block shape) available for the blocks, and in this regard, boundary values can be set. Only blocks less than or equal to / less than the boundary value k can be supported, where k can be defined according to the aspect ratio A / B (A is the longer or equal value between the width and the height, and B is the other value). k can be a real number equal to or greater than 1, such as 1.5, 2, 3, 4, etc. As in the above example, constraints on the shape of a block in the image can be supported, or one or more constraints can be supported according to the partitioning configuration.
[0099] In summary, whether block partitioning is supported can be determined based on the above ranges and constraints and the partitioning configuration described later. For example, when the candidate (sub-block) split from a block (parent block) meets the block support conditions, the partitioning can be supported; otherwise, the partitioning can be not supported.
[0100] The block splitting unit can be configured for each component of the image encoding device and the image decoding device, and in this process, the size and shape of the block can be determined. Different blocks can be configured according to the components. The blocks can include prediction blocks for the prediction unit, transform blocks for the transform unit, and quantization blocks for the quantization unit. However, the present disclosure is not limited thereto, and block units can be additionally defined for other components. Although in the present disclosure, the shape of each of the inputs and outputs in each component is described as rectangular, the inputs and outputs of some components can have any other shape (e.g., triangular).
[0101] The size and shape of the initial (or starting) block in the block splitting unit can be determined from a higher unit. The initial block can be split into smaller blocks. Once the optimal size and shape are determined according to the block division, the block can be determined as the initial block for a lower unit. The higher unit can be an encoding block, and the lower unit can be a prediction block or a transform block. The present disclosure is not limited thereto. On the contrary, various modification examples are possible. Once the initial block of the lower unit is determined as in the above example, the division process can be performed to detect the optimal size and shape of the block like the higher unit.
[0102] In summary, the block splitting unit can split a basic encoding block (or the largest encoding block) into at least one encoding block, and the encoding block can be split into at least one prediction block / transform block / quantization block. In addition, the prediction block can be split into at least one transform block / quantization block, and the transform block can be split into at least one quantization block. Some blocks can have a dependency relationship with other blocks (i.e., defined by a higher unit and a lower unit), or can have an independent relationship with other blocks. For example, the prediction block can be a higher unit above the transform block, or can be a unit independent of the transform block. Various relationships can be established according to the type of the block.
[0103] According to the encoding settings, it can be determined whether to combine a higher unit and a lower unit. The combination between units means that: the block of the higher unit undergoes the encoding process of the lower unit (e.g., in the prediction unit, transform unit, inverse transform unit, etc.) without being split into the lower unit. That is to say, this may mean that the division process is shared among multiple units, and division information is generated in one of the units (e.g., the higher unit).
[0104] For example, (when the encoding block is combined with the prediction block or the transform block), the encoding block can undergo prediction, transformation, and inverse transformation.
[0105] For example, (when the encoding block is combined with the prediction block), the encoding block can undergo prediction, and the transform block that is equal to or smaller than the encoding block in size can undergo transformation and inverse transformation.
[0106] For example, (when the coding block is combined with the transform block), a prediction block that is equal to or smaller than the coding block in size may undergo prediction, and the coding block may undergo transform and inverse transform.
[0107] For example, (when the prediction block is combined with the transform block), a prediction block that is equal to or smaller than the coding block in size may undergo prediction, transform, and inverse transform.
[0108] For example, (when there is no block combination), a prediction block that is equal to or smaller than the coding block in size may undergo prediction, and a transform block that is equal to or smaller than the coding block in size may undergo transform and inverse transform.
[0109] Although various cases of coding blocks, prediction blocks, and transform blocks have been described in the above examples, the present disclosure is not limited thereto.
[0110] For the combination between units, a fixed configuration may be supported in the image, or an adaptive configuration may be supported in consideration of various coding factors. The coding factors include image type, color component, coding mode (intra / inter), partitioning configuration, block size / shape / position, aspect ratio, prediction-related information (e.g., intra prediction mode, inter prediction mode, etc.), transform-related information (e.g., transform scheme selection information, etc.), quantization-related information (e.g., quantization region selection information and quantization transform coefficient coding information), etc.
[0111] When the block of the optimal size and shape has been detected as described above, mode information (e.g., partitioning information) of the block may be generated. The mode information may be included in the bitstream together with information generated from the component to which the block belongs (e.g., prediction-related information and transform-related information) and sent to the decoder, and may be parsed by the decoder at the same unit level for video decoding processing.
[0112] Now, the partitioning scheme will be described. Although, for ease of description, it is assumed that the initial block is shaped as a square, the present disclosure is not limited thereto, and the description may be applied in the same or similar manner to the case where the initial block is rectangular.
[0113] The block splitting unit may support various types of partitioning. For example, tree-based partitioning or index-based partitioning may be supported, and other methods may also be supported. In tree-based partitioning, the partitioning type may be determined based on various types of information (e.g., information indicating whether partitioning is performed, tree type, partitioning direction, etc.), while in index-based partitioning, specific index information may be used to determine the partitioning type.
[0114] Figure 4 is an exemplary diagram showing various partitioning types that can be obtained in the block splitting unit of the present disclosure. In this example, it is assumed that Figure 4The partitioning types shown are obtained by a partitioning operation (or process), which should not be construed as limiting the present disclosure. The partitioning types can also be obtained in multiple partitioning operations. Additionally, Figure 4 Other partitioning types not shown are also available.
[0115] (Tree-based partitioning)
[0116] In the tree-based partitioning of the present disclosure, QT, BT, and TT can be supported. If one tree method is supported, this can be referred to as single-tree partitioning, and if two or more tree methods are supported, this can be referred to as multi-tree partitioning.
[0117] In QT, a block is divided into two partitions (n) in each of the horizontal and vertical directions, while in BT, a block is divided into two partitions (b to g) in either the horizontal or vertical direction. In TT, a block is divided into three partitions (h to m) in either the horizontal or vertical direction.
[0118] In QT, a block can be divided into four partitions (o and p) by restricting the partitioning direction to one of the horizontal and vertical directions. Additionally, in BT, it is possible to support dividing a block only into partitions of equal size (b and c), only into partitions of different sizes (d to g), or both of these partitioning types. Additionally, in TT, it is possible to support dividing a block into partitions that are concentrated only in a specific direction (1:1:2 or 2:1:1 in the left -> right or up -> down direction) (h, j, k, and m), dividing a block into partitions that are concentrated in the center (1:2:1) (i and l), or both of these partitioning types. Additionally, it is possible to support dividing a block into four partitions in each of the horizontal and vertical directions (i.e., a total of 16 partitions) (q).
[0119] In the tree method, it is possible to support dividing a block into z partitions only in the horizontal direction (b, d, e, h, i, j, o), only in the vertical direction (c, f, g, k, l, m, p), or both of these partitioning types. Here, z can be an integer equal to or greater than 2, for example, 2, 3, or 4.
[0120] In the present disclosure, it is assumed that partitioning type n is supported for QT, partitioning types b and c are supported for BT, and partitioning types i and l are supported for TT.
[0121] Depending on the coding settings, one or more of the tree partitioning schemes can be supported. For example, QT, QT / BT, or QT / BT / TT can be supported.
[0122] In the above example, the basic tree partitioning scheme is QT, and BT and TT are included as additional partitioning schemes according to whether other trees are supported. However, various modifications can be made. Information indicating whether other trees are supported (bt_enabled_flag, tt_enabled_flag, and bt_tt_enabled_flag, where 0 indicates not supported and 1 indicates supported) can be implicitly determined according to the coding settings or explicitly determined in units such as sequences, pictures, slices, tile groups, tiles, or bricks.
[0123] The partitioning information can include information indicating whether partitioning is performed (tree_part_flag or qt_part_flag, bt_part_flag, tt_part_flag, and bt_tt_part_flag, which can have values of 0 or 1, where 0 indicates no partitioning and 1 indicates partitioning). Additionally, according to the partitioning schemes (BT and TT), information about the partitioning direction can be added (dir_part_flag or bt_dir_part_flag, tt_dir_part_flag, and bt_tt_dir_part_flag, which have values of 0 or 1, where 0 indicates <width / horizontal> and 1 indicates <height / vertical>). This can be information that may be generated when partitioning is performed.
[0124] When multi-tree partitioning is supported, various partitioning information can be configured. For ease of description, the following description gives an example of how to configure partitioning information at one depth level (that is, however, recursive partitioning can be performed by setting one or more supported partitioning depths).
[0125] In Example 1, the information indicating whether partitioning is performed is checked. If partitioning is not performed, the partitioning ends.
[0126] If partitioning is performed, the selection information about the partitioning type is checked (for example, tree_idx. It is 0 for QT, 1 for BT, and 2 for TT). The partitioning direction information is additionally checked according to the selected partitioning type, and the process proceeds to the next step (if additional partitioning is possible due to reasons such as when the partitioning depth has not reached the maximum value, the process starts again from the beginning; if additional partitioning is not possible, the partitioning process ends).
[0127] In Example 2, check the information indicating whether the partitioning is performed in a specific tree scheme (QT), and the process proceeds to the next step. If the partitioning is not performed in the tree scheme (QT), check the information indicating whether the partitioning is performed in another tree scheme (BT). In this case, if the partitioning is not performed in this tree scheme, check the information indicating whether the partitioning is performed in a third tree scheme (TT). If the partitioning is not performed in the third tree scheme (TT), the partitioning process ends.
[0128] If the partitioning is performed in the tree scheme (QT), the process proceeds to the next step. In addition, perform the partitioning in the second tree scheme (BT), check the partitioning direction information, and the process proceeds to the next step. If the partitioning is performed in the third tree scheme (TT), check the partitioning direction information, and the process proceeds to the next step.
[0129] In Example 3, check the information indicating whether the partitioning is performed in the tree scheme (QT). If the partitioning is not performed in the tree scheme (QT), check the information indicating whether the partitioning is performed in other tree schemes (BT and TT). If the partitioning is not performed, the partitioning process ends.
[0130] If the partitioning is performed in the tree scheme (QT), the process proceeds to the next step. In addition, perform the partitioning in other tree schemes (BT and TT), check the partitioning direction information, and the process proceeds to the next step.
[0131] Although in the above examples, the tree partitioning schemes are prioritized (Examples 2 and 3) or no priorities are assigned to the tree partitioning schemes (Example 1), various modified examples are also available. In addition, in the above examples, the partitioning in the current step is independent of the partitioning result of the previous step. However, the partitioning in the current step can depend on the partitioning result of the previous step.
[0132] In Examples 1 to 3, if some tree partitioning scheme (QT) was performed in the previous step and thus the process proceeds to the current step, the same tree partitioning scheme (QT) can also be supported in the current step.
[0133] On the other hand, if a certain tree partitioning scheme (QT) was not performed in the previous step and another tree partitioning scheme (BT or TT) was performed, then other tree partitioning schemes (BT and TT) can be supported in the current step and subsequent steps in addition to the certain tree partitioning scheme (QT).
[0134] In the above case, the tree configuration supported for block partitioning can be adaptive, and thus the aforementioned partitioning information may also be configured differently. (The example to be described later is assumed to be Example 3). That is to say, if the partitioning is not performed with a certain tree scheme (QT) in the previous step, the partitioning process may be performed in the current step without considering the tree scheme (QT). In addition, the partitioning information related to a certain tree scheme can be removed (for example, the information indicating whether the partitioning is performed, the information about the partitioning direction, etc. In this example <qt>in, information indicating whether to perform partitioning).
[0135] The above example relates to the adaptive partitioning information configuration for the case where block partitioning is allowed (e.g., the block size is within the range between the maximum value and the minimum value, the partitioning depth of each tree scheme does not reach the maximum depth (allowed depth), etc.). Even when block partitioning is restricted (e.g., the block size does not exist in the range between the maximum value and the minimum value, the partitioning depth of each tree scheme has reached the maximum depth, etc.), the partitioning information can be configured adaptively.
[0136] As already mentioned, in the present disclosure, tree-based partitioning can be performed recursively. For example, if the partitioning flag of a coded block with a partitioning depth k is set to 0, coded block coding is performed in the coded block with the partitioning depth k. If the partitioning flag of a coded block with a partitioning depth k is set to 1, coded block coding is performed in N sub-coded blocks with a partitioning depth k + 1 according to the partitioning scheme (where N is an integer equal to or greater than 2, such as 2, 3, and 4).
[0137] In the above process, the sub-coded block can be set as the coded block (k + 1) and partitioned into sub-coded blocks (k + 2). This hierarchical partitioning scheme can be determined according to partitioning settings such as the partitioning range and the allowed partitioning depth.
[0138] In this case, the bitstream structure representing the partitioning information can be selected from one or more scanning methods. For example, the bitstream of the partitioning information can be configured based on the order of the partitioning depth or based on whether partitioning is performed.
[0139] For example, in the case based on the partitioning depth order, the partitioning information is obtained based on the initial block at the current depth level and then at the next depth level. In the case based on whether partitioning is performed, additional partitioning information is first obtained in the blocks split from the initial block, and other additional scanning methods can be considered.
[0140] Regardless of the tree type (or all trees), the maximum block size and the minimum block size can have a common setting, or can have separate settings for each tree, or can have a common setting for two or more trees. In this case, the maximum block size can be set to be equal to or less than the maximum coded block. If the maximum block size of a predetermined first tree is different from the maximum coded block, the partitioning is implicitly performed using a predetermined second tree method until the maximum block size of the first tree is reached.
[0141] In addition, regardless of the tree type, a common split depth can be supported, where each tree can support a separate split depth, or a common split depth can be supported for two or more trees. Alternatively, split depth can be supported for some trees and not supported for other trees.
[0142] Explicit syntax elements for setting information can be supported, and some setting information can be determined implicitly.
[0143] (Index-based partitioning)
[0144] In the index-based partitioning of the present disclosure, a constant split index (CSI) scheme and a variable split index (VSI) scheme can be supported.
[0145] In the CSI scheme, k sub-blocks can be obtained by partitioning in a predetermined direction, where k can be an integer equal to or greater than 2, such as 2, 3, or 4. Specifically, regardless of the size and shape of the block, the size and shape of the sub-blocks can be determined based on k. The predetermined direction can be one direction among the horizontal direction, the vertical direction, and the diagonal direction (upper left -> lower right direction or lower left -> upper right direction), or a combination of two or more directions among the horizontal direction, the vertical direction, and the diagonal direction (upper left -> lower right direction or lower left -> upper right direction).
[0146] In the index-based CSI partitioning scheme of the present disclosure, z candidates can be obtained by partitioning in the horizontal direction or the vertical direction. In this case, z can be an integer equal to or greater than 2, such as 2, 3, or 4. One of the width and height of the sub-blocks can be the same, and the other of the width and height of the sub-blocks can be the same or different. The length ratio of the width or height of the sub-blocks is A1:A2:……:AZ, and each of A1 to AZ can be an integer equal to or greater than 1, such as 1, 2, or 3.
[0147] In addition, candidates can be obtained by partitioning into x partitions and y partitions along the horizontal direction and the vertical direction, respectively. Each of x and y can be an integer equal to or greater than 1, such as 1, 2, 3, or 4. However, candidates where both x and y are 1 may be restricted (because a already exists). Although Figure 4 The case where the sub-blocks have the same width ratio or height ratio is shown, but candidates with different width ratios or height ratios can also be included.
[0148] In addition, a candidate can be split into w partitions in one of the diagonal directions - upper left -> lower right and lower left -> upper right. Here, w can be an integer equal to or greater than 2, such as 2 or 3.
[0149] Refer to Figure 4 , according to the length ratio of each sub-block, the partitioning types can be classified into a symmetric partitioning type (b) and an asymmetric partitioning type (d and e). In addition, the partitioning types can be classified into partitioning types concentrated in a specific direction (k and m) and a central partitioning type (k). The partitioning types can be defined by various coding factors including the sub-block shape and the sub-block length ratio, and the supported partitioning types can be determined implicitly or explicitly according to the coding settings. Therefore, the candidate group can be determined based on the supported partitioning types in the index-based partitioning scheme.
[0150] In the VSI scheme, with the width w or height h of each sub-block fixed, one or more sub-blocks can be obtained by partitioning in a predetermined direction. Herein, each of w and h can be an integer equal to or greater than 1, for example, 1, 2, 4, or 8. Specifically, the number of sub-blocks can be determined based on the size and shape of the block and the value of w or h.
[0151] In the index-based VSI partitioning scheme of the present disclosure, a candidate can be partitioned into sub-blocks, one of the width and length of each sub-block being fixed. Alternatively, a candidate can be partitioned into sub-blocks, both the width and length of each sub-block being fixed. Since the width or height of the sub-block is fixed, equal partitioning in the horizontal or vertical direction can be allowed. However, the present disclosure is not limited thereto.
[0152] When the size of the block is M×N before partitioning, if the width w of each sub-block is fixed, the height h of each sub-block is fixed, or both the width w and height h of each sub-block are fixed, the number of sub-blocks obtained can be (M*N) / w, (M*N) / h, or (M*N) / w / h.
[0153] According to the coding settings, only one or both of the CSI scheme and the VSI scheme can be supported, and information about the supported scheme can be determined implicitly or explicitly.
[0154] The present disclosure will be described in the case where the CSI scheme is supported.
[0155] According to the coding settings, in the index-based partitioning scheme, the candidate group can be constructed to include two or more candidates.
[0156] For example, candidate groups such as {a, b, c}, {a, b, c, n}, or {a to g and n} can be formed. The candidate group can be an example including block types predicted to appear multiple times based on general statistical features, such as a block divided into two partitions in the horizontal direction or the vertical direction or in each of the horizontal and vertical directions.
[0157] Alternatively, candidate groups such as {a, b}, {a, o}, or {a, b, o} or candidate groups such as {a, c}, {a, p}, or {a, c, p} can be constructed. The candidate groups can be examples including the following candidates: Each candidate is divided into two partitions and four partitions in the horizontal direction and the vertical direction, respectively. This can be the following example: Configuring block types predicted to be divided mainly in a specific direction as candidate groups.
[0158] Alternatively, candidate groups such as {a, o, p} or {a, n, q} can be constructed. This can be the following example: Constructing candidate groups to include block types predicted to be divided into many smaller partitions than the block before division.
[0159] Alternatively, candidate groups such as {a, r, s} can be constructed, and it can be the following example: Determining the best division result that can be obtained in a rectangular shape from the block before segmentation by other methods (tree method), and constructing a non-rectangular shape as a candidate group.
[0160] As noted from the above examples, various candidate group constructions can be available, and considering various coding factors, one or more candidate group constructions can be supported.
[0161] Once the candidate groups are fully constructed, various division information configurations can be available.
[0162] For example, for a candidate group including an undivided candidate a and divided candidates b to s, index selection information can be generated.
[0163] Alternatively, information indicating whether division is performed (information indicating whether the division type is a) can be generated. If division is performed (if the division type is not a), index selection information for a candidate group including divided candidates b to s can be generated.
[0164] Division information can be configured in many other ways than the ways described above. In addition to the information indicating whether division is performed, binary bits can be assigned to the indexes of each candidate in the candidate group in various ways such as fixed-length binarization, variable-length binarization, etc. If the number of candidates is 2, 1 bit can be assigned to the index selection information, and if the number of candidates is 3, one or more bits can be assigned to the index selection information.
[0165] Compared with the tree-based division scheme, division types predicted to occur multiple times can be included in the candidate groups in the index-based division scheme.
[0166] Since the number of bits used to represent index information can increase according to the number of supported candidate groups, this scheme may be suitable for single-layer partitioning (e.g., the partitioning depth is limited to 0), rather than tree-based hierarchical partitioning (recursive partitioning). That is, a single partitioning operation can be supported, and the sub-blocks obtained by index-based partitioning can be not further divided.
[0167] This may mean that further partitioning into smaller blocks of the same type is not possible (e.g., the coded blocks obtained by index-based partitioning cannot be further divided into coded blocks), and it also means that further partitioning into blocks of different types is not possible (e.g., it is not possible to divide a coded block into a prediction block and a coded block). Obviously, the present disclosure is not limited to the above examples, and other modification examples are also available.
[0168] Now, a description will be given of determining block partitioning settings mainly based on the block type among coding factors.
[0169] First, a coded block can be obtained in the partitioning process. A tree-based partitioning scheme can be adopted for the partitioning process, and partitioning types such as Figure 4 a (no split), n (QT), b, c (BT), i or l (TT) can be generated according to the tree type. According to the coding settings, various combinations of tree types such as QT / QT+BT / QT+BT+TT may be available.
[0170] The following example is a process of finally dividing the coded block obtained in the above process into a prediction block and a transform block. Assume that prediction, transform, and inverse transform are performed based on the size of each partition.
[0171] In Example 1, prediction can be performed by setting the size of the prediction block to be equal to the size of the coded block, and transform and inverse transform can be performed by setting the size of the transform block to be equal to the size of the coded block (or the prediction block).
[0172] In Example 2, prediction can be performed by setting the size of the prediction block to be equal to the size of the coded block. The transform block can be obtained by partitioning the coded block (or the prediction block), and transform and inverse transform can be performed based on the size of the obtained transform block.
[0173] Here, a tree-based partitioning scheme can be adopted for the partitioning process, and partitioning types such as Figure 4 a (no split), n (QT), b, c (BT), i or l (TT) can be generated according to the tree type. According to the coding settings, various combinations of tree types such as QT / QT+BT / QT+BT+TT may be available.
[0174] Here, the partitioning process can be an index-based partitioning scheme. The partitioning type can be obtained according to the index type, for example, Figure 4 a (no split), b, c, or d of Figure 4 . According to the coding settings, various candidate groups such as {a, b, c} and {a, b, c, d} can be constructed.
[0175] In Example 3, the prediction block can be obtained by partitioning the coding block, and prediction is performed based on the size of the obtained prediction block. For the transform block, its size is set to the size of the coding block, and the transform and inverse transform can be performed on the transform block. In this example, the prediction block and the transform block can be in an independent relationship.
[0176] The index-based partitioning scheme can be used for the partitioning process, and the partitioning type can be obtained according to the index type, for example, Figure 4 a (no split), b to g, n, r, or s of Figure 4 . Various candidate groups such as {a, b, c, n}, {a to g, n}, and {a, r, s} can be constructed according to the coding settings.
[0177] In Example 4, the prediction block can be obtained by partitioning the coding block, and prediction is performed based on the size of the obtained prediction block. For the transform block, its size is set to the size of the prediction block, and the transform and inverse transform can be performed on the transform block. In this example, the transform block can have a size equal to the size of the obtained prediction block, and vice versa (the size of the transform block is set to the size of the prediction block).
[0178] The tree-based partitioning scheme can be used for the partitioning process, and the partitioning type can be generated according to the tree type, for example, Figure 4 a (no split), b, c (BT), i, l (TT), or n (QT) of Figure 4 . According to the coding settings, various combinations of tree types such as QT / BT / QT+BT may be available.
[0179] Here, the index-based partitioning scheme can be used for the partitioning process, and the partitioning type can be generated according to the index type, for example Figure 4 a (no split), b, c, n, o, or p of Figure 4 . Various candidate groups such as {a, b}, {a, c}, {a, n}, {a, o}, {a, p}, {a, b, c}, {a, o, p}, {a, b, c, n}, and {a, b, c, n, p} can be constructed according to the coding settings. In addition, in the index-based partitioning scheme, candidate groups can be constructed individually in the VSI scheme or in a combination of the CSI scheme and the VSI scheme.
[0180] In Example 5, prediction blocks can be obtained by partitioning coding blocks, and prediction is performed based on the sizes of the obtained prediction blocks. Transform blocks can also be obtained by partitioning coding blocks, and transformation and inverse transformation are performed based on the sizes of the obtained transform blocks. In this example, each of the prediction blocks and transform blocks can be obtained by partitioning coding blocks.
[0181] Here, a tree-based partitioning scheme and an index-based partitioning scheme can be used for the partitioning process, and candidate groups can be constructed in the same manner as or in a similar manner to Example 4.
[0182] In this case, the above examples are situations that may occur depending on whether the processing of partitioning each block type is shared, and this should not be construed as limiting the present disclosure. Various modification examples may also be available. In addition, various coding factors and block types can be considered to determine the block partitioning settings.
[0183] Coding factors can include image type (I / P / B), color component (YCbCr), block size / shape / location, block aspect ratio, block type (coding block, prediction block, transform block, or quantization block), partitioning state, coding mode (intra / inter), information related to prediction (intra prediction mode or inter prediction mode), information related to transformation (transformation scheme selection information), information related to quantization (quantization region selection information and quantization transform coefficient coding information).
[0184] (Description of the relationship between blocks)
[0185] Figure 5 This is an example for describing the genetic traits of members in a family and the number of families of the people in the blood relationship. For ease of description, the horizontal relationship and vertical relationship according to a specific gender (male) will be described.
[0186] Referring to Figure 5 , the target person (main) can have a horizontal relationship (a) with an older brother and a younger brother, and a vertical relationship (b) with a grandfather, a father (ancestor), a son, and a grandson (descendant). In this case, the people in the horizontal relationship may have similar genetic factors, such as appearance, physique, personality, etc. Alternatively, some factors may be similar while some factors may be dissimilar. Whether all or part of the genetic factors are similar can be determined by various environmental factors, etc. (including the mother).
[0187] This description can also be applied equally or similarly to the vertical relationship. For example, there may be a situation where the genetic factors (appearance, physique, personality) of the target person are similar to those of the father. Alternatively, some genetic factors (appearance, physique) of the target person may be similar to those of the father, but some genetic factors (personality) may be dissimilar to the father (and similar to the mother).
[0188] In another example, a target person may be genetically similar to a grandfather (or grandson), and the grandfather may be genetically similar to the grandson, but the degree of similarity between the persons can be determined based on the relationship between the persons. In other words, the similarity between the grandfather and the target person (a difference of two generations) may be high, and the similarity between the grandfather and the grandson (a difference of four generations) may be low.
[0189] Generally, direct analysis may be the first priority for grasping the characteristics of a target person, but when the target person does not exist, direct analysis is impossible. In such a case, as in the example, the characteristics of the target person can be approximately grasped by indirect analysis of persons in various relationships. Of course, it may be necessary to analyze a person having a high similarity to the target person.
[0190] This example describes the relationship between persons based on various blood relationships, which may be equally or similarly applicable to image compression coding. In this case, block-based coding will be used as an example. Information of blocks (related blocks) having various relationships with a predetermined block (target block) can be used / quoted / referred to for encoding the target block.
[0191] Here, the information of the related blocks can be data based on pixel values, data based on mode information used in the encoding process, or data based on setting information used in the encoding process. For example, it can be pixel values in the spatial domain or coefficient values (or quantized coefficients) in the frequency domain of the related blocks. Alternatively, it can be mode information generated in the process of encoding the related blocks. Alternatively, it can be information about reference settings used in the encoding process of the related blocks (e.g., reference candidate groups). Here, the data based on pixel values or the data based on mode information can be information for configuring the reference settings.
[0192] In the present invention, the relationship between blocks (the related blocks described by the target block and the reference target block) can be defined as follows.
[0193] - Horizontal relationship: In the case where there is no overlapping area between the target block and the related block (independent relationship between blocks)
[0194] - Vertical relationship: In the case where the target block is larger than the related block and includes the related block. Or in the case where the target block is smaller than the related block and is included in the related block (dependency relationship between blocks)
[0195] Here, in the case of having a horizontal relationship, the related block can be located regardless of the space to which the target block belongs. That is, the related block can belong to the same space as the target block in terms of time, or can belong to a different space from the target block in terms of time.
[0196] Here, in the case of a vertical relationship, the relevant block can be located in the space to which the target block belongs. That is, the relevant block does not belong to a space that is different from the target block in terms of time, but according to the coding settings, the relevant block can have a vertical relationship based on regions corresponding to the target block in spaces that are different in terms of time.
[0197] Figure 6 Examples of various arrangements of relevant blocks having a horizontal relationship with the target block are shown. Referring to Figure 6 , blocks placed in a horizontal relationship with the target block can be classified into blocks belonging to the same space (Curr) in terms of time and blocks belonging to different spaces (Diff) in terms of time.
[0198] Here, even if the relevant block belongs to a color component different from that of the target block (X), it is considered to belong to the same space in terms of time, but some limitations of the horizontal relationship are changed (there is a relevant block having the same size and position as the target block). Here, blocks belonging to the same space can be classified into blocks adjacent (or closest) to the target block (UL, U, UR, L, DL) and non - adjacent (or far - away) blocks (F0, F1, F2).
[0199] Among the blocks belonging to the same space, the blocks adjacent to the target block can be the blocks closest to the left, top, upper - left, upper - right, lower - left, etc. These are the blocks that have been encoded by considering the raster scan order (or Z - scan. In the case of 2×2, upper - left -> upper - right -> lower - left -> lower - right). That is, the positions of the adjacent blocks can be determined according to a predetermined scan order, and changes such as removing blocks at the above positions or adding blocks at new positions (right, bottom, lower - right, etc.) may occur according to the type of scan order (reverse Z - scan <lower - right -> lower - left -> upper - right -> upper - left>, clockwise scan <upper - left -> upper - right -> lower - right -> lower - left>, counter - clockwise scan <upper - left -> lower - left -> lower - right -> upper - right>, etc.).
[0200] In addition, the blocks not adjacent to the target block can be the blocks that have been encoded. In this case, it can belong to the same block unit as the target block (e.g., the largest coded block), or it can belong to the same segmentation unit (slice, tile, etc.). That is, limited settings can be supported, such as setting a range for the region that is not adjacent but can be included as a relevant block (based on the target block existing within the ranges of x_offset and y_offset in the horizontal and vertical directions). In the present invention, it is assumed that the blocks having a horizontal relationship with the target block have been encoded, but this is not limited thereto.
[0201] For the coding of the target block, the coding information / reference settings of the relevant blocks having a horizontal relationship can be used (referred to).
[0202] For example, the pixel values of the relevant blocks can be used to generate the predicted value of the target block. Specifically, in intra prediction, the predicted value of the target block can be obtained by applying methods such as extrapolation, interpolation, averaging, or methods such as block matching or template matching to the pixel values of the relevant blocks. In addition, in inter prediction, the predicted value of the target block can be obtained by using methods such as block matching or template matching with the pixel values of the relevant blocks. In this case, in terms of finding the predicted value in the same space, block matching or template matching can be defined as intra prediction (Mode_Intra), or can be defined as inter prediction (Mode_Inter) according to the prediction method, or it can be classified as other coding modes defined otherwise.
[0203] Here, only the pixel values in the spatial domain are used as the target, but all or some of the coefficient values in the frequency domain of the relevant blocks can be used as the predicted value of the target block (i.e., for the prediction of frequency components).
[0204] For example, the mode information of the relevant blocks can be used to encode the mode information of the target block. Specifically, in intra prediction, the prediction information (direction mode, non-direction mode, motion vector, etc.) of the relevant blocks can be used to encode the prediction information (MPM, non-MPM, etc.) of the target block. In addition, in inter prediction, the prediction information (motion vector, reference picture, etc.) of the relevant blocks can be used to encode the prediction information of the target block.
[0205] Here, according to the prediction method of intra prediction, not only can the relevant blocks belonging to the same space and the same color component as the target block in time be used as the target (prediction modes using extrapolation, interpolation, averaging, etc.), but also the relevant blocks belonging to the same space and different color components as the target block in time can be used as the target (prediction mode in which the data of different color components are copied).
[0206] Here, in the case of inter prediction, the motion vector and the reference picture are taken as examples of prediction information, but various information such as motion information coding mode, motion prediction direction, and motion model can be included.
[0207] For example, for the reference setting of the target block, the reference setting of the relevant blocks can be used. Specifically, in intra prediction, the MPM candidate group of the relevant blocks can be used as the MPM candidate group of the target block. In addition, in inter prediction, the motion prediction candidate group of the relevant blocks can be used as the motion prediction candidate group of the target block. That is to say, even if the candidate group is constructed based on the relevant blocks, this means that the candidate group of the relevant blocks can be used as it is without separate candidate group construction in the target block.
[0208] In the above example, the description has been made under the assumption that the relevant block is a block having a horizontal relationship with the target block. However, there may be many relevant blocks in the image, and at least one relevant block to be used for encoding the target block must be specified. The following case classifications are merely examples, and it is necessary to understand that various case configurations and definitions are possible and are not limited thereto.
[0209] Here, a block belonging to the same space and adjacent to the target block can be specified as a relevant block (Case 1). Alternatively, a block belonging to the same space and not adjacent to the target block can be specified as a relevant block (Case 2). Alternatively, a block belonging to a different space can be specified as a relevant block (Case 3). Alternatively, blocks belonging to all or some of (Case 1) to (Case 3) can be specified as relevant blocks.
[0210] Here, in the case of (Case 1), all or some of the upper left block, upper block, upper left block, upper right block, and lower left block (L, U, UL, UR, DL) adjacent to the target block can be specified as relevant blocks. In the case of (Case 2), one or more of the blocks not adjacent to the target block can be specified as relevant blocks. In the case of (Case 3), all or some of the central block, left block, right block, upper block, lower block, upper left block, upper right block, lower left block, and lower right block (C, L, R, U, D, UL, UR, DL, DR) adjacent to the target block and one or more of the blocks not adjacent to the target block can be specified as relevant blocks.
[0211] To specify relevant blocks, there are various verification methods. First, a block including coordinates based on a predetermined position of the target block can be specified as a relevant block. First, assume that the target block (m×n) has a range of (a + m - 1, b + n - 1) with the upper left coordinates based on (a, b).
[0212] (Case 1 or Case 3) The C block is described under the assumption that the target block and the position are the same in each picture. Therefore, the descriptions of the blocks having the same letter in the same image Curr and different images Diff may be common. However, (in the case of Case 3), the position of the C block may be different from the position of the target block in the picture, and in the example described later (i.e., the block belonging to Diff), the pixel position can be changed according to the position of the C block.
[0213] Block C refers to a block including pixels at predetermined positions among the internal pixels of the target block, such as (a, b), (a, b + n - 1), (a + m - 1, b), (a + m - 1, b + n - 1), (a + m / 2 - 1, b + n / 2 - 1), (a + m / 2 + 1, b + n / 2 - 1), (a + m / 2 - 1, b + n / 2 + 1), (a + m / 2 + 1, b + n / 2 + 1). And, block L refers to a block including pixels at predetermined positions among the pixels outside the left boundary of the target block, such as (a - 1, b), (a - 1, b + n - 1); block U refers to a block including pixels at predetermined positions among the pixels outside the upper boundary of the target block, such as (a, b - 1), (a + m - 1, b - 1). In addition, block UL refers to a block including pixels at predetermined positions among the pixels outside the upper left boundary of the target block, such as (a - 1, b - 1); block UR refers to a block including pixels at predetermined positions among the pixels outside the upper right boundary of the target block, such as (a + m, b - 1), (a - 1, b + n); block DL refers to a block including pixels at predetermined positions among the pixels outside the lower left boundary of the target block, such as (a - 1, b + n). In the cases of the right direction, the downward direction, and the lower right direction, since they can be derived from the above description, they are omitted.
[0214] In the above description, examples are given of designating a block including a pixel at one position among the pixels in the blocks existing in each direction as a relevant block, but two or more relevant blocks can be designated in all directions or some directions, and two or more pixel positions can be defined therefor.
[0215] (Case 2) Block Fk (k is from 0 to 2) may refer to a block including pixels separated by a predetermined length (off_x, off_y, etc.) in a predetermined direction of horizontal / vertical / diagonal, such as (a - off_x, b), (a, b - off_y), (a - off_x, b - off_y). Here, the predetermined length can be an integer of 1 or greater, such as 4, 8, 16, etc., and can be set based on the horizontal length and vertical length of the target block. Alternatively, it can be set based on the horizontal length and vertical length of the maximum coding block, and various modifications thereof are possible. The predetermined length can be implicitly set as in the above examples, or relevant syntax elements can be generated as explicit values.
[0216] As another method of specifying a relevant block, a block having pattern information identical / similar to the coding information of the target block may be specified as the relevant block. In this case, the pattern information does not refer to the information to be coded (or used / predicted, candidate group construction, etc.) at the current stage, but rather information that has already been coded or has different attributes / meanings that have already been determined. The determined information may be information determined in a previous step in the pattern determination process, or information recovered in a previous step in the decoding process.
[0217] For example, when the target block performs inter prediction using a non-translational motion model, a block that has been coded using the non-translational motion model among the previously coded blocks may be specified as the relevant block. In this case, the motion vector information of the non-translational motion model according to the specified relevant block may be used to construct a candidate group for motion vector prediction according to the non-translational motion model of the target block.
[0218] In the above example, the motion model information refers to information having different attributes / meanings for checking identity / similarity with the target block, and the motion vector according to the motion model (non-translational motion model) of the relevant block may be information for constructing a motion vector prediction candidate group according to the motion model of the target block. In this case, the relevant block may or may not be adjacent to the target block. If there are few patterns having identity / similarity with the target block, it may be useful when blocks in a region not adjacent to the target block are also specified as relevant blocks and used / referred to.
[0219] The following may be considered to determine the relevant block used / referred to for coding the target block.
[0220] The relevant block may be determined based on the information to be used / referred to for coding the target block. Here, the information to be used / referred to for coding the target block is pixel value information for prediction, pattern information related to prediction / transformation / quantization / loop filtering / entropy coding, etc., and reference candidate group information related to prediction / transformation / quantization / loop filtering / entropy coding, etc.
[0221] In addition, the relevant block may be determined based on the state information of the target block and the image information to which the target block belongs. Here, the state information of the target block may be defined based on the block size, block shape, horizontal / vertical length ratio of the block, and position in a unit such as a picture / segmentation unit (slice, tile, etc.) / largest coding block. Here, the image information to which the target block belongs may be defined based on the image type (I / P / B), color component (Y / Cb / Cr), etc.
[0222] In addition, relevant blocks can be determined based on the coding information of the target block. Specifically, relevant blocks can be determined based on whether the relevant blocks have information regarding the existence of identity / similarity with the target block. Here, the information for identity / similarity reference can be pattern information related to prediction / transformation / quantization / loop filtering / entropy coding, etc.
[0223] Considering all or some of the factors mentioned in the above examples, the category (the aforementioned case), quantity, and position of relevant blocks can be determined. Specifically, it can be determined which category is selected, and the quantity and position of relevant blocks supported within the selected category can be determined. In this case, the quantity of blocks supported in each category can be m, n, o, and the quantity of blocks supported in each category can be an integer of 0 or greater, such as 0, 1, 2, 3, 5, etc.
[0224] Relevant blocks (block positions) can be determined in directions such as left, right, up, down, upper left, upper right, lower left, lower right, center, etc. centered on the target block (or a block corresponding to the target block in an image different in time from the image to which the target block belongs). For example, relevant blocks can be determined based on the block closest to this direction. Alternatively, relevant blocks can be determined among blocks that additionally satisfy the direction and a specific range / condition. It can belong to a largest coding block different from the largest coding block to which the target block belongs, or can be a block at a position (e.g., left direction, up direction, upper left direction) having a difference based on the horizontal length or vertical length of the target block.
[0225] In addition, relevant blocks can be determined based on the coding order, and in this case, the coding order can be defined by various scanning methods such as raster scan and z-scan. As an example, a predetermined quantity of blocks can be included as relevant blocks (based on the proximity of the coding order), and the predetermined quantity can be an integer greater than or equal to 0, such as 0, 1, 2, 3, etc. That is, relevant blocks can be managed according to a memory management method such as FIFO (First In First Out) based on the coding order, and it can be an example of determining relevant blocks that may appear in (Case 2).
[0226] When one relevant block is supported, this may mean that only the information of the corresponding block can be used / referred to. In addition, even when multiple relevant blocks are supported, one piece of information can be derived based on the multiple relevant blocks according to the coding settings. For example, in inter-frame prediction, three relevant blocks such as the left block, the upper block, and the upper right block are specified for motion vector prediction to support three motion vectors, but one motion vector can be derived by a method such as the median (or average value) of the motion vectors of the three blocks according to the coding settings to be used as the motion vector prediction value of the target block.
[0227] In the case of the above example, it can be an encoding setting that can reduce the occurrence of the best candidate selection information generated by supporting two or more candidates. However, it may be difficult to expect to obtain one candidate that has a high correlation with the encoding information of the target block. Therefore, a method of constructing a candidate group having a plurality of candidates may be more efficient. Of course, as the number of candidates included in the candidate group increases, the amount of information for expressing this may increase, so it is important to construct an efficient candidate group.
[0228] Therefore, it is possible to support related blocks such as the various examples described above, but it may be necessary to consider general image characteristics, etc. to specify the best related block and construct a candidate group based on the best related block. In the present invention, it is assumed that candidates are constructed from two or more pieces of information from one or more related blocks.
[0229] The following shows the candidate group construction based on a block having a horizontal relationship with the target block and encoding / decoding processing.
[0230] Specify the block (1) that is referred to for the encoding information of the target block. Construct a candidate group in a predetermined order based on the encoding information of the specified block (2). Select one candidate from the candidate group based on the encoding information of the target block (3). Perform image encoding / decoding processing based on the selected candidate (4).
[0231] In (1), specify the related block for constructing the candidate group for the encoding information of the target block. In this case, the related block can be a block having a horizontal relationship with the target block. It has been described that: various categories of related blocks as described above can be included; in addition to the encoding information of the target block, various information such as the status information of the target block can be considered to specify the related block.
[0232] In (2), based on the encoding information of the related block specified by the above processing, construct a candidate group according to a predetermined order. Here, the information obtained based on the encoding information of one related block can be included in the candidate group, or the information obtained based on the encoding information of a plurality of related blocks can be included in the candidate group. In this case, for the candidate group construction order, a fixed order can be supported, or an adaptive order based on various encoding elements (elements to be considered when specifying related blocks, etc.) can be supported.
[0233] In (3), select one candidate from the candidate group based on the encoding information of the target block, and in (4), image encoding / decoding processing can be performed based on the selected candidate.
[0234] The flowchart can be a process that is checked and executed in units of blocks. Here, in the case of some steps (1, 2), it can be a process that is checked and executed in the initial stage of encoding. Even if this content is not mentioned in the above (1) to (4), since it can be obtained through the above various embodiments, its detailed description will be omitted. Generally, it is difficult to pre-confirm which block has a high correlation with the target block among blocks with a horizontal relationship. A method for pre-confirming what correlation it has with the target block among blocks with a horizontal relationship will be described later. In addition, although (3) describes the case of selecting one candidate from the candidate group, two or more pieces of information can also be selected according to the type of encoding information, encoding settings, etc., and this can be a description generally applicable to the present invention.
[0235] The target block and the related block can be one of the units such as encoding / prediction / transformation / quantization / loop filtering, etc., and the target block and the related block can be set in the same unit. For example, when the target block is an encoding block, the related block can also be an encoding block, and modifications with different units set according to the encoding settings are also possible.
[0236] Figure 7 Examples of various arrangements of related blocks in a vertical relationship with the target block are shown.
[0237] Figure 7 This is the case of performing recursive tree-based partitioning (QT), and the description will be given centered on blocks X and A to C. It is possible to start from the basic encoding block (CTU.C block) with a segmentation depth of 0, and obtain blocks B (1), X (2), and A (3) as the segmentation depth increases. Here, the blocks placed in a vertical relationship with the target block can be classified into higher blocks (or ancestor blocks) and lower blocks (or descendant blocks). In this case, the higher block of the target block X can be block B or block C, and the lower block can be block A. Here, the target block and the related block can be set as the higher block and the lower block respectively, or can be set as the lower block and the higher block respectively.
[0238] For example, in this example, if the related block has a value greater than the segmentation depth (k) of the target block, the related block can be a child block (k + 1) or a grandchild block (k + 2), and if the related block has a smaller value, the related block can be a parent block (k - 1) or a grandparent block (k - 2). That is, in addition to defining the vertical relationship between existing blocks, the detailed relationship between blocks can also be confirmed by the segmentation depth.
[0239] Here, as in the above example, one tree method is supported and comparison can be made through the common segmentation depth. When multiple tree methods are supported and more than one segmentation depth according to each tree method is supported, the detailed relationship can be checked by considering the number of segmentations and each segmentation depth, rather than the simple classification as in the above example.
[0240] For example, if QT is executed once in a 4M×4N block, 2M×2N blocks can be obtained at the (QT) split depth of 1. However, when BT is executed twice, 2M×2N blocks can be obtained. However, the (BT) split depth is the same as the split depth that can be obtained at 2. In this case, the 4M×4N block can be the parent block (QT) or grandparent block (BT) for the 2M×2N block; conversely, the 4M×4N block can be the child block (QT) or grandchild block (BT) for the 2M×2N block, and the detailed relationship can be determined based on the block split result.
[0241] In the above example, the starting unit of the split is the maximum coding unit (the highest ancestor block, the maximum size that a block can have. Here, it is assumed that the starting unit of the split is a coding unit or a block. When a block is a unit for prediction or transformation, the block can be understood as the maximum prediction block or the maximum transformation block). It is impossible to have a vertical relationship exceeding the maximum coding unit, but the block area having a vertical relationship can be freely specified according to a coding setting separate from the block split setting such as the maximum coding unit. In the present invention, it is assumed that the vertical relationship does not deviate from the maximum coding unit. In addition, the relationship between blocks will be described centering on the tree-based partitioning method later, but it will be mentioned in advance that the same or similar application is possible for the index-based partitioning method.
[0242] For the coding of the target block, the coding information / reference setting of the related block having a vertical relationship can be used. For ease of explanation, it is assumed that the target block is a lower block of the related block. In this case, the higher block is not an independent unit that performs coding, prediction, transformation, etc., but can be a temporary unit composed of multiple lower blocks. That is to say, it must be understood that the higher block is the starting unit or intermediate unit of the block partitioning process for obtaining an independent unit for performing coding (i.e., a coding / prediction / transformation block, etc. that no longer performs partitioning).
[0243] For example, the reference pixels of the related block can be used to generate the prediction value of the target block. Specifically, in intra-frame prediction, the prediction value of the target block can be obtained by applying methods such as extrapolation, interpolation, averaging, or template matching to the reference pixels of the related block. In addition, in inter-frame prediction, the prediction value of the target block can be obtained by using methods such as template matching with the reference pixels of the related block.
[0244] Here, the reference pixels of the related block are not the pixels located in the related block, but refer to the pixels obtained by assuming that the related block is a unit for performing intra-frame prediction / inter-frame prediction. That is to say, this means that the pixels of the block having a horizontal relationship with the related block (higher block) (for example, the pixels closest to the left, upper, upper-left, upper-right, lower-left directions) are used for intra-frame prediction / inter-frame prediction of the target block (lower block).
[0245] For example, for the reference setting of a target block, the reference setting of a related block can be used. Specifically, in intra prediction, the MPM candidate group of the related block can be used as the MPM candidate group of the target block. In addition, in inter prediction, the motion prediction candidate group of the related block can be used as the motion prediction candidate group of the target block. That is, even if the candidate group is constructed based on the related block, this means that the candidate group of the related block can be used as it is without going through a separate candidate group construction for the target block.
[0246] In the above example, the prediction value and reference setting are determined based on the related block rather than the target block, so there may be a problem of performing encoding using information with poor correlation to the target block. However, since the related information is not obtained from completely separated regions, there may be a certain degree of correlation. Since the processing to be performed in each lower block unit is integrated into a common processing in the higher block, the complexity can be reduced. In addition, parallel processing of the lower blocks belonging to the higher block is possible.
[0247] In the above example, the description has been made based on the assumption that the related block is a block having a vertical relationship with the target block, but there may be many related blocks in the image, and at least one related block to be used for encoding the target block must be specified.
[0248] The following shows a description of the support conditions / range that a block having a vertical relationship can have, and the support conditions / range can be determined based on all or some of the factors mentioned in the examples to be described later.
[0249] (Case 1) The higher block can be less than or equal to a predetermined first threshold size. Here, the first threshold size may mean the maximum size that the higher block can have. Here, the first threshold size can be expressed as width (W), height (H), W×H, W*H, etc., and W and H can be integers of 8, 16, 32, 64 or more. Here, the block having the first threshold size can be set based on the sizes of the maximum coding block, the maximum prediction block, and the maximum transform block.
[0250] (Case 2) The lower block can be greater than or equal to a predetermined second threshold size. Here, the second threshold size may mean the minimum size that the lower block can have. Here, the second threshold size can be expressed as width (W), height (H), W×H, W*H, etc., and W and H can be integers of 4, 8, 16, 32 or more. However, the second threshold size can be set to be less than or equal to the first threshold size. Here, the block having the second threshold size can be set based on the sizes of the minimum coding block, the minimum prediction block, and the minimum transform block.
[0251] (Case 3) The minimum size of the lower block can be determined based on the size of the higher block. Here, the minimum size of the lower block (e.g., W%p or H>>q, etc.) can be determined by applying a predetermined division value (p) or a shift operation value (q. right shift operation) to at least one of the width (W) or height (H) of the higher block. Here, the division value can be an integer such as 2, 4, 8, or a larger integer, and the shift operation value q can be an integer such as 1, 2, 3, or a larger integer.
[0252] (Case 4) The maximum size of the higher block can be determined based on the size of the lower block. Here, the maximum size of the higher block (e.g., W*r or H<<s, etc.) can be determined by applying a predetermined multiplication value (r) or a shift operation value (s. left shift operation) to at least one of the width (W) or height (H) of the lower block. Here, the multiplication value can be an integer such as 2, 4, 8, or a larger integer, and the shift operation value s can be an integer such as 1, 2, 3, or a larger integer.
[0253] (Case 5) The minimum size of the lower block can be determined considering the size of the higher block and the division settings. Here, the division settings can be determined by the split type (tree type), split depth (common depth, individual depth for each tree), etc. For example, if QT is supported in the higher block, the size of the block after m divisions can be determined as the minimum size of the lower block; if BT (or TT) is supported, the size of the block after n divisions can be determined as the minimum size of the lower block; if both QT and BT (or TT) are supported, the size of the block after 1 division can be determined as the minimum size of the lower block. Here, m to l can be integers such as 1, 2, 3, or a larger integer. The split depth (m) of the tree that is split into smaller blocks (or a larger number) (due to one split operation) can be set to be less than or equal to the split depth (n) of the tree that is not split into smaller blocks. In addition, the split depth (l) when tree splitting is mixed can be set to be greater than or equal to the split depth (m) of the tree that is split into smaller blocks, and can be set to be less than or equal to the split depth (n) of the tree that is not split into smaller blocks.
[0254] Alternatively, the maximum size of the higher block can be determined considering the size of the lower block and the division settings. In this description, it can be derived conversely from the above examples, so the detailed description will be omitted.
[0255] The following can be considered to determine the relevant blocks used / referred to for encoding the target block.
[0256] The relevant blocks can be determined based on the information to be used / referred to for encoding the target block. Here, the information to be used / referred to for encoding the target block can be pixel value information for prediction and information on a group of reference candidates related to prediction / transformation / quantization / loop filtering / entropy coding, etc.
[0257] In addition, the relevant blocks can be determined based on the state information of the target block and the image information to which the target block belongs. Here, the state information of the target block can be defined based on the block size, block shape, horizontal / vertical length ratio of the block, and the position in units such as pictures / segmentation units (slices, tiles, etc.) / largest coding blocks. Here, the image information to which the target block belongs can be defined based on the image type (I / P / B), color components (Y / Cb / Cr), etc.
[0258] All or some of the factors mentioned in the above examples can be considered to determine the number, size, position, etc. of the relevant blocks. Specifically, it can be determined whether to use / reference the information of the blocks having a vertical relationship for encoding the target block, and (when used / referred to) the position and size of the relevant blocks can be determined. Here, the position of the relevant block can be expressed in terms of a predetermined coordinate within the block (e.g., the upper left coordinate), and the size of the relevant block can be expressed in terms of the width (W) and height (H). The relevant blocks can be specified by combining these.
[0259] For example, if there is no special range limit for the lower blocks, all the lower blocks (target block) belonging to the relevant blocks can use / reference the coding information of the relevant blocks. Alternatively, if the range for the lower blocks is limited, the coding information of the relevant blocks can be used / reference if it belongs to the relevant blocks and is larger than the size of the lower blocks. In addition, when two or more relevant blocks are supported, selection information for the relevant blocks can be generated additionally.
[0260] The following shows a candidate group structure based on the blocks having a vertical relationship with the target block and the coding / decoding process.
[0261] Determine the basic block (1) for specifying the block to be referred to. Based on the determined basic block, specify the block (2) whose coding information is to be referred to for encoding the target block. Use the coding information of the specified block to construct a candidate group in a predetermined order (3). Select one candidate from the candidate group based on the coding information of the target block (4). Perform image coding / decoding processing based on the selected candidate (5).
[0262] In (1), determine the block (basic block) that serves as the criterion for constructing a candidate group related to the coding information of the target block from the target block or the first relevant block. Here, the first relevant block can be a block having a vertical relationship with the target block (here, the higher block).
[0263] In (2), when determining the basic block through the above processing, a second related block for specifying a candidate group related to the coding information of the target block is designated. Here, the second related block may be a block having a horizontal relationship with the basic block. In (3), the coding information of the second related block designated through the above processing is used to construct the candidate group in a predetermined order. Since blocks having a horizontal relationship can be obtained not only through the construction of the candidate group based on the blocks having a horizontal relationship and the encoding / decoding process, but also through the above various embodiments, the description of the blocks having a horizontal relationship will be omitted.
[0264] In (4), one candidate in the candidate group is selected based on the coding information of the target block, and in (5), image encoding / decoding processing can be performed based on the selected candidate.
[0265] In this case, when determining the higher block and the lower block based on the vertical relationship setting between blocks, since the processing of constructing the candidate group based on the higher block is only performed once, the lower block can use / borrow them. That is to say, the flowchart can be a configuration that can be generated in the block where encoding / decoding is first performed. If the construction of the candidate group based on the higher block has been completed and the basic block is determined as the related block in a certain order (2, 3), the already constructed candidate group can be simply used / borrowed.
[0266] In the above example, if the candidate group is constructed based on the higher block, the lower block is described as a simple use / borrow configuration, but it is not limited to this.
[0267] For example, even if the candidate group is constructed based on the higher block, some candidates can be fixed regardless of the lower block, while other candidates can be adaptive based on the lower block. That is to say, this means that some candidates can be deleted / added / changed based on the lower block. In this case, the deletion / adding / change can be performed based on the position and size of the lower block within the higher block.
[0268] That is to say, even if the basic block is determined as the related block in some steps (2, 3), a candidate group in which all or some modifications to the already constructed candidate group are reflected can be constructed.
[0269] The target block and the related block can be one of the units in encoding / prediction / transformation / quantization / loop filtering, etc., and the target block can be the same unit as the related block or a higher unit. For example, when the target block is an encoding block, the related block can be an encoding block, and when the target block is an encoding block, the related block can be a prediction block or a transformation block.
[0270] Figure 8 Various layout examples of the related blocks having vertical and horizontal relationships with the target block are shown.
[0271] Figure 8 A case where recursive tree-based partitioning (quadtree) is performed is shown, and the description will be centered around blocks X, A to G, and p to t. Starting from a basic coding block (CTU) with a partition depth of 0, as the partition depth increases, blocks q / r / t (1), p / D / E / F / s (2), A / B / C / G / X (3) can be obtained. Here, it can be classified into blocks having a vertical relationship with the target block and blocks having a horizontal relationship with the target block. In this case, the related blocks (higher blocks) having a vertical relationship with the target block X can be blocks p and q (excluding CTU), and the related blocks having a horizontal relationship can be blocks A to G.
[0272] Here, in the case of some related blocks, there are not only blocks having a smaller value than the segmentation depth (k) of the target block, but also blocks having a larger value than the segmentation depth (k) of the target block. It is assumed that the present embodiment targets the above-mentioned block and the target block has the maximum segmentation depth (i.e., no longer segmented).
[0273] In block segmentation, the segmentation result is determined according to the characteristics of the image. In flat parts such as background or areas with small time changes, block segmentation can be minimized. In parts with complex patterns or areas with fast time changes, block segmentation can be performed in large quantities.
[0274] Through the above-mentioned examples of the present invention, it has been mentioned that blocks with many horizontal or vertical relationships can be used / referenced for encoding the target block. Through the examples to be described later, various examples of methods for using them more efficiently by considering the blocks with horizontal relationships and blocks with vertical relationships around the target block will be presented. Therefore, the premise is that the description of the above-mentioned horizontal and vertical relationships can be applied to the content to be described later in the same or similar manner.
[0275] Next, various cases of correlation between blocks placed in a horizontal relationship according to the type (method) of block partitioning will be described. In this case, it is assumed that QT, BT, and TT are supported as partitioning schemes, BT is symmetric partitioning (SBT), and TT is partitioning at a ratio of 1:2:1. In addition, it is assumed that blocks placed in a general horizontal relationship can have high correlation or low correlation (general relationship). In each case, it is assumed that only one of the described tree methods is supported.
[0276] Figure 9 is an exemplary diagram of block partitioning obtained according to the tree type. Here, p to r represent examples of block partitioning of QT, BT, and TT. It is assumed that block partitioning is performed when the encoding of the block itself is not efficient due to different characteristics of some areas in the block.
[0277] In the case of QT(p), it is divided into two horizontally and vertically respectively, and it can be seen that at least one of the four sub-blocks has different characteristics. However, since four sub-blocks are obtained, it is impossible to know which specific sub-block has different characteristics.
[0278] For example, blocks A to D can all have different characteristics, only one of blocks A to D can have different characteristics, and the remaining blocks can have the same characteristics. Blocks A and B have the same characteristics, and blocks C and D can have the same characteristics. Blocks A and C can have the same characteristics, and blocks B and D can have the same characteristics.
[0279] If blocks A and B, blocks C and D each have the same characteristics and BT is supported, then horizontal division can be performed in BT. However, if it is divided by QT like p, it can be seen that blocks A and B, blocks C and D have different characteristics from each other. However, in this example, since it is assumed that only QT is supported, the correlation between the blocks cannot be accurately determined.
[0280] If only QT is supported and QT is performed on a higher block with a division depth difference of 1 around the target block (one of the sub-blocks), the correlation between the target block and the related blocks (blocks other than the target block among the sub-blocks) may be high or low.
[0281] BT(q) is divided into two in the horizontal or vertical direction, and it can be seen that the two sub-blocks (blocks E and F) have different characteristics. If the characteristics between the sub-blocks are the same or similar, it can be obtained as explained above under the assumption that it will not be divided. If only BT is supported and BT is performed on a higher block with a division depth difference of 1 around the target block, the correlation between the target block and the related blocks may be low.
[0282] In the case of TT(r), TT(r) is divided into three in the horizontal or vertical direction, and it can be seen that at least one of the three sub-blocks has different characteristics. However, since three sub-blocks are obtained, it is impossible to know which specific sub-block has different characteristics.
[0283] For example, blocks G to I may all have different characteristics. Blocks G and H have the same characteristics, while block I can have different characteristics. Blocks H and I have the same characteristics, while block G can have different characteristics.
[0284] If block G and block H have the same characteristics, block I has different characteristics, and BT (asymmetric) is also supported, a vertical split (3:1) can be performed in BT. However, if it is split by TT like r, it can be seen that block G, block H, and block I have different characteristics. However, since only TT is assumed to be supported in this example, the correlation between blocks cannot be accurately determined. If only TT is supported and TT is performed on higher blocks with a depth difference of 1 around the target block, the correlation between the target block and the related blocks may be high or low.
[0285] As described above: In the case of QT and TT in the partitioning scheme, the correlation between blocks may be high or low. Assume that only a partitioning scheme (e.g., QT) is supported, and the coding information of the remaining sub-blocks except for one sub-block (e.g., block D) is known. If the coding information of the remaining sub-blocks A to C except for block D is the same or similar, block D may have different characteristics and thus can be split by QT. In this way, the correlation information can also be identified by checking the coding information of the lower blocks belonging to the higher block. However, since the possibility of occurrence is low and this may be a complex situation, the description thereof will only mention the possibility and its detailed description will be omitted. For reference, in the above case, since block D has a low correlation with the blocks having a horizontal relationship, the coding information of the blocks having a vertical relationship can be used / referred to, which refers to the candidate group construction information of the higher block (the block including A and D) of block D.
[0286] In the above example, the correlation between sub-blocks has been described when one tree split is supported for the higher block. That is, when using / referring to the coding information of the blocks having a horizontal relationship for encoding the target block, the coding information of the related blocks can be used / referred to more efficiently by checking the split state (path) between the target block and the related blocks. For example, when constructing the candidate group, the information about the blocks determined to have a low correlation can be excluded, or a low priority can be assigned to the blocks determined to have a low correlation. Alternatively, the information and reference settings of the blocks having a vertical relationship can be used.
[0287] In this case, the situation of supporting one tree split may include the case of only supporting one tree method for block splitting. Even if multiple tree splits are supported, it may include the case of only supporting one tree split based on the maximum value, minimum value, and maximum split depth of the blocks according to each tree method, and the block split setting of the tree not allowed at the previous split depth is not supported at a later split depth. That is, only QT is supported and splitting can be performed using QT. Moreover, QT / BT / TT are supported, but only BT is possible at this stage, so splitting can be performed using BT.
[0288] The case of checking the correlation between each block when multiple tree partitions are supported will be described below.
[0289] Figure 10 is an exemplary diagram of partitions obtained due to QT, BT, and TT. In this example, it is assumed that the maximum coding block is 64×64 and the minimum coding block is 8×8. Additionally, it is assumed that the maximum value of the block supported by QT is 64×64, the minimum value is 16×16, the maximum value of the block supported by BT is 32×32, the minimum value is 4, where 4 is one of the horizontal length / vertical length of the block, and the maximum partition depth is 3. In this case, it is assumed that the partition setting for TT is determined together with BT (they are used bundled). It is assumed that QT and BT (QT, BT) are supported for the higher blocks (A to M), additionally asymmetric BT (ABT) is supported for the lower left blocks (N to P) (QT, BT <or SBT>, ABT), and TT is additionally supported for the lower right blocks of TT (Q to S) (QT, SBT, ABT, TT).
[0290] (Basic block: the block including B, C, D, E)
[0291] Blocks B to E can be sub - blocks that can be obtained by QT (partitioned once), and can also be sub - blocks that can be obtained by BT (the number of partitions is 3 times, one vertical partition + two horizontal partitions, or one horizontal partition + two vertical partitions).
[0292] Since the maximum block size supported by QT is 16×16, blocks B to E cannot be obtained by QT, and blocks B to E can be sub - blocks obtained by BT partitioning. In this example, a horizontal partition (B + C / D + E) is performed in BT, and vertical partitions are performed in each region (B / C / D / E). Therefore, as described above, the correlation between block B and block C, and between block D and block E obtained from the same higher block (parent block, partition depth difference is 1) by BT may be low.
[0293] Furthermore, since block B and block D, or block C and block E are partitioned, the correlation between block B and block D or between block C and block E may be low. This is because: if the correlation is high, only vertical partitioning is performed in BT, and partitioning may not be performed in each region.
[0294] In the above example, it is mentioned that the correlation between sub - blocks obtained by BT is low, but this is only limited to sub - blocks belonging to the same higher block (parent block, partition depth difference is 1). However, in this example, the correlation between blocks can be checked by extending to the same higher block (ancestor block, partition depth difference is 2).
[0295] (Basic block: the block including J, K, L, M)
[0296] Block J to block M can be sub - blocks obtainable through QT or sub - blocks obtainable through BT.
[0297] Since it is within the range that can support both QT and BT, the tree method can be selected as the optimal segmentation type. This is an example of executing QT. Through the above example, it is mentioned that the correlation between sub - blocks obtained through QT may be high or low. However, in this example, since multiple tree segmentations are supported, the correlation between sub - blocks can be determined differently.
[0298] Each of block J and block K or block L and block M can have a low correlation, and each of block J and block L or block K and block M can have a low correlation. This is because: if the correlation between the blocks adjacent to block J to block M in the horizontal and vertical directions is high, then even if QT is not executed and BT is executed, the region with high correlation may not be segmented.
[0299] In the above example, it is mentioned that the correlation between sub - blocks obtained through QT may be high or low, but this is the case of supporting a single tree method. In this example, when multiple tree methods are supported, the correlation between blocks can be checked based on the number of various cases of block segmentation.
[0300] (Basic block: the block including N, O, P)
[0301] Blocks N to P can be sub - blocks obtained through BT (one - time horizontal segmentation + one - time horizontal segmentation) (ratio is 2:1:1).
[0302] (When only symmetric BT is supported <sbt>At that time), blocks N and O may have high or low correlation. In the case of block N, since it is obtained from a higher block through BT, its correlation with the area where blocks O and P are bundled may be low. However, it cannot be said that block N has low correlation with blocks O and P. Of course, block N can have low correlation with blocks O and P. Alternatively, block N can have high correlation with block O and low correlation with block P, and block N can have low correlation with block O and high correlation with block P.
[0303] In this example, asymmetric BT can be supported <abt>If block N and block O have a high correlation, the regions of block N and block O are grouped, and an asymmetric BT horizontal split can be performed at a ratio of 3:1. However, since BT (SBT) is performed twice, the correlation between block N and block O may be low.
[0304] (Basic block: the block including Q, R, S)
[0305] Blocks Q to S can be sub-blocks obtained by TT (one horizontal split).
[0306] (When only TT is supported) Blocks Q and S may have a high or low correlation. In this example, asymmetric BT can be supported. If the correlation between block Q and block R is high, the regions of block Q and block R are grouped, and an asymmetric BT horizontal split can be performed at a ratio of 3:1. However, since TT is performed, the correlation between block Q and block R may be low.
[0307] As in the above examples, the correlation between the target block and the related blocks having a horizontal relationship with the target block can be estimated based on the supported split method and split settings. Please view the various cases regarding the correlation between blocks through the following content.
[0308] Figure 11 Parts (a) to (h) are exemplary diagrams for checking the correlation between blocks based on the split method and split settings.
[0309] Figure 11 Parts (a) to (c) are cases where QT, BT, and TT are respectively performed and only QT, BT, and TT are supported for the higher blocks. As described above, in the cases of QT and TT, the correlation between adjacent blocks (A and B or A and C) in the horizontal or vertical direction may be high or low. This is called a general relationship. Meanwhile, in the case of BT, the correlation between adjacent blocks A and B in the horizontal or vertical direction may be low. This is called a special relationship.
[0310] Figure 11 Part (d) is a case where QT is performed, and it can be a case where QT and BT are supported. In this example, the correlation between adjacent blocks (A and B or A and C) in the horizontal or vertical direction may be a low special relationship. If the correlation is high, BT may have been applied instead of QT.
[0311] Figure 11 Part (e) is the case where BT (one vertical split + one horizontal split) is performed, and it can be a case that supports both QT and BT. In the cases of A and B, since they are split from the same higher block by BT, this may be a special relationship with low correlation. In the cases of A and C, if the regions below A and C are highly correlated, they can be grouped together and splitting can be performed on them. However, due to the encoding cost, it may be the case that it is split into c. Of course, since it can be other cases, A and C can be a general relationship with high or low correlation.
[0312] Figure 11 Part (f) is the case where BT (one vertical split + one vertical split) is performed, and it can be a case that supports both BT and TT. In the cases of A and C, since they are split from the same higher block by BT, this may be a special relationship with low correlation. In the cases of A and B, there are the following situations: A and B are grouped together and splitting is performed (the part corresponding to 2 in the 1:2:1 region of TT), but additional splitting of the left region occurs due to TT. In this case, since it is difficult to accurately determine the correlation, A and B can have a general relationship with high or low correlation.
[0313] Figure 11 Part (g) is the case where TT (one vertical split) is performed, and it can be a case that supports BT (or SBT), ABT, and TT. In this example, the correlation between adjacent blocks (A and B or A and C) in the horizontal or vertical direction may be a special relationship with low correlation. If the correlation is high, ABT may have been applied instead of TT.
[0314] Figure 11 Part (h) is the case where QT and BT (two vertical splits based on BT) are performed and it can support QT and BT <1>, and it can also be a case that supports additional TT <2>. In the cases of A and B, in the situation of <1>, there is no case of splitting the basic block in the state where A and B are bound, so the correlation may be a general relationship with high or low. However, in the situation of <2>, there are the following situations: The basic block is split in the state where A and B are bound (after BT horizontal split, upper BT vertical split, and lower TT vertical split), but nevertheless, it is still split using QT and BT. Therefore, this may be a special relationship with low correlation. In this example, to check the relationship between blocks, the number of cases of block splitting that can be obtained from the same higher block with a difference of 1 or more (the difference is 2 in this example) can be checked.
[0315] Through the above various examples, it has been confirmed that the correlation between blocks is measured to use / reference the correlated blocks having a horizontal relationship with the target block. In this case, the correlated blocks that belong to the same space as the target block and are adjacent to the target block can be targeted. In particular, the target block and the correlated blocks can be blocks that are adjacent to each other in the horizontal direction or the vertical direction.
[0316] The correlation between blocks can be grasped / estimated based on various information. For example, the correlation between blocks can be examined based on the status information (size, shape, position, etc.) of the target block and the correlated blocks.
[0317] Here, as an example of determining the correlation based on the size of the block, if the predetermined length (horizontal length or vertical length) of the correlated block adjacent to the boundary (horizontal or vertical) in contact with the target block is greater than or equal to the predetermined length of the target block, the correlation between the blocks may be very high or slightly low, which is called general relationship A. If the predetermined length of the correlated block is less than the predetermined length of the target block, the correlation between the blocks may be slightly high or very low, which is called general relationship B. In this case, the horizontal lengths of each block can be compared when contacting the horizontal boundary (the upper block), and the vertical lengths of each block can be compared when contacting the vertical boundary (the left block).
[0318] Here, as an example of determining the correlation based on the shape of the block, when the target block has a rectangular shape, the correlation with the correlated block adjacent to the longer boundary of the horizontal length and the vertical length is general relationship A, and the correlation with the correlated block adjacent to the shorter boundary can be general relationship B.
[0319] The above description can be some examples of determining the correlation between blocks based on the status information of the blocks, and various modified examples are possible. The correlation between blocks can be grasped not only based on the status information of the blocks but also based on various information.
[0320] The following describes the process of examining the correlation of the correlated blocks having a horizontal relationship with the target block.
[0321] (Examination of block segmentation settings within the image)
[0322] <1> Various setting information regarding block segmentation within the image can be examined. Examine range information supported by units such as coding / prediction / transformation (assuming the target block is a coding unit in this example), such as the maximum block size, the minimum block size, etc. As an example, it can be examined that the maximum coding block is 128×128 and the minimum coding block is 8×8.
[0323] <2> Check the supported splitting schemes and check the conditions supported by each splitting scheme, such as the maximum block size, minimum block size, and maximum splitting depth. For example, the maximum size of a block supported by QT can be 128×128, the minimum size can be 16×16, the maximum sizes of blocks supported by BT and TT can be 128×128 and 64×64 respectively, the minimum size can be 4×4 jointly, and the maximum splitting depth can be 4.
[0324] <3> Check the settings to be assigned to the splitting scheme, such as the priority. For example, if splitting is performed by QT, QT can be supported for its lower blocks (sub - blocks); if splitting is not performed by QT and is performed by a different scheme, QT may not be supported for the lower blocks.
[0325] <4> When multiple splitting schemes are supported, the conditions for prohibiting certain splits can be checked to avoid overlapping according to the results of the splitting schemes. For example, after performing TT, the vertical split of BT can be prohibited for the central region. That is, in order to prevent the overlapping of splitting results that may occur according to each splitting scheme, the setting information for the prohibited splits is checked in advance.
[0326] By checking all or some of <1> to <4> and additional other setting information, the block candidates that can be obtained in the image can be checked. This can be referred to identify the block candidates that can be obtained based on the target block and related blocks to be described later.
[0327] (Check the information of the block)
[0328] The status information such as the size, shape, position, etc. of the target block and related blocks can be checked. Here, check whether the position of the block is at the boundary of units such as pictures, slices, tile groups, tiles, brick - shaped blocks, or blocks or inside.
[0329] The block in the unit can be set as the maximum coding block, and the maximum coding block can be the higher block (the highest ancestor block) of the target block, but it is the unit jointly divided in the form of picture units rather than in the form obtained according to the characteristics of the image. Therefore, if it belongs to a different maximum coding block from the maximum coding block to which the target block belongs, the correlation between the blocks cannot be checked, so it is necessary to check whether it belongs to the boundary processing.
[0330] In addition, since other units such as pictures and slices are composed of integer multiples of the maximum coding block or have settings that cannot be referred to, the process of checking the correlation can be performed only when the position of the block is not at the boundary. In other words, the correlation can be checked only when the position of the block is not at the boundary.
[0331] The dimensions, shapes, and positions of the target block and related blocks can be used to check the segmentation status of each block or to obtain the segmentation path of each block from it. This will be described in detail later.
[0332] (Check the division status and check the common higher block)
[0333] The segmentation status of the target block and related blocks can be checked. Here, the segmentation status can mean obtaining the segmentation path of each block from it. The process of checking the higher block of each block can be performed by checking the segmentation status, where the higher block can mean a block having a vertical relationship with each block. The process of checking the higher block obtained based on state information such as the dimensions, shapes, and positions of each block and the segmentation path is executed.
[0334] For example, state information such as the target block position (at the upper left) being (32, 32), width and height being 8×8, segmentation depth being p, and segmentation path being (QT / 1 - BT / h / 0 - BT / v / 1) can be obtained. State information such as the related block position being (24, 32), width and height being 8×8, segmentation depth being q, and segmentation path being (QT / 1 - BT / h / 0 - BT / v / 0) can be checked. Here, the segmentation path can be expressed as segmentation scheme / segmentation direction (h is horizontal, v is vertical, omitted if not present) / segmentation position (0 to 3 for QT, 0 to 1 for BT, etc.).
[0335] In this example, state information such as the position of the higher block (parent block, segmentation depth difference of 1) of the target block being (24, 32), width and height being 16×8, segmentation depth being p - 1, and segmentation path being (QT / 1 - BT / h / 0) can be obtained. In this example, the higher block (segmentation depth of q - 1) of the related block can be the same as the higher block of the target block.
[0336] For example, state information such as the target block position being (128, 64), width and height being 16×32, segmentation depth being p, and segmentation path being (QT / 3 - QT / 2 - BT / v / 1) can be obtained. State information such as the related block position being (120, 64), width and height being 8×32, segmentation depth being q, and segmentation path being (QT / 3 - QT / 2 - BT / v / 0 - BT / v / 1) can be obtained.
[0337] In this example, state information such as the position of the higher block (parent block, segmentation depth difference of 1) of the target block being (112, 64), width and height being 32×32, segmentation depth being p - 1, and segmentation path being (QT / 3 - QT / 2) can be obtained.
[0338] On the other hand, status information can be obtained. For example, the position of the higher block (parent block, with a segmentation depth difference of 1) of the relevant block is (112, 64), the width and height are 16×32, the segmentation depth is q - 1, and the segmentation path is (QT / 3 - QT / 2 - BT / v / 0). Status information can be obtained. For example, the position of the higher block (grandparent block, with a segmentation depth difference of 2) of the relevant block is (112, 64), the width and height are 32×32, the segmentation depth is q - 2, and the segmentation path is (QT / 3 - QT / 2). It can be seen that this is the same higher block as the higher block (parent block) of the target block.
[0339] As in the above example, based on the segmentation status, the process of checking the higher blocks of each block can be performed, and the process of checking the common higher blocks can be performed.
[0340] In summary, the higher blocks with a segmentation depth difference of 1 or greater from the target block and the relevant block can be checked. As an example, the higher block with a segmentation depth difference of c from the target block can be the same as the higher block with a segmentation depth difference of d from the relevant block. In this case, c and d can be integers of 1, 2 or greater, and c and d can be the same or different.
[0341] Here, since it is difficult to grasp the complexity or correlation, it may not be necessary to check the higher blocks with a large segmentation depth difference. For example, when the higher blocks are common in the maximum coding block, it may be difficult to check the correlation between the blocks.
[0342] For this reason, there can be a predetermined first threshold (maximum value) for c and d, and the first threshold can be an integer of 1, 2 or greater. Alternatively, there can be a predetermined second threshold related to the sum of c and d, and the second threshold can be an integer of 2, 3 or greater. That is, when the threshold condition is not satisfied, the correlation between the blocks is not checked.
[0343] There can be various methods for checking whether the target block and the relevant block have the same higher block. For example, it can be checked whether the higher blocks are the same through information about the predetermined position of the higher block or information about the width and height of the block. Specifically, it can be checked whether the higher blocks are the same based on the upper left coordinates of the higher block and information about the width and height of the block.
[0344] (Check the available candidates)
[0345] When the common higher block of the target block and the relevant block can be obtained, the number of cases of various block segmentations that can be obtained from the corresponding higher block can be checked. This can be checked based on the block segmentation settings and the segmentation status of the higher block. Since this has been mentioned in the above various examples, the detailed description will be omitted.
[0346] (Check the correlation)
[0347] To check the correlation between blocks, the following can be checked in this example. In this example, it is assumed that for each block, the maximum value of the difference in segmentation depth is 2.
[0348] <1> If the higher block has a segmentation depth difference of 1 with the target block and the related block, check the supported segmentation schemes.
[0349] When only one segmentation scheme is available, the correlation can be determined according to the segmentation scheme. If the segmentation scheme is QT or TT, the correlation can be set to a general relationship (the correlation may be high or low), and if the segmentation scheme is BT, the correlation can be set to a special relationship.
[0350] If multiple segmentation schemes are available, check whether they are segmented in a form where the target block and the related block are combined. If so, set the correlation to a special relationship, and if not, set the correlation to a general relationship.
[0351] <2> In the case where the higher block has a segmentation depth difference of 2 with at least one of the target block and the related block, check the supported segmentation schemes.
[0352] If only one segmentation scheme is available, regardless of the segmentation scheme, the correlation can be set to a general relationship.
[0353] If multiple segmentation schemes are available, check whether they are segmented in a form where the target block and the related block are combined. If so, set the correlation to a special relationship, and if not, set the correlation to a general relationship.
[0354] The above examples are some cases for checking the correlation between blocks, and are not limited to this, and various modifications and additions are possible. The information of the related block can be used / referred to for encoding the target block by referring to the correlation between the blocks checked through the above processing.
[0355] In summary, to determine the correlation between the target block and the related block, all or some of the processes such as (checking the block segmentation within the image), (checking the information of the block), (checking the segmentation status and the common higher block), (checking the available candidates), (checking the correlation), etc. can be used, and the process of determining the correlation can be executed in various orders rather than the order listed above. In addition, it is not limited to those mentioned above, and the correlation can be determined by changing some components or combining other components. Alternatively, the process of determining the correlation of other configurations can be executed, and the information of the related block can be used / referred to for encoding the target block based on the correlation between the blocks determined through the above processing.
[0356] The correlation determined through the above processing may not be an absolute fact regarding the characteristics between blocks, but rather predictive information for estimating the correlation between blocks considering block segmentation and the like. Therefore, since the correlation can be information that is referred to for constructing a candidate group of encoding information for the target block, correlated blocks with low correlation may not be included in the candidate group. Considering the possibility that the determined correlation is inaccurate, the priority for candidate group construction can be set to a lower priority, or the candidate group information of a higher block with a vertical relationship can be borrowed. Additionally, in the above example, although the correlation between blocks is assumed to be classified into two types, two, three, or more classification categories can be supported.
[0357] Whether to use / reference the correlation between blocks for encoding the target block (such as candidate group construction) can be explicitly determined in units such as sequences, pictures, slices, tile groups, tiles, brick-shaped blocks, blocks, etc., or can be implicitly determined based on the encoding settings. Next, examples of various information constituting the encoding settings will be described.
[0358] Here, whether to reference the correlation between blocks can be determined based on the information to be used / referred to for encoding the target block. For example, for constructing an intra prediction mode candidate group, the correlation between blocks can be considered; while for constructing a candidate group for representing motion vector prediction of a non-translational motion model in inter prediction, the correlation between blocks can be not considered.
[0359] Here, whether to reference the correlation between blocks can be determined based on the state information of the target block and the image information to which the target block belongs. Here, the state information of the target block can be defined based on the block size, block shape, the horizontal length / vertical length ratio of the block, and the position in units such as pictures / segmentation units (slices, tiles, etc.) / maximum coding blocks. Here, the image information to which the target block belongs can be defined based on the image type (I / P / B), color components (Y / Cb / Cr), etc. For example, the correlation between blocks can be referenced only when it has a block size belonging to a predetermined range; while when the block size is outside the predetermined range, the correlation between blocks can be not referenced. In this case, the predetermined range can be defined by a first threshold size (minimum value) and a second threshold size (maximum value), and each threshold size can be expressed as W, H, W×H, W*H based on the width (W) and height (H). W and H can be integers of 1 or greater, such as 4, 8, 16.
[0360] Here, it is possible to determine whether to refer to the relevance between blocks according to the categories of related blocks (which can be derived from the description of the positions of related blocks having a horizontal relationship). For example, related blocks can be in the same space as the target block and can be adjacent blocks. Even if related blocks are in the same space as the target block, in the case of non - adjacent related blocks, the relevance may not be referred to.
[0361] All or some of the factors mentioned in the above example can be considered to define the coding settings, so it is possible to implicitly determine whether to refer to the relevance between blocks.
[0362] The following shows a candidate group construction based on blocks having a horizontal relationship with the target block and encoding / decoding processing.
[0363] Check the relevance between the target block and the blocks where there is a possibility of reference (1). Based on the relevance, specify the blocks to be referred to for the encoding information of the target block (2). Use the specified encoding information to construct a candidate group in a predetermined order (3). Select one candidate from the candidate group based on the encoding information of the target block (4). Perform image encoding / decoding processing based on the selected candidate (5).
[0364] In (1), check the relevance between the target block and the blocks that can be considered as related blocks. In (2), based on the relevance checked in (1), specify the blocks for constructing the candidate group of the encoding information for the target block. That is to say, this may mean determining whether to include as related blocks based on the result of the checked relevance. Of course, in this example, the content describing the specifications of related blocks having a horizontal relationship as described above can be considered together.
[0365] In (3), use the encoding information of the related blocks specified by the above - mentioned processing to construct a candidate group according to a predetermined order. In this case, an adaptive order considering the related blocks included or not included in (2) can be supported. In (4), select one candidate from the candidate group based on the encoding information of the target block. In (5), perform image encoding / decoding processing based on the selected candidate.
[0366] In the flowchart, blocks determined to have low relevance based on relevance may not be included as related blocks.
[0367] The following shows another example of candidate group construction based on blocks having a horizontal relationship with the target block and encoding / decoding processing.
[0368] Specify a block (1) whose coding information is to be referred to for the target block. Check the correlation between the target block and the specified block (2). Determine a predetermined order based on the coding information of the target block and the correlation checked in (2), and accordingly construct a candidate group (3). Select one candidate from the candidate group based on the coding information of the target block (4). Perform image encoding / decoding processing based on the selected candidate (5).
[0369] In (1), specify a relevant block to be used for constructing a candidate group for the coding information of the target block. In (2), check the correlation between the target block and the relevant block. In (3), the order including candidates can be determined based on the correlation checked in (2).
[0370] For example, if the correlation is high or low, a predetermined order can be applied. If the correlation is high, an order with high priority for the relevant block can be applied, and if the correlation is low, an order with low priority for the relevant block can be applied.
[0371] Subsequently, in (3), after determining the order for constructing the candidate group through the above processing, the candidate group can be constructed according to this order. In (4), select one candidate from the candidate group based on the coding information of the target block. In (5), perform image encoding / decoding processing based on the selected candidate.
[0372] In the flowchart, the order including candidates can be adaptively set based on the correlation.
[0373] The following shows an example of candidate group construction based on blocks having a horizontal or vertical relationship with the target block and encoding / decoding processing. Here, it is assumed that the basic block for specifying the block to be referred to is set as the target block.
[0374] Check the correlation between the target block and the blocks where there is a possibility of reference (1).
[0375] (The number of blocks determined to have a low correlation with the target block is less than or equal to a predetermined number)
[0376] Specify a block (2A) whose coding information is to be referred to for the target block based on the correlation. Construct a candidate group in a predetermined order using the specified coding information (3A).
[0377] (The number of blocks determined to have a low correlation with the target block exceeds a predetermined number)
[0378] The basic block for specifying the block to be referred to is changed to a predetermined higher block (2B). Specify a block (3B) whose coding information is to be referred to for the target block based on the changed basic block. Construct a candidate group in a predetermined order using the coding information of the specified block (4B).
[0379] Select one (5) from the candidate groups based on the coding information of the target block. Perform image encoding / decoding processing (6) based on the selected candidate group.
[0380] In the flowchart, one of the orders 1-2A-3A-5-6 (P) or 1-2B-3B-4-5-6 (Q) can be determined according to the result of correlation determination. Specifically, when there are few blocks determined to have low correlation with the target block, the remaining blocks except the corresponding blocks are designated as relevant blocks; when there are many blocks determined to have low correlation with the target block, by changing the basic block of the candidate group construction to a block higher than the target block, the blocks having a horizontal relationship with the higher block are designated as relevant blocks.
[0381] In the case of the P order, since some of the flowcharts in the above flowchart in which the blocks determined to have low correlation are not included in the relevant blocks are the same, the detailed description is omitted. The Q order can be a configuration based on the combination of the blocks having a vertical relationship and the candidate group construction. It can be an example of constructing a candidate group by changing the block unit that is the basis of the candidate group construction when the blocks adjacent to the target block are composed of blocks with low correlation. In the following description, the redundant descriptions described previously are omitted and the focus is on the differences.
[0382] In (2B), the block used as the candidate group criterion is changed to the first relevant block. Here, the first relevant block can be a block having a vertical relationship with the target block (here, the higher block).
[0383] In (3B), the second relevant block for constructing the candidate group for the coding information of the target block is designated. Here, the second relevant block can be a block having a horizontal relationship with the basic block, and the basic block is the higher block. In (4B), the coding information of the second relevant block designated through the above processing is used to construct the candidate group in a predetermined order.
[0384] Here, the criterion for determining low correlation with the target block is the case of being divided by the number of blocks in the flowchart, but various criteria to be determined can be set.
[0385] Various relationships between blocks are described through the above various examples, and the cases of performing encoding / decoding using these are described. When describing the algorithms based on the relationships between the above blocks in various encoding / decoding processes to be described later, it should be understood that the settings proposed through the above various embodiments can be applied in the same or similar manner even without adding detailed descriptions.
[0386] (Inter-frame prediction)
[0387] In an image encoding method according to an embodiment of the present disclosure, inter-frame prediction may be configured as follows. Inter-frame prediction in a prediction unit may include a reference picture construction stage, a motion estimation stage, a motion compensation stage, a motion information determination stage, and a motion information encoding stage. In addition, an image encoding apparatus may be configured to include a reference picture construction unit, a motion estimation unit, a motion compensation unit, a motion information determination unit, and a motion information encoding unit that implement the reference picture construction stage, the motion estimation stage, the motion compensation stage, the motion information determination stage, and the motion information encoding stage. Some of the above-mentioned processes may be omitted or other processes may be added, and the order may be changed to an order other than the above-mentioned order.
[0388] In an image decoding method according to an embodiment of the present disclosure, inter-frame prediction may be configured as follows. Inter-frame prediction in a prediction unit may include a motion information decoding stage, a reference picture construction stage, and a motion compensation stage. In addition, an image decoding apparatus may be configured to include a motion information decoding unit, a reference picture construction unit, and a motion compensation unit that implement the motion information decoding stage, the reference picture construction stage, and the motion compensation stage. Some of the above-mentioned processes may be omitted or other processes may be added, and the order may be changed to an order other than the above-mentioned order.
[0389] Since the reference picture construction unit and the motion compensation unit of the image decoding apparatus perform the same functions as those corresponding to the configuration of the image encoding apparatus, detailed descriptions thereof are omitted, and the motion information decoding unit may be executed by reversely using the method used in the motion information encoding unit. In this case, the predicted block generated in the motion compensation unit may be sent to the addition unit.
[0390] Figure 12 is an exemplary diagram showing various cases of obtaining a predicted block through inter-frame prediction.
[0391] Referring to Figure 12 , for uni-directional prediction, a predicted block (A. forward prediction) may be obtained from a previously encoded reference picture (T-1, T-2), or a predicted block (B. backward prediction) may be obtained from a subsequently encoded reference picture (T+1, T+2). For bi-directional prediction, a predicted block (C, D) may be generated from a plurality of previously encoded reference pictures (T-2 to T+2). Generally, picture type P may support uni-directional prediction, while picture type B may support bi-directional prediction.
[0392] As in this example, a picture referred to for encoding of the current picture may be obtained from a memory. The reference picture list may be configured to include a reference picture having a temporal order or a display order before the current picture and a reference picture having a temporal order or a display order after the current picture based on the current picture (T).
[0393] Inter-frame prediction (E) can be performed in the current image and in images before or after the current image. When inter-frame prediction is performed in the current image, it can be referred to as non-directional prediction. It can be supported in picture type I or picture type P / B, and the supported picture types can be determined according to the encoding settings. Performing inter-frame prediction using the current image is different from performing inter-frame prediction using other images to utilize temporal correlation because performing inter-frame prediction using the current image generates a prediction block by using spatial correlation, but the prediction methods between them (e.g., reference pictures, motion vectors, etc.) can be the same.
[0394] In this case, it is assumed that the picture types for which inter-frame prediction can be performed are P and B, but inter-frame prediction can also be applied to various other picture types that are added or substituted. For example, a predetermined picture type may not support intra-frame prediction, may support only inter-frame prediction, may support only inter-frame prediction in a predetermined direction (backward), and may support only inter-frame prediction in a predetermined direction.
[0395] The reference picture construction unit can configure and manage the reference pictures for encoding the current picture through a reference picture list. At least one reference picture list can be configured according to the encoding settings (e.g., picture type, prediction direction, etc.), and a prediction block can be generated from the reference pictures included in the reference picture list.
[0396] For uni-directional prediction, inter-frame prediction can be performed using at least one reference picture included in reference picture list 0 (L0) or reference picture list 1 (L1). In addition, for bi-directional prediction, inter-frame prediction can be performed using at least one reference picture included in the combined list (LC) generated by combining L0 and L1.
[0397] For example, the uni-directional direction can be classified into forward prediction (Pred_L0) using the forward reference picture list (L0) and backward prediction (Pred_L1) using the backward reference picture list (L1). Bi-directional prediction (Pred_BI) can use both the forward reference picture list (L0) and the backward reference picture list (L1).
[0398] Alternatively, in bi-directional prediction, two or more forward predictions can be included by copying the forward reference picture list (L0) to the backward reference picture list (L1). Two or more backward predictions can be included in bi-directional prediction by copying the backward reference picture list (L1) to the forward reference picture list (L0).
[0399] The prediction direction can be represented by flag information indicating the corresponding direction (e.g., inter_pred_idc. Assume that this value can be adjusted by predFlagL0, predFlagL1, predFlagBI). predFlagL0 indicates whether forward prediction is performed, and predFlagL1 indicates whether backward prediction is performed. Whether bidirectional prediction is performed can be indicated by predFlagBI or by activating predFlagL0 and predFlagL1 simultaneously (e.g., when each flag is 1).
[0400] In this disclosure, the case of unidirectional prediction using forward prediction with a forward reference picture list is mainly described, but it can also be applied to other cases in the same way or by modification.
[0401] Generally, the following method can be used: determine the best reference picture for the picture to be encoded in the encoder, and explicitly send the information about the corresponding reference picture to the decoder. For this purpose, the reference picture construction unit can manage the picture list referred to for inter-frame prediction of the current picture, and set the rules for reference picture management by considering the limited memory size.
[0402] The information sent can be limited to the RPS (reference picture set), and the pictures selected from the RPS can be classified as reference pictures and stored in the memory (or DPB), and the pictures not selected from the RPS can be classified as non-reference pictures and removed from the memory after a certain period of time. A preset number of pictures (e.g., 14, 15, 16 or more pictures) can be stored in the memory, and the memory size can be set according to the level and image resolution.
[0403] Figure 13 It is an exemplary diagram of configuring a reference picture list according to an embodiment of this disclosure.
[0404] Refer to Figure 13 , generally, the reference pictures (T-1, T-2) existing before the current picture can be assigned to L0 and managed, and the reference pictures (T+1, T+2) existing after the current picture can be assigned to L1 and managed. When the allowed number of reference pictures for L0 is not reached when configuring L0, the reference pictures of L1 can be assigned. Similarly, when the allowed number of reference pictures for L1 is not reached when configuring L1, the reference pictures of L0 can be assigned.
[0405] In addition, the current picture can be included in at least one reference picture list. For example, the current picture can be included in L0 or L1, and L0 can be configured by adding a reference picture (or the current picture) with a time order of T before the current picture, and L1 can be configured by adding a reference picture with a time order of T after the current picture.
[0406] The reference picture list configuration can be determined according to the coding settings.
[0407] The current picture can be managed in a separate memory, distinct from the reference picture list, without including the current picture in the reference picture list, or the current picture can be managed by including it in at least one reference picture list.
[0408] For example, it can be determined by a signal (curr_pic_ref_enabled_flag) indicating whether the current picture is included in the reference picture list. In this case, the signal can be information that is implicitly determined or explicitly generated.
[0409] Specifically, when the signal is deactivated (e.g., curr_pic_ref_enabled_flag = 0), the current picture may not be included in all reference picture lists as a reference picture. When the signal is activated (e.g., curr_pic_ref_enabled_flag = 1), it can be implicitly determined whether the current picture is included in a predetermined reference picture list (e.g., the current picture can be added only to L0, only to L1, or to both L0 and L1 simultaneously), or it can be explicitly determined whether the current picture is included in a predetermined reference picture list by generating relevant signals (e.g., curr_pic_ref_from_l0_flag, curr_pic_ref_from_l1_flag). The signal can be supported in units such as video, sequence, picture, slice, tile group, tile, block, etc.
[0410] In this case, the current picture can be positioned at the first or last position in the reference picture list as in Figure 13 , and the arrangement order in the list can be determined according to the coding settings (e.g., picture type information, etc.). For example, in the case of an I-type picture, the current picture can be positioned at the first position, while in the case of a P / B-type picture, the current picture can be positioned at the last position, and not limited to this, other modified examples are also possible.
[0411] Alternatively, a separate reference picture memory may be supported according to a signal (ibc_enabled_flag) indicating whether block matching (or template matching) is supported in the current picture. In this case, the signal may be information that is implicitly determined or explicitly generated.
[0412] Specifically, when the signal is deactivated (e.g., ibc_enabled_flag = 0), this may mean that block matching is not supported in the current picture; when the signal is activated (e.g., ibc_enabled_flag = 1), block matching may be supported in the current picture, and its reference picture memory may be supported. In this example, it is assumed that an additional memory is set, but without setting an additional memory, the current picture may also be set to directly support block matching in the existing memory supported by the current picture.
[0413] The reference picture construction unit may include a reference picture interpolation unit, and whether to perform interpolation processing on pixels in fractional units may be determined according to the interpolation accuracy of inter-frame prediction. For example, when the interpolation accuracy is in integer units, the reference picture interpolation processing may be omitted, while when the interpolation accuracy is in fractional units, the reference picture interpolation processing may be performed.
[0414] The interpolation filter used in the reference picture interpolation processing may be determined according to the encoding settings, and one preset interpolation filter (e.g., DCT-IF (Discrete Cosine Transform-based Interpolation Filter), etc.) may be used, or one of multiple interpolation filters may be used. For the former, the selection information about the interpolation filter may be implicitly omitted, while for the latter, the selection information about the interpolation filter may be included in units such as video, sequence, picture, slice, tile group, tile, brick-shaped block, etc. For the latter, the information about the interpolation filter (e.g., filter coefficient information, etc.) may also be information that can be explicitly generated.
[0415] The same type of filter may be used according to the interpolation position (e.g., fractional units such as 1 / 2, 1 / 4, 1 / 8). For example, the filter coefficient information may be obtained from one filter formula according to the interpolation position. Alternatively, different types of interpolation filters may be used according to the interpolation position. For example, a 6-tap Wiener filter may be applied to the 1 / 2 unit, an 8-tap Kalman filter may be applied to the 1 / 4 unit, and a linear filter may be applied to the 1 / 8 unit.
[0416] The interpolation precision (the maximum precision for performing interpolation) can be determined according to the coding settings, and the interpolation precision can be a precision in integer units or decimal units (e.g., 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, etc.). In this case, the interpolation precision can be determined according to the image type, reference picture settings, supported inter-frame prediction methods, etc.
[0417] For example, the interpolation precision can be set to integer units in image type I, and the interpolation precision can be set to decimal units in image types P / B. Alternatively, when the current picture is included in the reference pictures, the interpolation precision can be set to one of integer units or decimal units according to the pictures to be referenced. Alternatively, when block matching or template matching is supported for the current picture, the interpolation precision can be set to one of integer units or decimal units, otherwise the interpolation precision can be set to decimal units.
[0418] Alternatively, the interpolation process can be performed by selecting one of multiple interpolation precisions. When the interpolation process according to the adaptive interpolation precision is supported (e.g., when adaptive_ref_resolution_enabled_flag is 0, a preset interpolation precision is used; when adaptive_ref_resolution_enabled_flag is 1, one of multiple interpolation precisions is used), precision selection information (e.g., ref_resolution_idx) can be generated.
[0419] The settings and information related to the interpolation precision (e.g., whether adaptive interpolation precision is supported, precision selection information, etc.) can be implicitly determined or explicitly generated, and the settings and information can be included in units such as videos, sequences, pictures, slices, tile groups, tiles, and brick-shaped blocks. Alternatively, whether adaptive interpolation precision, precision selection information, precision candidate groups, etc. are supported can be determined based on the coding settings defined by one or more coding elements such as image type, reference picture settings, supported intra-frame prediction methods, and supported motion models.
[0420] The motion estimation and compensation processes can be performed according to the interpolation precision, and the representation unit and storage unit of the motion vector can also be determined based on the interpolation precision.
[0421] For example, when the interpolation precision is 1 / 2 unit, the motion estimation and compensation processes can be performed in 1 / 2 unit, the motion vector can be represented in 1 / 2 unit, and the motion vector can be used in the coding process. In addition, the motion vector can be stored in 1 / 2 unit and the motion vector can be referenced in the motion information coding process of other blocks.
[0422] Alternatively, when the interpolation precision is 1 / 8 unit, motion estimation and compensation processing can be performed in 1 / 8 unit, the motion vector can be represented in 1 / 8 unit, the motion vector can be used in the encoding process in 1 / 8 unit, and the motion vector can be stored in 1 / 8 unit.
[0423] In addition, motion estimation and compensation processing and the motion vector can be performed, represented, and stored in units different from the interpolation precision, such as 1 / 2 unit, 1 / 4 unit, etc., and motion estimation and compensation processing and the motion vector can be adaptively determined according to the inter-frame prediction method / settings (e.g., motion estimation / compensation method, motion model, motion information encoding mode <content mentioned later>, etc.).
[0424] In the example, when the interpolation precision is assumed to be 1 / 8 unit, for the translational motion model, motion estimation and compensation processing can be performed in 1 / 4 unit, and the motion vector can be represented in 1 / 4 unit (this example assumes the unit in the encoding process) and stored in 1 / 8 unit. For the non-translational motion model, motion estimation and compensation processing can be performed in 1 / 8 unit, and the motion vector can be represented in 1 / 4 unit and stored in 1 / 8 unit.
[0425] In the example, when the interpolation precision is assumed to be 1 / 8 unit, for block matching, motion estimation and compensation processing can be performed in 1 / 4 unit, and the motion vector can be represented in 1 / 4 unit and stored in 1 / 8 unit. For template matching, motion estimation and compensation processing can be performed in 1 / 8 unit, and the motion vector can be represented in 1 / 8 unit and stored in 1 / 8 unit.
[0426] In the example, when the interpolation precision is assumed to be 1 / 16 unit, for the competition mode, motion estimation and compensation processing can be performed in 1 / 4 unit, and the motion vector can be represented in 1 / 4 unit and stored in 1 / 16 unit. For the merge mode, motion estimation and compensation processing can be performed in 1 / 8 unit, and the motion vector can be represented in 1 / 4 unit and stored in 1 / 16 unit. For the skip mode, motion estimation and compensation processing can be performed in 1 / 16 unit, and the motion vector can be represented in 1 / 4 unit and stored in 1 / 16 unit.
[0427] In summary, the motion estimation and compensation, and the representation and storage units of motion vectors can be adaptively determined based on the inter-frame prediction method or settings and interpolation accuracy. Specifically, the motion estimation and compensation, and the representation unit of motion vectors are usually determined adaptively based on the inter-frame prediction method or settings, and the storage unit of motion vectors is determined based on the interpolation accuracy, but this is not limited thereto, and various modified examples are possible. In addition, this example is based on one category (e.g., motion model, motion estimation / compensation method, etc.), but it is also possible to mix two or more categories and determine the settings.
[0428] In addition, as described above, the interpolation accuracy information has a preset value or is selected as one of multiple accuracies. Conversely, the reference picture interpolation accuracy can be determined according to the motion estimation and compensation settings supported by the inter-frame prediction method or settings. For example, when the translational motion model supports up to 1 / 8 unit and the non-translational motion model supports up to 1 / 16 unit, the interpolation process can be performed according to the accuracy unit of the non-translational motion model with the highest accuracy.
[0429] In other words, the reference picture interpolation can be performed according to the settings of the accuracy information supported for the translational motion model, non-translational motion model, competition mode, merge mode, skip mode, etc. In this case, the accuracy information can be determined implicitly or explicitly, and when the relevant information is explicitly generated, the accuracy information can be included in units such as video, sequence, picture, slice, tile group, tile, brick-shaped block, etc.
[0430] The motion estimation unit performs the process of estimating (or searching) which block of which reference picture has a high correlation with the target block. The size and shape (M×N) of the target block for performing prediction can be obtained from the block partitioning unit. In the example, the target block can be determined in the range of 4×4 to 128×128. Generally, inter-frame prediction is performed in units of prediction blocks, but inter-frame prediction can be performed in units of coding blocks, transform blocks, etc. according to the settings of the block partitioning unit. When performing estimation within the available range of the reference area, at least one motion estimation method can be used. The estimation order and conditions, etc. can be defined in units of pixels in the motion estimation method.
[0431] Motion estimation can be performed based on the motion estimation method. For example, in the case of block matching, the area to be compared for the motion estimation process can be the target block; in the case of template matching, the area to be compared for the motion estimation process can be a predetermined area (template) based on the basic block settings. For the former, the block with the highest correlation with the target block can be found within the available range of the reference area, and for the latter, the area with the highest correlation with the template defined according to the coding settings can be found within the available range of the reference area.
[0432] In this case, templates can be set in one or more neighboring blocks such as the left, upper, upper-left, upper-right, lower-left blocks, etc. based on a basic block. The neighboring blocks can be blocks that have already been encoded. In an example, when the basic block is M×N, M×h and v×N can be configured as templates at the upper and left of the basic block respectively. In this case, the template can have settings such as a predetermined fixed area (e.g., the left block and the upper block) and length (w, h), or can have an adaptive setting according to the encoding. In this case, the encoding settings can be defined by the size, shape, position, aspect ratio, image type, color components, etc. of the basic block. Alternatively, information about the template area, length, etc. can be explicitly generated in units such as video, sequence, picture, slice, tile group, tile, brick-shaped block, etc., or can be implicitly determined according to the encoding settings.
[0433] In this case, block matching can be one of the methods for explicitly generating all or some of the motion information, and template matching can be one of the methods for implicitly obtaining all or some of the motion information. The motion information (or type of motion information) explicitly generated or implicitly obtained in the motion estimation method can be determined by the inter-frame prediction settings, and the inter-frame prediction settings can be defined by a motion model, a motion information coding mode, etc. In this case, for template matching, information based on the estimated start position, modified information (related to the x and y vectors) of the motion vector that completes the final estimation, etc. can be implicitly determined, but the related information can be explicitly generated.
[0434] Information about the support range of template matching can be explicitly generated or implicitly determined according to the encoding settings. In this case, the encoding settings can be defined by one or more factors such as the size, shape, position, image type, color components, etc. of the target block. In an example, template matching can be supported in the range from A×B to C×D, where A to D can be integers such as 4, 8, 16, 32, 64 or larger, and A and B can be less than C and D respectively or the same as C and D. In addition, the support range of template matching can belong to the support range of block matching, or an exception configuration (e.g., when the minimum size of the block is small) is enabled.
[0435] When multiple motion estimation methods are supported (e.g., when adaptive_motion_comp_enabled_flag is 0, a preset motion estimation method is used, and when adaptive_motion_comp_enabled_flag is 1, one of multiple motion estimation methods is used), motion estimation method selection information can be generated and can be included in units of blocks. For example, when motion_comp_idx is 0 in a predetermined division unit such as a picture, a slice, etc., only block matching is supported; while when motion_comp_idx is 1, both block matching and template matching are supported. Alternatively, when motion_comp_idx is 0 in units of blocks, block matching is supported; while when motion_comp_idx is 1 in units of blocks, template matching is supported. As described above, multiple related selection information can be generated according to units, but the candidates indicated by the corresponding information can have different meanings.
[0436] The above example can be a configuration where the classification (selection) of the motion estimation method in units of blocks is prior. In the example, when template matching is selected, there may be no information or merge mode to be further confirmed, and the competition mode can be supported as a candidate for the motion information coding mode. In this case, when the merge mode is selected, the motion vector obtained by the final estimation can be set as the motion vector of the target block; when the competition mode is selected, offset information for modifying the motion vector obtained by the final estimation in the horizontal or vertical direction is additionally generated, and the motion vector obtained by adding the offset information to the motion vector obtained by the final estimation can be set as the motion vector of the target block.
[0437] Alternatively, template matching can be supported by being included as some of the inter-frame prediction configurations. In the example, the motion information coding mode for template matching can be supported, and the motion information coding mode for template matching can be supported by being included as a candidate in a group of motion information prediction candidates configured in a predetermined motion information coding mode. For the former, it can be a configuration where template matching is performed by the classification (selection) of the motion information coding mode, while for the latter, it can be a configuration where template matching is performed by selecting a candidate in a group of motion information prediction candidates for representing the best motion information in a predetermined motion information coding mode.
[0438] The template in the above example can be set based on a basic block, and the basic block can be an encoding block or a prediction block (or a transform block). It is described that when an encoding block is determined by a block partitioning unit, the encoding block can be set as a prediction block as it is, or the encoding block can be partitioned into two or more prediction blocks. In this case, when the encoding block is not partitioned or is partitioned into two or more prediction blocks (e.g., a non-square or right-angled triangle) to perform inter-frame prediction, the encoding block is referred to as inter-frame prediction in units of sub-blocks. In other words, the basic block can be set as an encoding block or a prediction block, and the template can be set based on the basic block.
[0439] In addition to the above units, basic blocks for template setting can also be supported. In the example, blocks in a vertical relationship or a horizontal relationship with the basic block can be targets.
[0440] Specifically, it can be assumed that the encoding block is the target block, and the higher block in a vertical relationship with the encoding block is the related block <1>. Alternatively, it can be assumed that the encoding block is the target block, and the block in a horizontal relationship with the encoding block is the related block. In this case, the encoding block can be changed to a prediction block and applied. In addition, it can be assumed that the prediction block is the target block, and the encoding block is the related block. In the examples mentioned later, the case in <1> is assumed.
[0441] The method of configuring a template by setting a basic block among multiple candidate blocks can be explicitly determined by information on whether it is supported (for 0, not supported; for 1, supported) in units such as sequence, picture, slice, tile group, tile, brick-shaped block, block, etc., and when the corresponding information is not confirmed, a predetermined value (0 or 1) can be assigned to it. Alternatively, it can be implicitly determined whether it is supported, and it can be determined whether it is supported based on the encoding settings. In this case, the encoding settings can be defined by one or more elements of status information such as the size, shape, position, etc. of a block (target block), image type (I / P / B), color component, whether inter-frame prediction in units of sub-blocks is applied, etc.
[0442] For example, when the size of the target block is greater than or the same as a predetermined first threshold, a method of setting a basic block among multiple candidate blocks can be supported. Alternatively, when the size of the target block is less than or equal to a predetermined second threshold, this method can be supported. In this case, the threshold size can be represented by width (W) and height (H) as W, H, W×H, W*H, and the threshold size is represented as a pre-promised value in the encoding / decoding device. W and H can be integers equal to or greater than 1 such as 4, 8, 16. When the threshold size is represented as the sum of the width and height, W*H can be an integer such as 16, 32, 64 or greater. The first threshold size is less than or equal to the second threshold size. In this case, when this method is not supported, it means that the basic block is set to a predetermined block (target block).
[0443] When multiple candidate blocks are supported, the candidate blocks (related blocks) can be defined differently. (When the coding block is the target block), the related block (e.g., the higher block) can be a block whose partitioning depth is 1 or more less than the partitioning depth of the coding block. Alternatively, it can be a block having a predetermined width (C) and height (D) at a predetermined upper-left coordinate (e.g., to the left or above the upper-left coordinate of the target block). In this case, C and D can be integers such as 8, 16, 32 or greater, and can be greater than or equal to the width and height of the target block. In addition, C and D can be determined based on information about the block size (e.g., the maximum size of the transform block, the maximum size of the coding block, etc.).
[0444] When multiple candidate blocks are supported, candidate selection information can be explicitly generated, and the corresponding candidate block can be set as the basic block. Alternatively, the basic block can be implicitly determined and it can be based on the coding settings.
[0445] For example, when it is determined that the coding setting (first category) supports multiple candidate blocks, the related block can be set as the basic block, otherwise when it is determined that the coding setting (second category) supports multiple candidate blocks, the target block can be set as the basic block.
[0446] Alternatively, when the coding setting belongs to the first category, the target block can be set as the basic block, when the coding setting belongs to the second category, the first related block (e.g., the higher block) can be set as the basic block, and when the coding setting belongs to the third category, the second related block (e.g., the neighboring block) can be set as the basic block. In addition, when some categories are replaced or added, information for selecting one of the target block and the related blocks can be generated.
[0447] The coding settings can be defined by one or more elements of status information such as the size, shape, aspect ratio, position, etc. of the target block, image type, color component, and whether inter-frame prediction on a sub-block basis is applied.
[0448] In summary, the basic block can be set among various candidate blocks including the target block, and the basic block can be determined explicitly or implicitly. This example describes that the basic block for template setting can be set differently, but the basic block can be applied to various situations for inter-frame prediction. In other words, this description can be equally or similarly applied to the setting of the basic block for the configuration of the motion information prediction candidate group and the like in the inter-frame prediction mentioned later. However, it should be understood that the limitations on the basic block or the support range, etc. in this example can be set equally or differently.
[0449] Motion estimation can be performed based on a motion model. Motion estimation and compensation can be performed by using an additional motion model other than the translational motion model that only considers parallel translation. For example, motion estimation and compensation can be performed by using a motion model that considers motions such as rotation, perspective, zoom-in / zoom-out, etc. as well as parallel translation. By generating a prediction block that reflects various types of motions generated according to the regional characteristics of the image, it can be supported to improve the coding performance.
[0450] Figure 14 is a conceptual diagram showing a non-translational motion model according to an embodiment of the present disclosure. Refer to Figure 14 , as some examples of the affine model, an example is described in which the motion information is represented based on the motion vectors V0 and V1 at a predetermined position. Since the motion can be represented based on multiple motion vectors, accurate motion estimation and compensation are possible.
[0451] As in the example, inter-frame prediction can be performed based on a predefined motion model, but inter-frame prediction based on an additional motion model can also be supported. In this case, it is assumed that the predefined motion model is a translational motion model, and the additional motion model is an affine model, but it is not limited thereto, and various modifications are possible.
[0452] The translational motion model can represent the motion information based on one motion vector (assuming unidirectional prediction), and it is assumed that the control point (base point) for representing the motion information is the upper left coordinate, but it is not limited thereto.
[0453] The non-translational motion model can be represented by various configurations of motion information. This example assumes a configuration represented by additional information (based on the upper left coordinate) in addition to one motion vector. Some of the motion estimation and compensation mentioned in the later examples may not be performed in units of blocks, but some of the motion estimation and compensation mentioned in the later examples can be performed in units of predetermined sub-blocks. In this case, the predetermined size and position of the sub-block can be determined based on each motion model.
[0454] Figure 15 It is an exemplary diagram showing motion prediction in units of sub - blocks according to an embodiment of the present disclosure. Specifically, motion prediction in units of sub - blocks according to an affine model (two motion vectors) is described.
[0455] For a translational motion model, the motion vector in units of pixels included in a target block can be the same. In other words, it can have a motion vector that is equally applied in units of pixels, and motion estimation and compensation are performed by using one motion vector (V0).
[0456] For a non - translational motion model (affine model), the motion vectors in units of pixels included in a target block may not be the same, and separate motion vectors in units of pixels may be required. In this case, motion vectors in units of pixels or in units of sub - blocks can be derived based on the motion vectors (V0, V1) at predetermined control point positions of the target block, and motion estimation and compensation can be performed by using the derived motion vectors.
[0457] For example, it can be derived by the equation according to V x =(V 1x -V 0x )×x / M-(V 1y -V 0y )×y / N+V 0x 、V y =(V 1y -V 0y )×x / M+(V 1x -V 0x )×y / N+V 0y to obtain motion vectors in units of sub - blocks or in units of pixels in a target block {for example, (V x ,V y )}. In this equation, V0 {in this example, (V 0X ,V 0Y )} is the motion vector at the upper left of the target block, and V1 {in this example, (V 1X ,V 1Y )} is the motion vector at the upper right of the target block. Considering the complexity, motion estimation and motion compensation of the non - translational motion model can be performed in units of sub - blocks.
[0458] In this case, the size (M×N) of the sub-block can be determined according to the coding settings, and the size (M×N) of the sub-block can have a fixed size or can be set to an adaptive size. In this case, M and N can be integers such as 2, 4, 8, 16 or larger, and M and N can be the same or different. The size of the sub-block can be explicitly generated in units such as sequences, pictures, slices, tile groups, tiles, brick-shaped blocks, etc. Alternatively, the size of the sub-block can be implicitly determined by a common commitment between the encoder and the decoder, or the size of the sub-block can be determined by the coding settings.
[0459] In this case, the coding settings can be defined by one or more elements of state information such as the size, shape, position, etc. of the target block, image type, color component, inter-frame prediction setting information (e.g., motion information coding mode, reference picture information, interpolation accuracy, motion model <type>, etc.).
[0460] The above example describes the following process: obtaining the size of the sub-block according to some non-translational motion models, and performing motion estimation and compensation based on the size of the sub-block. As in this example, motion estimation and compensation can be performed in units of sub-blocks or pixels according to the motion model, and its detailed description is omitted.
[0461] Next, various examples regarding the motion information configured according to the motion model will be described.
[0462] In the example, the motion model representing rotational motion can represent the translational motion of a block with one motion vector, and can represent the rotational motion with rotation angle information. The rotation angle information (0 degrees) can be measured based on a predetermined position (e.g., the upper left coordinate), and the rotation angle information can be represented by k (k is an integer such as 1, 2, 3 or larger) candidates having a predetermined interval (e.g., the angle difference is 0 degrees, 11.25 degrees, 22.25 degrees, etc.) within a predetermined angle range (e.g., between -90 degrees and 90 degrees).
[0463] In this case, the rotation angle information can be encoded as it is in the motion information coding process, or can be encoded based on the motion information of neighboring blocks (e.g., motion vector, rotation angle information) (e.g., prediction + difference information).
[0464] Alternatively, the translational motion of the block can be represented by one motion vector, and the rotational motion of the block can be represented by one or more additional motion vectors. In this case, the number of additional motion vectors can be an integer such as 1, 2 or larger, and the control points of the additional motion vectors can be determined in the upper right, lower left or lower right coordinates, or other coordinates in the block can be set as the control points.
[0465] In this case, the additional motion vector can be encoded as is in the motion information encoding process, or can be encoded based on the motion information of neighboring blocks (e.g., the motion vector according to a translational motion model or a non-translational motion model) (e.g., prediction + difference information), or can be encoded based on other motion vectors representing rotational motion of the block (e.g., prediction + difference information).
[0466] In an example, for a motion model representing a size adjustment or scaling motion such as a zoom-in / zoom-out situation, it can represent the translational motion of a block having one motion vector, and can represent the size adjustment motion having scaling information. The scaling information can be represented by scaling information indicating expansion or contraction in the horizontal direction or the vertical direction based on a predetermined position (e.g., the upper left coordinate).
[0467] In this case, scaling can be applied in at least one of the horizontal direction and the vertical direction. In addition, the scaling information applied in each of the horizontal direction and the vertical direction can be supported, or the scaling information applied in common can be supported. The position of motion estimation and compensation can be determined by adding the width and height of the scaled block to a predetermined position (the upper left coordinate).
[0468] In this case, the scaling information can be encoded as is in the motion information encoding process, or can be encoded based on the motion information of neighboring blocks (e.g., the motion vector, the scaling information) (e.g., prediction + difference information).
[0469] Alternatively, the translational motion of the block can be represented by one motion vector, and the size adjustment of the block can be represented by one or more additional motion vectors. In this case, the number of additional motion vectors can be an integer such as 1, 2, or greater, and the control point of the additional motion vector can be determined in the upper right, lower left, or lower right coordinate, or other coordinates in the block can be set as the control point.
[0470] In this case, the additional motion vector can be encoded as is in the motion information encoding process, or can be encoded based on the motion information of neighboring blocks (e.g., the motion vector according to a translational motion model or a non-translational motion model) (e.g., prediction + difference information), or can be encoded based on a predetermined coordinate in the block (e.g., the lower right coordinate) (e.g., prediction + difference information).
[0471] The above examples describe the situation regarding the representation for representing some motions, and it can be represented as motion information for representing multiple motions.
[0472] For example, for a motion model representing various motions or complex motions, it can represent the translational motion of a block with a motion vector, represent the rotational motion with rotation angle information, and represent the size adjustment with scaling information. Descriptions regarding each motion can be derived from the examples mentioned above, and thus detailed descriptions are omitted.
[0473] Alternatively, it can represent the translational motion of a block with a motion vector and represent other motions of the block with one or more additional motion vectors. In this case, the number of additional motion vectors can be an integer such as 1, 2, or greater, and the control points of the additional motion vectors can be determined in the upper right, lower left, or lower right coordinates, or other coordinates in the block can be set as the control points.
[0474] In this case, the additional motion vectors can be encoded as they are in the motion information encoding process, or can be encoded based on the motion information of neighboring blocks (e.g., according to the motion vectors of the translational motion model or non - translational motion model) (e.g., prediction + difference information), or can be encoded based on other motion vectors representing various motions of the block (e.g., prediction + difference information).
[0475] The description can be regarding the affine model, and mainly describes the case where the number of additional motion vectors is 1 or 2. In summary, the number of motion vectors used according to the motion model can be 1, 2, 3, and it can be regarded as a separate motion model assuming according to the number of motion vectors used to represent the motion information. Additionally, when the number of motion vectors is 1, it is assumed to be a predefined motion model.
[0476] Multiple motion models for inter - frame prediction can be supported and can be determined by a signal indicating support for additional motion models (e.g., adaptive_motion_mode_enabled_flag). In this case, when the signal is 0, a predefined motion model can be supported; when the signal is 1, multiple motion models can be supported. The signal can be generated in units such as video, sequence, picture, slice, tile group, tile, brick - shaped block, block, etc., but when it is not possible to confirm the signal separately, the value of the signal can be assigned according to a predetermined setting. Alternatively, it can be implicitly determined whether to support based on the encoding settings. Alternatively, it can be determined whether it is an implicit or explicit case according to the encoding settings. In this case, the encoding settings can be defined by one or more factors such as image type, image category (e.g., for 0, general image; for 1, 360 - degree image), color component, etc.
[0477] It is possible to determine whether multiple motion models are supported in the above processing. Next, assuming that two or more motion models are additionally supported, and it is determined to support multiple motion models in units such as sequences, pictures, slices, tile groups, tiles, brick-shaped blocks, etc., but there may be some exception configurations. In the examples mentioned later, it is assumed that motion models A, B, and C can be supported, and A is the substantially supported motion model, while B and C are the motion models that can be additionally supported.
[0478] Configuration information regarding the supported motion models can be generated in the above units. In other words, supported motion model configurations such as {A, B}, {A, C}, {A, B, C} are possible.
[0479] For example, indices (0 to 2) can be assigned to candidates of the configuration and selected. When index 2 is selected, it can be determined that the motion model configuration supporting {A, C} is supported, and when index 3 is selected, it can be determined that the motion model configuration supporting {A, B, C} is supported.
[0480] Alternatively, information indicating whether a predetermined motion model is supported can be supported individually. In other words, a flag indicating whether B is supported or a flag indicating whether C is supported can be generated. When both flags are 0, it may be the case that only A is supported. This example can be an example processed without generating information indicating whether multiple motion models are supported.
[0481] As in the above examples, when a candidate group of supported motion models is configured, one motion model of the candidate group can be explicitly determined and used, or can be implicitly used in units of blocks.
[0482] Generally, the motion estimation unit can be a configuration existing in the encoding device, but it can be a configuration that can be included in the decoding device according to a prediction method (e.g., template matching, etc.). For example, for template matching, this is because motion information of the target block can be obtained by performing motion prediction based on neighboring templates of the target block in the decoder. In this case, information related to motion estimation (e.g., motion estimation range, motion estimation method <scan order>, etc.) can be implicitly determined or explicitly generated, and can be included in units such as videos, sequences, pictures, slices, tile groups, tiles, brick-shaped blocks, etc.
[0483] The motion compensation unit performs processing for obtaining data of some blocks of some reference pictures of the prediction block determined as the target block in the motion estimation processing. Specifically, based on the motion information (e.g., reference picture information, motion vector information, etc.) obtained in the motion estimation processing, a prediction block of the target block can be generated from at least one region (or block) of at least one reference picture.
[0484] Motion compensation can be performed based on the following motion compensation methods.
[0485] For block matching, based on the motion vector (V x , V y ) of the target block (M×N) explicitly obtained in the reference picture and the upper left coordinates (P x , P y ) of the target block, the coordinates (P x + V x , P y + V y ) are obtained. The data in the area corresponding to the right M and the lower N can be compensated as the predicted block of the target block.
[0486] For template matching, based on the motion vector (V x , V y ) of the target block (M×N) implicitly obtained in the reference picture and the upper left coordinates (P x , P y ) of the target block, the coordinates (P x + V x , P y + V y ) are obtained. The data in the area corresponding to the right M and the lower N can be compensated as the predicted block of the target block.
[0487] In addition, motion compensation can be performed based on the following motion models.
[0488] For the translational motion model, based on one motion vector (V x , V y ) of the target block (M×N) explicitly obtained in the reference picture and the upper left coordinates (P x , P y ) of the target block, the coordinates (P x + V x , P y + V y ) are obtained. The data in the area corresponding to the right M and the lower N can be compensated as the predicted block of the target block.
[0489] For the non-translational motion model, based on multiple motion vectors (V 0x , V 0y ), (V 1x , V 1y ) of the target block (M×N) explicitly obtained in the reference picture, the motion vectors (V mx , V my ) of the m×n sub-blocks implicitly obtained, and the upper left coordinates (P mx , P ny ) of each sub-block, the coordinates (P mx +V nx ,P my +V ny The data in the region corresponding to the right M / m and the lower N / n can be compensated with the predicted block of the target block. In other words, the predicted block of the target block can be utilized for compensation by collecting the predicted blocks of the sub-blocks.
[0490] In the motion information determination unit, a process for selecting the best motion information of the target block can be performed. Generally, the best mode information can be determined from the perspective of encoding cost by using block distortion {e.g., the distortion between the target block and the reconstructed block, SAD (Sum of Absolute Differences), SSD (Sum of Squared Differences), etc.} and a rate-distortion method that takes into account the amount of bits generated according to the corresponding motion information. The predicted block generated based on the motion information determined in the above process can be sent to the subtraction unit and the addition unit. Additionally, it can be a configuration that can be included in the decoding device according to some prediction methods (e.g., template matching, etc.), and in this case, it can be determined based on block distortion.
[0491] For the motion information determination unit, setting information related to inter-frame prediction such as motion compensation methods, motion models, etc. can be considered. For example, when multiple motion compensation methods are supported, the motion compensation method selection information and the obtained motion vectors, reference picture information, etc. can be the best motion information. Alternatively, when multiple motion models are supported, the motion model selection information and the obtained motion vectors, reference picture information, etc. can be the best motion information.
[0492] In the motion information encoding unit, the motion information of the target block obtained in the motion information determination process can be encoded. In this case, the motion information can be configured with information about the image and region that are referred to for the prediction of the target block. Specifically, the motion information can be configured with information about the referred image (e.g., reference image information, etc.) and information about the referred region (e.g., motion vector information, etc.).
[0493] In addition, the setting information related to inter-frame prediction (or selection information, etc., e.g., motion estimation / compensation method, selection information of motion models, etc.) can also be included in the motion information of the target block. Information about the referred image and region (e.g., the number of motion vectors, etc.) can be configured based on the settings related to inter-frame prediction.
[0494] The motion information can be encoded by configuring the information about the referred image and the referred region as a combination, and the combination of the information about the referred image and the referred region can be configured as the motion information encoding mode.
[0495] In this case, information about the reference image and the reference region can be obtained based on neighboring blocks or predetermined information (e.g., an image encoded before or after the current picture, a zero motion vector, etc.), and the neighboring blocks can be blocks having a horizontal relationship with the basic block (related blocks). In other words, when the categories are classified as <inter_blk_A> of the block that belongs to the same space as the basic block and is closest to the target block, <inter_blk_B> of the block that belongs to the same space as the basic block and is far from the target block, and <inter_blk_C> of the block that does not belong to the same space as the basic block, the blocks belonging to one or more of these categories can be designated as related blocks.
[0496] For example, the motion information of the target block can be encoded based on the motion information or reference picture information of the related blocks, and the motion information of the target block can be encoded based on the information derived from the motion information or reference picture information of the related blocks (or information through intermediate values, transformation processing, etc.). In other words, for the motion information of the target block, prediction can be performed from the neighboring blocks to encode the information thereon.
[0497] The motion information of the target block can be predicted and encoded, or the motion information itself can be encoded, and it can be based on a signal indicating whether the prediction of the motion information has been performed (e.g., when mvp_enabled_flag is 0, the motion information is encoded as it is; while when mvp_enabled_flag is 1, the motion information is predicted and encoded. In other words, only for 1, the motion information encoding modes mentioned later, such as skip mode, merge mode, competition mode, etc., can be used). In the present disclosure, the description is made under the assumption that the signal is 1. In other words, in the examples mentioned later, it is assumed that all or some of the motion information of the target block is encoded based on prediction.
[0498] In the above description, the basic block can be set in the target block and the blocks having a horizontal or vertical relationship with the target block. Specifically, this means that the following blocks can be set differently, which become the standard when the motion information of the target block is encoded based on the reference picture information or motion information of the related blocks (through prediction). Since its content can be derived from the various examples mentioned above, the detailed description is omitted.
[0499] In summary, the target block can be a block that is a related party having motion information to be encoded, and the basic block can be a block that becomes the standard when a candidate group for motion information prediction is configured (e.g., a block that becomes the standard when neighboring blocks in the left and upper directions are designated). In this case, the basic block can be set as the target block, or can be set as a related block (a block in a vertical / horizontal relationship). The basic block mentioned in the examples mentioned later can be derived from the various examples mentioned above including this example in the present disclosure.
[0500] In the present disclosure, motion information of a target block may be encoded based on one or more motion information encoding modes. In this case, the motion information encoding modes may be defined differently and may include one or more of a skip mode, a merge mode, a competition mode (Comp mode), etc.
[0501] It may be combined with the motion information encoding mode based on the template matching (tmp) mentioned above, or may be supported as a separate motion information encoding mode, or may be included in all or some of the detailed configurations of the motion information encoding mode. On the premise that it is determined to support the template matching in a higher unit (e.g., a picture, a slice, etc.), but the flag regarding whether to support may be regarded as a partial element for the inter-frame prediction setting.
[0502] It may be combined with the motion information encoding mode based on the method of performing block matching in the current picture (ibc) mentioned above, or may be supported as a separate motion information encoding mode, or may be included in all or some of the detailed configurations of the motion information encoding mode. On the premise that it is determined to support the block matching in a higher unit for the current picture, but the flag regarding whether to support may be regarded as a partial element for the inter-frame prediction setting.
[0503] It may be combined with the motion information encoding mode based on the motion model (affine) mentioned above, or may be supported as a separate motion information encoding mode, or may be included in all or some of the detailed configurations of the motion information encoding mode. On the premise that it is determined to support the non-translational motion model in a higher unit, but the flag regarding whether to support may be regarded as a partial element for the inter-frame prediction setting.
[0504] For example, a separate motion information encoding mode such as temp_inter, temp_tmp, temp_ibc, temp_affine may be supported. Alternatively, a combined motion information encoding mode such as temp_inter_tmp, temp_inter_ibc, temp_inter_affine, temp_inter_tmp_ibc, etc. may be supported. Alternatively, it may be configured by including a template-based candidate, a candidate based on the method of performing block matching in the current picture, and an affine-based candidate in the motion information prediction candidate group configuration temp.
[0505] In this case, temp can represent a skip mode, a merge mode, or a competition mode. In the example, for the skip mode, motion information coding modes such as skip_inter, skip_tmp, skip_ibc, and skip_affine can be supported; for the merge mode, motion information coding modes such as merge_inter, merge_tmp, merge_ibc, and merge_affine can be supported; and for the competition mode, motion information coding modes such as comp_inter, comp_tmp, comp_ibc, and comp_affine can be supported.
[0506] When candidates that support the skip mode, the merge mode, and the competition mode and consider the above elements are included in the motion information prediction candidate group for each mode, one mode can be selected by identifying the flags for the skip mode, the merge mode, and the competition mode. In the example, when a flag indicating whether it is the skip mode is supported and the value of the flag is 1, the skip mode can be selected; when a flag indicating whether it is the merge mode when the value is 0 is supported and the value of the flag is 1, the merge mode can be selected; when the value of the flag is 0, the competition mode can be selected. Moreover, candidates based on inter, tmp, ibc, and affine can be included in the motion information prediction candidate group for each mode.
[0507] Alternatively, when multiple motion information coding modes are supported in a common mode, in addition to the flag for selecting one of the skip mode, the merge mode, and the competition mode, an additional flag for identifying the detailed mode of the selected mode can also be supported. In the example, when the merge mode is selected, this means that a flag for selecting among detailed modes of the merge mode such as merge_inter, merge_tmp, merge_ibc, and merge_affine is additionally supported. Alternatively, a flag indicating whether it is merge_inter is supported, and when it is not merge_inter, a flag for selecting among merge_tmp, merge_ibc, and merge_affine can be additionally supported.
[0508] All or some of the motion information coding mode candidates can be supported according to the coding settings. In this case, the coding settings can be defined by one or more elements of status information such as the size, shape, aspect ratio, position, etc. of a basic block (e.g., the target block), image type, image category, color component, inter-frame prediction support settings (e.g., whether template matching is supported, whether block matching is supported in the current picture, non-translational motion model support elements, etc.).
[0509] In an example, a supported motion information encoding mode can be determined according to the size of a block. In this case, for the size of the block, a support range can be determined by a first threshold size (minimum value) or a second threshold size (maximum value), and each threshold size can be represented as W, H, W×H, or W*H using the width (W) and height (H) of the block. For the first threshold size, W and H can be integers such as 4, 8, 16, or greater, and W*H can be an integer such as 16, 32, 64, or greater. For the second threshold size, W and H can be integers such as 16, 32, 64, or greater, and W*H can be an integer such as 64, 128, 256, or greater. The range can be determined by one of the first threshold size or the second threshold size, or can be determined by using both the first threshold size and the second threshold size.
[0510] In this case, the threshold size can be fixed or can be adaptive according to an image (e.g., image type, etc.). In this case, the first threshold size can be set based on the size of a minimum coding block, a minimum prediction block, a minimum transform block, etc., and the second threshold size can be set based on the size of a maximum coding block, a maximum prediction block, a maximum transform block, etc.
[0511] In an example, a supported motion information encoding mode can be determined according to an image type. In this case, for image type I, at least one of a skip mode, a merge mode, and a competition mode can be included. In this case, a separate motion information encoding mode regarding a method of performing block matching (or template matching) and an affine model (hereinafter, referred to as a term "element") in a current picture can be supported, or a motion information encoding mode can be supported by combining two or more elements. Alternatively, elements according to a predetermined motion information encoding mode can be configured in a motion information prediction candidate group.
[0512] For image type P / B, at least one of a skip mode, a merge mode, or a competition mode can be included. In this case, a separate motion information encoding mode regarding general inter-frame prediction, template matching, block matching, and an affine model (hereinafter, referred to as a term "element") in a current picture can be supported, or a motion information encoding mode can be supported by combining two or more elements. Alternatively, elements according to a predetermined motion information encoding mode can be configured in a motion information prediction candidate group.
[0513] Next, the definition and configuration according to the motion information encoding mode are described.
[0514] The skip mode, merge mode, and competition mode can use / refer to the motion information of related blocks (candidate blocks) set based on the basic blocks encoded with the motion information for the target block. In other words, motion information (e.g., reference image or reference region) can be derived from the related blocks, and the motion information (predicted motion information) obtained based on the related blocks can be predicted as the motion information of the target block.
[0515] In this case, the difference information between the motion vector of the target block and the predicted motion vector may not be generated in the skip mode and merge mode, and the difference information may be generated in the competition mode. In other words, this may mean that the predicted motion vector is used as the motion vector of the target block as it is in the skip mode and merge mode, but a modification to generate the difference information in the merge mode is possible.
[0516] In this case, the information about the reference image of the target block can be used as the predicted reference image information as it is in the skip mode and merge mode, and the information about the reference image of the target block can be encoded as the reference image information without prediction or based on the predicted reference image information (e.g., encoding the difference value after prediction) in the competition mode.
[0517] In this case, the residual component of the target block may not be generated in the skip mode, and the residual component of the target block may be generated in the merge mode and competition mode. In other words, this means that the residual component can be generated in the merge mode and competition mode, and processing for the residual component (transformation, quantization, inverse processing of the residual component) can be performed (e.g., the processing may not be performed by checking whether there is a residual signal in a block such as coded_block_flag), but a modification not to generate the residual component in the merge mode is possible.
[0518] The motion information prediction candidate groups of the motion information encoding mode can be configured differently. In an example, the skip mode and merge mode can commonly configure the candidate group, and the competition mode can configure a separate candidate group. The candidate group configuration setting can be determined based on the motion information encoding mode.
[0519] In this case, the candidate group configuration setting can be defined by the number of candidate groups, the category and position of the candidate blocks (related blocks), the candidate configuration method, etc.
[0520] The number of candidate groups can be k, and k can be an integer from 1 to 6 or greater. In this case, when the number of candidate groups is 1, this means that no candidate group selection information is generated, and the motion information of a predetermined candidate block is set as the predicted motion information, and when the number of candidate groups is equal to or greater than 2, candidate group selection information can be generated.
[0521] The category of the candidate block may be one or more of inter_blk_A, inter_blk_B, and inter_blk_C. In this case, inter_blk_A may be a substantially included category, and the other categories may be additionally supported categories, but are not limited thereto.
[0522] For inter_blk_A or inter_blk_C, the position of the candidate block may be adjacent to the basic block in the left, upper, upper-left, upper-right, lower-left directions, and may be derived from the description regarding Figure 6 For inter_blk_B, it may be specified as a block having pattern information identical / similar to the coding information of the target block, and the relevant description may be derived from the part specifying the block having the above-mentioned horizontal relationship. In the example, when configuring a motion information prediction candidate group in a motion information coding mode based on a motion model, the motion information of blocks having the same pattern may be included in the candidate group.
[0523] As described, the position of the candidate block referred to for configuring the motion information prediction candidate group is described. Next, a method of obtaining predicted motion information based on the corresponding candidate block will be described.
[0524] The candidate block for configuring the motion information prediction candidate group for inter-frame prediction may represent the relevant block among the above-mentioned relationships between blocks. In other words, the block specified among many relevant blocks based on the basic block may be called a candidate block, and the motion information prediction candidate group may be configured based on the motion information of the candidate block, etc. In the examples mentioned later, it should be understood that the candidate block may represent the relevant block.
[0525] Figure 16 is an exemplary diagram of the arrangement of blocks spatially or temporally adjacent to the basic block according to an embodiment of the present disclosure. Specifically, Figure 16 may be an exemplary diagram of the arrangement of blocks (relevant blocks or candidate blocks) having a horizontal relationship with the basic block belonging to the categories of inter_blk_A and inter_blk_C. In the examples mentioned later, the block belonging to inter_blk_A is called a spatial candidate, and the block belonging to inter_blk_C is called a temporal candidate.
[0526] Referring to Figure 16 , the blocks adjacent to the basic block in the left, upper, upper-left, upper-right, lower-left directions, etc., and the blocks adjacent to the basic block corresponding to the basic block in a different spatial (Col_Pic) in terms of time in the center, left, right, upper, lower, upper-left, upper-right, lower-left, lower-right directions, etc., may be configured as candidate blocks for predicting the motion information of the target block.
[0527] In addition, blocks adjacent in the above direction can be divided (classified) in units of one or more sub-blocks (for example, the L block in the left direction can be divided into sub-blocks l0, l1, l2, l3). Figure 16 Figure 16 is an example where adjacent blocks in each direction are configured with 4 sub-blocks (the center of the corresponding block is configured with 16 sub-blocks), but it can be divided into various p sub-blocks, but not limited to this. In this case, p can be an integer such as 1, 2, 3, 4 or larger. In addition, p can be adaptively determined based on the position (direction) of adjacent blocks.
[0528] In this case, Col_Pic can be an image adjacent before or after the current image (for example, when the interval between images is 1), and the corresponding block can be set to have the same position in the image as the basic block.
[0529] Alternatively, Col_Pic can be an image that predefines the interval between images based on the current image (for example, the interval between images is z. z is an integer such as 1, 2, 3), and the corresponding block can be set to have a position that moves a predetermined disparity vector on the predetermined coordinates (for example, upper left) of the basic block, and the disparity vector can be set to a predefined value.
[0530] Alternatively, Col_Pic can be set based on the motion information of adjacent blocks (for example, the reference image), and the disparity vector can be set based on the motion information of adjacent blocks (for example, the motion vector) to determine the position of the corresponding block.
[0531] In this case, k adjacent blocks can be referred to, and k can be an integer such as 1, 2 or larger. When k is equal to or greater than 2, Col_Pic and the disparity vector can be obtained based on operations such as the maximum value, minimum value, median value, weighted average value, etc. of the motion information (for example, the reference image or the motion vector) of adjacent blocks. For example, the disparity vector can be set to the motion vector of the left block or the upper block, and can be set to the median value or average value of the motion vectors of the left block and the lower left block.
[0532] The above settings of the time candidate can be determined based on the motion information configuration settings, etc. For example, the position of Col_Pic, the position of the corresponding block, etc. can be determined according to whether the motion information to be included in the motion information prediction candidate group is block-based motion information or sub-block-based motion information configured in units of blocks or sub-blocks. In the example, when sub-block-based motion information is obtained, the block at the position where a predetermined disparity vector is moved can be set to have the position of the corresponding block.
[0533] The above examples represent the following situation: Information regarding the position of the block corresponding to Col_Pic is implicitly determined, and relevant information can be explicitly generated in units such as sequences, pictures, slices, tile groups, tiles, and brick-shaped blocks.
[0534] According to the motion information coding mode, the above-mentioned motion information of spatially or temporally neighboring blocks (or candidate blocks) (spatial candidates and temporal candidates respectively) can be included in the motion information prediction candidate group.
[0535] In example (1), singular motion information can be included in the candidate group as it is. In other words, without change, all or some of the singular motion vector, singular reference picture, singular prediction direction, etc. can be used as the prediction values of all or some of the motion vector, reference picture, prediction direction, etc. of the target block.
[0536] In this case, singular motion information can be obtained from one candidate block of the spatial candidate or the temporal candidate. In the examples mentioned later, it is assumed that the candidate obtained in this example is a spatial candidate. Additionally, it can be the case applied to all modes among the skip mode, merge mode, and competition mode.
[0537] In example (2), after the adjustment (or transformation) process, singular motion information can be included in the candidate group. This example also assumes the case where singular motion information is obtained from one candidate block. Specifically, the adjustment process regarding the motion information (e.g., motion vector) of the candidate block can be performed based on the distance interval between the current picture and the reference picture and the distance interval between the picture to which the candidate block belongs and the reference picture of the candidate block.
[0538] In other words, the motion vector of the candidate block can be adjusted in the scaling process based on the distance interval between the current picture and the reference picture. And the adjusted motion vector of the candidate block can be included in the candidate group. Regarding the motion information, prediction direction, etc. of the reference picture, the reference picture or prediction direction information of the candidate block can be included in the candidate group without change, or can be included in the candidate group based on the distance interval of the reference picture. Alternatively, information set based on predetermined information (e.g., the reference picture is a picture before or after the current picture <reference picture index is 0>, the prediction direction is one of forward, backward, bi-directional, etc.) or the prediction direction information of the reference picture or the target block can be included in the candidate group.
[0539] In this case, the singular motion information can be a time candidate. In the examples mentioned later, it is assumed that the candidate obtained in the example is a time candidate. In addition, it can be a case applied to all or some of the skip mode, merge mode, and competition mode. For example, it can be applied to the skip mode, merge mode, and competition mode, or it can not be applied to the skip mode and merge mode and can be applied to the competition mode.
[0540] In example (3), after the combination process, multiple motion information can be included in the candidate group. The combination process can be performed on all or some of the motion vector, reference image, and prediction direction. The combination process represents a process such as the median value, weighted average, etc. of the motion information. Specifically, q motion information obtained based on the median value, weighted average, etc. of p motion information can be included in the candidate group. In this case, p can be an integer such as 2, 3, 4, or greater, and q can be an integer such as 1, 2, or greater. In the example, p and q can be 2 and 1 respectively. In this case, for the average value, the same weight (e.g., 1:1, 1:1:1, etc.) can be applied to the motion information or different weights (e.g., 1:2, 1:3, 2:3, 1:1:2, 1:2:3, etc.) can be applied.
[0541] Multiple motion information can be obtained from multiple candidate blocks. In this case, it is assumed that one motion information is obtained from one candidate block, but it is not limited to this. The motion information of the candidate block (input value of the combination process) can be information restricted to one of (1) or (2) or belonging to both (1) and (2).
[0542] In addition, multiple motion information can be obtained from either the spatial candidate or the time candidate, and can be obtained from both the spatial candidate and the time candidate. For example, the combination process can be performed by using the spatial candidate, or the combination process can be performed by using the time candidate. Alternatively, the combination process can be performed by using one or more spatial candidates and one or more time candidates.
[0543] Multiple motion information can be obtained from the candidates pre-included in the candidate group. In other words, multiple candidate motion information (e.g., spatial candidate or time candidate) previously included in the candidate group configuration process can be used as the input value of the combination process. In this case, the candidate motion information used as the input value of the combination process can be obtained from either the spatial candidate or the time candidate, and can be obtained from both the spatial candidate and the time candidate.
[0544] This description can be a case that applies to all or some of the skip mode, merge mode, and competition mode. For example, it can apply to the skip mode, merge mode, and competition mode, or it can apply to the skip mode and merge mode and may not apply to the competition mode.
[0545] For the motion information obtained in the combined processing, the motion information derived from the spatial candidate or the temporal candidate is called the spatially-derived candidate or the temporally-derived candidate, and the motion information derived from the spatial candidate and the temporal candidate is called the spatially and temporally-derived candidate. In the following examples, the processing of deriving candidates will be described.
[0546] As an example of the spatially-derived candidate, the spatially-derived candidate can be obtained by applying a process of obtaining a weighted average or a median value based on all or some of the motion vectors in the blocks in the left, top, top-left, top-right, bottom-left directions of the basic block. Alternatively, the spatially-derived candidate can be obtained by applying a process of obtaining a weighted average or a median value based on all or some of the spatial candidates already included in the candidate group. Next, assume that the motion vectors l3, t3, tl, tr, bl are regarded as the input values of the process of obtaining Figure 16 the weighted average or the median value in the Curr_Pic.
[0547] For example, the x component of the motion vector can be derived from avg(l3_x, t3_x, tl_x, tr_x, bl_x), the y component of the motion vector can be derived from avg(l3_y, t3_y, tl_y, tr_y, bl_y), and assume that avg is a function for calculating the average value of the motion vectors in the parentheses. Alternatively, the x component of the motion vector can be derived from median(l3_x, t3_x, tl_x, tr_x, bl_x), and the y component of the motion vector can be derived from median(l3_y, t3_y, tl_y, tr_y, bl_y), and assume that median is a function for calculating the median value of the motion vectors in the parentheses.
[0548] As an example of the temporally-derived candidate, the temporally-derived candidate can be obtained by applying a process of obtaining a weighted average or a median value based on all or some of the motion vectors in the blocks in the left, right, top, bottom, top-left, top-right, bottom-left, bottom-right, center directions of the block corresponding to the basic block. Alternatively, the temporally-derived candidate can be obtained by applying a process of obtaining a weighted average or a median value based on all or some of the temporal candidates already included in the candidate group. Next, assume that the motion vectors br, r2, b2 are regarded as the input values of the process of obtaining Figure 16 the weighted average or the median value in the Col_Pic.
[0549] For example, the x component of the motion vector can be derived from avg(br_x, r2_x, b2_x), and the y component of the motion vector can be derived from avg(br_y, r2_y, b2_y). Alternatively, the x component of the motion vector can be derived from median(br_x, r2_x, b2_x), and the y component of the motion vector can be derived from median(br_y, r2_y, b2_y).
[0550] As an example of obtaining candidates in terms of space and time, the space and time obtained candidates can be obtained by applying a process of obtaining a weighted average or a median value by using motion vectors of all or some of the spatial neighboring blocks of a basic block and the blocks corresponding to the basic block. Alternatively, the space and time obtained candidates can be obtained by applying a process of obtaining a weighted average or a median value by using all or some of the spatial candidates or the temporal candidates that have already been included in the candidate group. Examples thereof are omitted.
[0551] The motion information obtained in the above process can be added to (included in) the candidate group for predicting the motion information of the target block.
[0552] The embodiments in (1) to (3) describe the following examples: the spatial candidates or the temporal candidates are included in the candidate group for predicting the motion information without change, or the space obtained candidates, the time obtained candidates, and the space and time obtained candidates obtained based on a plurality of spatial candidates or temporal candidates are included in the candidate group for predicting the motion information. In other words, the embodiments in (1) and (2) can be descriptions of spatial or temporal candidates, while the embodiment in (3) can be a description of the space obtained candidates, the time obtained candidates, and the space and time obtained candidates.
[0553] The above-mentioned examples can be the case where the motion vector in units of pixels is the same in the target block. As an example of the case where the motion vector in units of pixels or in units of sub-blocks is different in the target block, the affine model has been described above. In the examples mentioned later, the motion vector in units of sub-blocks has partial variations in the target block, but examples of the case where it is represented as a candidate based on spatial or temporal correlation will be described.
[0554] Next, the case of obtaining a predicted value of the motion information in units of sub-blocks based on one or more pieces of motion information is described. It can be the following example: a process similar to the process (3) of obtaining motion information based on a plurality of pieces of motion information is performed, but the motion information in units of sub-blocks is obtained (4). The motion information obtained in the process (4) can be included in the candidate group for predicting the motion information. The embodiment (4) will be described later by using the space obtained candidates, the time obtained candidates, and the space and time obtained candidates which are terms used in this example.
[0555] When obtaining a predicted value of motion information in units of sub - blocks, the basic block can be restricted to the target block. In addition, when configuring a candidate group by mixing motion information in units of sub - blocks and motion information in units of blocks, it is assumed that the basic block is restricted to the target block.
[0556] As an example of obtaining candidates spatially, it is possible to obtain a motion information of a spatially neighboring block (in this example, various blocks in the Curr_Pic of Figure 16 which is configured with 4 sub - blocks <in the Curr_Pic of Figure 16 where c0 + c1 + c4 + c5 / c2 + c3 + c6 + c7 / c8 + c9 + c12 + c13 / c10 + c11 + c14 + c1 are respectively called D0, D1, D2, D3>) obtained based on sub - blocks in the target block, or a spatially derived candidate can be obtained by a weighted average or median value based on multiple motion informations. The obtained spatially derived candidate can be added to the motion information prediction candidate group of the target block.
[0557] For example, for sub - block D0, l1, t1, t2 (the lower region of the left block, the right region of the upper block, the upper - left block based on the sub - block) can be spatially neighboring blocks of the corresponding sub - block; for sub - block D1, l1, t3 can be spatially neighboring blocks of the corresponding sub - block; for sub - block D2, l3, t1, t2 can be spatially neighboring blocks of the corresponding sub - block; for sub - block D3, l3, t3, tr can be spatially neighboring blocks of the corresponding sub - block. In this case, what is added to the motion information prediction candidate group of the target block is the spatially derived candidate (assuming one is supported in this example), but the motion information thus obtained can be obtained in units of each sub - block (in this example, up to 4 sub - blocks). In other words, when the decoder parses the information for selecting the spatially derived candidate, the predicted value of the motion information of each sub - block can be obtained as in the above process.
[0558] As an example of obtaining candidates temporally, it is possible to obtain a motion information of a temporally neighboring block (in this example, various blocks in the Col_Pic of Figure 16 where the blocks used in this example are located at the center, lower - right, right, or lower positions neighboring the corresponding sub - blocks) obtained based on sub - blocks in the target block, or a temporally derived candidate can be obtained by a weighted average or median value based on multiple motion informations. The obtained temporally derived candidate can be added to the motion information prediction candidate group of the target block.
[0559] For example, for sub-block D0, c10, c6, c9 (the lower left block, the lower region of the right block, the right region of the lower block based on the corresponding sub-block) can be the temporally adjacent blocks of the corresponding sub-block; for sub-block D1, r2, r1, c11 can be the temporally adjacent blocks of the corresponding sub-block; for sub-block D2, b2, c14, b1 can be the temporally adjacent blocks of the corresponding sub-block; for sub-block D3, br, r3, b3 can be the temporally adjacent blocks of the corresponding sub-block. As the candidates added in this example are temporally derived, the motion information derived therefrom can be obtained for each sub-block.
[0560] As an example of spatially and temporally derived candidates, a weighted average value, a median value, etc. of motion information of spatially or temporally adjacent blocks obtained based on sub-blocks in a target block (the blocks used in this example are at spatially adjacent left and upper positions and at a temporally adjacent lower right position based on the sub-blocks) can be obtained as spatially and temporally derived candidates, and the weighted average value, the median value, etc. can be added to the group of motion information prediction candidates of the target block.
[0561] For example, for sub-block D0, l1, t1, c10 can be the adjacent blocks of the corresponding sub-block; for sub-block D1, l1, t3, r2 can be the adjacent blocks of the corresponding sub-block; for sub-block D2, l3, t1, b2 can be the adjacent blocks of the corresponding sub-block; for sub-block D3, l3, t3, br can be the adjacent blocks of the corresponding sub-block.
[0562] In summary, the motion information obtained based on one or more pieces of motion information in Embodiment (4) can be included in the candidate group. Specifically, q pieces of motion information obtained based on the median value, the weighted average value, etc. of p pieces of motion information can be included in the candidate group. In this case, p can be an integer such as 1, 2, 3, or larger, and q can be an integer such as 1, 2, or larger. In this case, for the average value, the same weight (e.g., 1:1, 1:1:1, etc.) can be applied to the motion information, or different weights (e.g., 1:2, 1:3, 2:3, 1:1:2, 1:2:3, etc.) can be applied to the motion information. The input values in the above processing can be limited to one of (1) or (2) or belong to the information of both (1) and (2).
[0563] The size of the sub-block can be m×n, where m and n can be integers such as 4, 8, 16, or larger, and m and n can be the same or different. The size of the sub-block can have a fixed value, or can be adaptively set based on the size of the target block. In addition, the size of the sub-block can be implicitly determined according to the coding settings, or relevant information can be explicitly generated in various units. In this case, for the limitations on the coding settings, various examples mentioned above can be referred to.
[0564] This description can be a case that is applied to all or some of the skip mode, merge mode, and competition mode. For example, it can be applied to the skip mode, merge mode, and competition mode, or it can be applied to the skip mode and merge mode and may not be applied to the competition mode.
[0565] Regarding whether to support the embodiments in (3) and (4), relevant information can be explicitly generated in units such as sequences, pictures, slices, tile groups, tiles, bricks, etc. Alternatively, it can be determined whether to support based on the coding settings. In this case, the coding settings can be defined by one or more elements of status information such as the size, shape, aspect ratio, position, etc. of a basic block (e.g., the target block), image type, image category, color component, settings related to inter-frame prediction (e.g., motion information coding mode, whether template matching is supported, whether block matching is supported in the current picture, non-translational motion model support elements, etc.).
[0566] For example, it can be determined whether to support a candidate group of the embodiments in (3) or (4) according to the size of the block. In this case, for the size of the block, the support range can be determined by a first threshold size (minimum value) or a second threshold size (maximum value), and each threshold size can be expressed as W, H, W×H, W*H using the width (W) and height (H) of the block. For the first threshold size, W and H can be integers such as 4, 8, 16 or larger, and W*H can be integers such as 16, 32, 64 or larger. For the second threshold size, W and H can be integers such as 16, 32, 64 or larger, and W*H can be integers such as 64, 128, 256 or larger. This range can be determined by one of the first threshold size or the second threshold size, or can be determined by using both the first threshold size and the second threshold size.
[0567] In this case, the threshold size can be fixed or can be adaptive according to the image (e.g., image type, etc.). In this case, the first threshold size can be set based on the size of the smallest coding block, the smallest prediction block, or the smallest transform block, etc., and the second threshold size can be set based on the size of the largest coding block, the largest prediction block, or the largest transform block, etc.
[0568] For the affine model, according to the motion information coding mode, the motion information of spatially or temporally adjacent blocks (candidate blocks) can be included in the motion information prediction candidate group. In this case, the positions of the spatially or temporally adjacent blocks can be the same as or similar to those in the previous embodiments.
[0569] In this case, the candidate group configuration can be determined according to the motion model of the candidate blocks.
[0570] For example, when the motion model of a candidate block is an affine model, the motion vector set configuration of the candidate block can be included in the candidate group without change.
[0571] Alternatively, when the motion model of a candidate block is a translational motion model, it can be included as a candidate for the motion vector at the position of the control point based on the relative position of the target block. In an example, when the upper left, upper right, and lower left coordinates are used as control points, the motion vector of the upper left control point can be predicted based on the motion vector of the left block, upper block, or upper left block of the basic target block (e.g., for a translational motion model), the motion vector of the upper right control point can be predicted based on the motion vector of the upper block or upper right block of the basic block (e.g., for a translational motion model), and the motion vector of the lower left control point can be predicted based on the motion vector of the left block or lower left block of the basic block.
[0572] In summary, when the motion model of a candidate block is an affine model, the motion vector set of the corresponding block can be included in the candidate group (A); and when the motion model of a candidate block is a translational motion model, it can be considered as a candidate for the motion vector of a predetermined control point of the target block, and the motion vector set obtained according to the combination of each control point can be included in the candidate group (B).
[0573] In this case, the candidate group can be configured by using only one of methods A or B, or the candidate group can be configured by using both methods A and B. And method A can be configured first, and method B can be configured subsequently, but it is not limited thereto.
[0574] Next, a method for configuring a motion information prediction candidate group according to a motion information coding mode will be described.
[0575] For ease of description, the merge mode and the competition mode are described under the assumption that the skip mode has the same candidate group configuration as the candidate group configuration of the merge mode, but some elements of the candidate group configuration of the skip mode can be different from the merge mode.
[0576] (Merge_inter mode)
[0577] The motion information prediction candidate group for the merge mode (hereinafter, the merge mode candidate group) can include k candidates, and k can be an integer such as 2, 3, 4, 5, 6, or greater. The merge mode candidate group can include at least one of spatial candidates or temporal candidates.
[0578] Spatial candidates can be derived based on the basic block from at least one of the blocks adjacent in the left, upper, upper left, upper right, lower left directions, etc. There can be a priority for configuring the candidate group, and priorities such as left - upper - lower left - upper right - upper left, left - upper - upper right - lower left - upper left, upper - left - lower left - upper left - upper right, etc. can be set. In an example, in Figure 16 In Curr_Pic of [[ID=]], the priorities can be set in the order of l3-t3-bl-tr-tl.
[0579] All or some of the candidates can be included in the candidate group based on the priorities, the availability of each candidate block (e.g., judged based on the coding mode, the position of the block, etc.), and the maximum allowable number of spatial candidates (an integer between p. 1 and the number of merge mode candidate groups). Depending on the maximum allowable number and the availability, they may not be included in the candidate group in the order of tl-tr-bl-t3-l3. When the maximum allowable number is 4 and the availability of the candidate blocks is completely true, the motion information of t1 may not be included in the candidate group, and when the availability of some candidate blocks is false, the motion information of t1 may be included in the candidate group.
[0580] Temporal candidates can be derived from at least one of the blocks adjacent in the center, left, right, top, bottom, upper left, upper right, lower left, lower right directions, etc. based on the block corresponding to the base block. There may be priorities for configuring the candidate group, and priorities such as center-lower left-right-bottom, lower left-center-left, etc. can be set. In the example, in Figure 16 Col_Pic of [[ID=]], the priorities can be set in the order of cl0-br.
[0581] Based on the priorities, the availability of each candidate block, and the maximum allowable number of temporal candidates (an integer between q. 1 and the number of merge mode candidate groups), all or some of the candidates can be included in the candidate group. When the maximum allowable number is 1 and the availability of c10 is true, the motion information of c10 can be included in the candidate group, and when the availability of c10 is false, the motion information of br can be included in the candidate group.
[0582] In this case, the motion vector of the temporal candidate can be obtained based on the motion vector of the candidate block, and the reference image of the temporal candidate can be obtained based on the reference image of the candidate block, or the reference image of the temporal candidate can be obtained as a predefined reference image (e.g., the reference picture index is 0).
[0583] For the priorities included in the merge mode candidate group, spatial candidates - temporal candidates can be set, and vice versa; and it is possible to support mixing and configuring the priorities of spatial candidates and temporal candidates. This example assumes the case of spatial candidates - temporal candidates.
[0584] In addition, motion information of blocks that belong to the same spatial location and are far neighbors (statistical candidates) or candidates derived on a per-block basis (spatially derived candidates, temporally derived candidates, spatially and temporally derived candidates) may be additionally included in the merge mode candidate group. The statistical candidates and candidates derived on a per-block basis may be configured after the spatial candidates and temporal candidates, but various priorities are possible and are not limited thereto. In the present disclosure, it is assumed that the candidate group is configured in the order of statistical candidates - derived candidates on a per-block basis, but the reverse order is possible and is not limited thereto. In this case, the candidates derived on a per-block basis may use only any one of the spatially derived candidates, temporally derived candidates, or spatially and temporally derived candidates, or the candidates derived on a per-block basis may use all of them. This example assumes the case of using spatially derived candidates.
[0585] In this example, the statistical candidates are motion information of blocks that are far from the base block (or are not the closest neighboring blocks), and it may be restricted to blocks having the same coding mode as the target block. Further, (when the statistical candidates are not motion information on a sub-block basis) it may be restricted to blocks having motion information on a per-block basis.
[0586] For the statistical candidates, up to n pieces of motion information may be managed by a method such as FIFO, and z pieces of the motion information may be included in the merge mode candidate group as statistical candidates. Depending on the candidate configuration already included in the merge mode candidate group, z may be variable, and z may be an integer such as 0, 1, 2, or greater and may be less than or equal to n.
[0587] Candidates derived on a per-block basis may be obtained by combining n candidates already included in the merge mode candidate group, and n may be an integer such as 2, 3, 4, or greater. The number (n) of the combined candidates may be information explicitly generated in units such as sequences, pictures, slices, tile groups, tiles, brick-shaped blocks, blocks, etc. Alternatively, n may be implicitly determined according to coding settings. In this case, the coding settings may be defined based on one or more factors such as the size, shape, or position of the base block, image type, color component, etc.
[0588] In addition, the number of the combined candidates may be determined based on the number of candidates not filled in the merge mode candidate group. In this case, the number of candidates not filled in the merge mode candidate group may be the difference between the number of the merge mode candidate group and the number of candidates already filled. In other words, when the configuration of the merge mode candidate group has been completed, candidates derived on a per-block basis may not be added. When the configuration of the merge mode candidate group has not been completed, candidates derived on a per-block basis may be added, but when the number of candidates filled in the merge mode candidate group is equal to or less than 1, candidates derived on a per-block basis are not added.
[0589] When the number of unfilled candidates is 1, candidates obtained in units of blocks can be added based on 2 candidates, and when the number of unfilled candidates is equal to or greater than 2, candidates obtained in units of blocks can be added based on 3 or more candidates.
[0590] The positions of n combined candidates can be preset positions in the combined mode candidate group. For example, an index can be assigned to each candidate belonging to the combined mode candidate group (e.g., 0 to <k-1>). In this case, k represents the number of merge mode candidate groups. In this case, the positions of the n combined candidates can correspond to indices 0 to index (n - 1) in the merge mode candidate group. In the example, the candidates obtained in units of blocks can be obtained according to the index combination of (0, 1)-(0, 2)-(1, 2)-(0, 3).
[0591] Alternatively, the n combined candidates can be determined by considering the prediction directions of each candidate belonging to the merge mode candidate group. In the example, among the candidates belonging to the merge mode candidate group, only bi-directional prediction candidates can be selectively used, or only uni-directional prediction candidates can be selectively used.
[0592] Through the above processing, the configuration of the merge mode candidate group may or may not be completed. When the configuration of the merge mode candidate group is not completed, the configuration of the merge mode candidate group can be completed by using zero motion vectors.
[0593] In the candidate group configuration process, the candidate group may not include posterior motion information redundant with the candidates already included in the candidate group, but this is not limited thereto, and an exception case including redundant motion information may also occur. In addition, redundancy means that the motion vector, reference picture, and prediction direction are the same, but an exception configuration allowing a predetermined error range (e.g., motion vector) is also possible.
[0594] In addition, the merge mode candidate group can be configured based on the correlation between blocks. For example, for spatial candidates, low priority can be assigned to the motion information of candidate blocks determined to have low correlation with the basic block.
[0595] For example, assume that the priority of the candidate group for configuring spatial candidates is left - up - lower left - upper right - lower left. When it is confirmed according to the correlation between blocks that the left block has low correlation with the basic block, the priority can be changed to set the left block as the posterior order. In the example, the priority can be changed to the order of up - lower left - upper right - lower left - left.
[0596] Alternatively, assume that the priority of the candidate group for configuring the candidates obtained in units of blocks is in the order of (0, 1)-(0, 2)-(1, 2)-(0, 3). When the candidate corresponding to the first index in the candidate group is the motion information of the left block (spatial candidate), the priority can be changed to set the candidate corresponding to the first index as the posterior order. In the example, the priority can be changed to the order of (0, 2)-(0, 3)-(0, 1)-(1, 2).
[0597] When the merge mode candidate group is configured based on the correlation between blocks, the candidate group can be effectively configured and the coding performance can be improved.
[0598] (comp_inter mode)
[0599] The motion information prediction candidate group for the competition mode (hereinafter, the competition mode candidate group) may include k candidates, where k may be an integer such as 2, 3, 4, or greater. The merge mode candidate group may include at least one of a spatial candidate or a temporal candidate.
[0600] Spatial candidates may be derived from at least one of the blocks adjacent in the left, up, upper left, upper right, lower left directions, etc. based on the basic block. Alternatively, at least one candidate may be derived from the blocks adjacent in the left direction (left, lower left blocks) and the blocks adjacent in the up direction (upper left, up, upper right blocks), which will be described later by assuming such a setting.
[0601] There may be two or more priorities for configuring the candidate group. The lower left - left priority may be set in the region adjacent in the left direction, and the upper right - up - upper left priority may be set in the region adjacent in the up direction.
[0602] The above example may be a configuration where spatial candidates are derived only from the blocks having the same reference picture as the target block, and spatial candidates may be derived by scaling processing (next, marked with *) based on the reference picture of the target block. In this case, the left – lower left – left* – lower left* or left – lower left – lower left* – left* priority may be set in the region adjacent in the left direction, and the upper right – up – upper left – upper right* – up* – upper left* or upper right – up – upper left – upper left* – up* – upper right* priority may be set in the region adjacent in the up direction.
[0603] Temporal candidates may be derived from at least one of the blocks adjacent in the center, left, right, up, down, upper left, upper right, lower left, lower right directions, etc. based on the block corresponding to the basic block. There may be priorities for configuring the candidate group, and priorities such as center - lower left - right - down, lower left - center - left, etc. may be set. In the example, in Figure 16 the Col_Pic of, the priorities may be set in the order of c10 - br.
[0604] When the sum of the maximum allowable number of spatial candidates and the maximum allowable number of temporal candidates is less than the number of the competition mode candidate group, the temporal candidates may be included in the candidate group regardless of the candidate group configuration of the spatial candidates.
[0605] Based on the priorities, the availability of each candidate block, and the maximum allowable number of temporal candidates (an integer between q. 1 and the number of the competition mode candidate group), all or some of the candidates may be included in the candidate group.
[0606] In this case, when the maximum allowable number of spatial candidates is set to be the same as the number of merge mode candidate groups, the temporal candidates may not be included in the candidate group, and when the maximum allowable number is not filled based on the spatial candidates, the temporal candidates may be included in the candidate group. This example assumes the latter case.
[0607] In this case, the motion vector of the temporal candidate can be obtained based on the motion vector of the candidate block and the distance interval between the current image and the reference image of the target block, and the reference image of the temporal candidate can be obtained based on the distance interval between the current image and the reference image of the target block, or the reference image of the temporal candidate can be obtained based on the reference image of the temporal candidate, or the reference image of the temporal candidate can be obtained as a predefined reference image (e.g., the reference picture index is 0).
[0608] Through the above processing, the configuration of the competition mode candidate group may or may not be completed. When the configuration of the competition mode candidate group is not completed, the competition mode candidate group can be configured by using a zero motion vector.
[0609] In the candidate group configuration process, the posterior motion information redundant with the candidates already included in the candidate group may not be included in the candidate group, but this is not limited to this, and an exception case including redundant motion information may also be generated. In addition, redundancy means that the motion vector, reference picture, and prediction direction are the same, but an exception configuration allowing a predetermined error range (e.g., the motion vector) is also possible.
[0610] In addition, the competition mode candidate group can be configured based on the correlation between blocks. For example, for spatial candidates, the motion information of candidate blocks determined to have low correlation with the basic block can be excluded from the candidate group.
[0611] For example, assume that in the blocks adjacent in the left direction, the priority of the candidate group for configuring spatial candidates is left - lower left - left* - lower left*; in the blocks adjacent in the upper direction, the priority of the candidate group for configuring spatial candidates is upper right - upper - upper left - upper right* - upper* - upper left*. When it is confirmed according to the correlation between blocks that the upper block has low correlation with the basic block, the candidates regarding the upper block can be removed from the priority. In the example, the priority can be changed to the order of upper right - upper left - upper right* - upper left*.
[0612] When the competition mode candidate group is configured based on the correlation between blocks, the candidate group can be effectively configured and the coding performance can be improved.
[0613] Next, the process of performing inter - frame prediction according to the motion information coding mode will be described.
[0614] Derive the motion information coding mode (1) of the target block. Specify a reference block (2) according to the derived motion information coding mode. Configure a motion information prediction candidate group (3) by using the motion information obtained based on the specified reference block and the motion information coding mode. Derive the motion information of the target block from the motion information prediction candidate group (4). Inter-frame prediction can be performed by using the motion information of the target block.
[0615] In (1), flag information regarding the motion information coding mode of the target block can be signaled. In this case, the motion information coding mode can be determined as one of a skip mode, a merge mode, and a competition mode; and according to the determined mode, the signaled flag can be configured with one or two or more pieces of information. In addition, additional flag information for selecting one of the detailed categories of the determined mode can be signaled.
[0616] In (2), a reference block is specified according to the derived motion information coding mode. The position of the reference block according to the motion information coding mode can be configured differently. In (3), it can be determined which motion information will be derived from the specified reference block based on the motion information coding mode, and the derived motion information can be included in the candidate group.
[0617] In (4), the motion information corresponding to the respective index in the candidate group can be derived based on the candidate selection information of the target block. In this case, the configuration of the derived motion information can be determined according to the motion information coding mode. In an example, for the skip mode or the merge mode, information regarding a motion vector, a reference image, and a prediction direction can be derived based on one piece of candidate selection information. Alternatively, for the competition mode, information regarding the motion vector can be derived based on one piece of candidate selection information, and information regarding the reference image and the prediction direction can be derived based on other predetermined flags. In (5), inter-frame prediction can be performed by using the motion information of the target block.
[0618] A prediction block can be generated by using the motion information obtained through the above processing. The target block can be reconstructed by adding the residual component of the target block. In this case, the residual component can be derived by performing at least one of dequantization or inverse transformation on the residual coefficients signaled from the bitstream.
[0619] Next, the differential motion vector, that is, the difference between the motion vector of the target block and the predicted motion vector, will be described later, and it can be a description corresponding to the competition mode.
[0620] The differential motion vector can be represented according to the motion vector precision (in this example, it is assumed that the motion vector precision is determined according to the interpolation precision of the reference picture. For the above-mentioned content <for example, the example where the motion vector rather than the differential motion vector is encoded unchanged>, when the motion vector precision is adaptively determined, it does not correspond to this example).
[0621] For example, when the motion vector precision is 1 / 4 unit, the motion vector of the target block is (2.5, -4), and the predicted motion vector is (3.5, -1), the differential motion vector can be (-1, -3).
[0622] When the differential motion vector precision is the same as the motion vector precision (i.e., 1 / 4 unit), the x differential component may need to be based on the codeword of the 4th index and the negative sign, while the y differential component may need to be based on the codeword of the 12th index and the positive sign.
[0623] In addition to the motion vector precision, the precision of the differential motion vector can also be supported. In this case, the differential motion vector precision can have a precision less than or equal to the motion vector precision as a candidate group.
[0624] The differential motion vector can be represented differently according to the differential motion vector precision.
[0625] For example, when the motion vector precision is 1 / 4 unit, the motion vector of the target block is (7.75, 10.25), and the predicted motion vector is (2.75, 3.25), the differential motion vector can be (5, 7).
[0626] When a fixed differential motion vector precision is supported (i.e., when the differential motion vector precision is the same as the motion vector precision. In this example, the motion vector precision is 1 / 4 unit), the x differential component may need to be based on the codeword of the 20th index and the positive sign, while the y differential component may need to be based on the codeword of the 28th index and the positive sign. This example can be an example according to the fixed differential motion vector precision.
[0627] When an adaptive differential motion vector precision is supported (i.e., when the differential motion vector precision is the same as or different from the motion vector precision. In other words, when one of multiple precisions is selected. In this example, it is assumed that the precision is an integer (1) unit), the x differential component may need to be based on the codeword of the 5th index and the positive sign, while the y differential component may need to be based on the codeword of the 7th index and the positive sign.
[0628] When it is assumed that the same codewords are assigned to the index order regardless of the accuracy in the example (e.g., the binarization method for differential motion vectors is the same); when a fixed differential motion vector is supported, long codewords such as the codewords according to the 20th index and the 28th index (in this example, it is assumed that short bits are assigned in the case of a small index and long bits are assigned in the case of a large index) are assigned to the x-differential component and the y-differential component; however, when an adaptive differential motion vector is supported, short codewords such as the codewords according to the 5th index and the 7th index can be assigned to the x-differential component and the y-differential component.
[0629] The above example can be a case where differential motion vector accuracy is supported for one differential motion vector. Alternatively, for the case where one differential motion vector accuracy is supported for the x-differential component and the y-differential component configuring one differential motion vector, it can be as follows.
[0630] For example, when the differential motion vector is (5, -1.75) and adaptive differential motion vector accuracy (in this example, integer (1) unit, 1 / 4 unit) is supported, the following configuration is possible.
[0631] For the integer (1) unit, the x-differential component may require the codeword according to the 5th index and a positive sign; while for the 1 / 4 unit, the y-differential component may require the codeword according to the 7th index and a negative sign. This example can be the best case, where the differential motion vector accuracy selection information for the x-differential component can be determined in integer units, and the differential motion vector accuracy selection information for the y-differential component can be determined in 1 / 4 units.
[0632] As in the above example, accuracy commonly applied to the differential motion vector can be supported, or accuracy applied separately to the components of the differential motion vector can be supported, and the accuracy can be determined according to the coding settings.
[0633] Additionally, as a description of the bitstream configuration, for the former, when at least one differential component is not 0, adaptive differential motion vector accuracy can be supported.
[0634] abs_mvd_x
[0635] if(abs_mvd_x)
[0636] mvd_sign_x
[0637] abs_mvd_y
[0638] if(abs_mvd_y)
[0639] mvd_sign_y
[0640] if((abs_mvd_x||abs_mvd_y)&&adaptive_mvd_precision_enabled_flag)
[0641] mvd_precision_flag
[0642] For the latter, when each differential component is not 0, adaptive differential motion vector precision can be supported.
[0643] abs_mvd_x
[0644] if(abs_mvd_x)
[0645] mvd_sign_x
[0646] abs_mvd_y
[0647] if(abs_mvd_y)
[0648] mvd_sign_y
[0649] if(abs_mvd_x&&adaptive_mvd_precision_enabled_flag)
[0650] mvd_precision_flag_x
[0651] if(abs_mvd_y&&adaptive_mvd_precision_enabled_flag)
[0652] mvd_precision_flag_y
[0653] In summary, the differential motion vector precision can be determined according to the coding settings, and the differential motion vector precision can be one of the integer unit and the decimal unit. The differential motion vector precision can have a preset differential motion vector precision or one of multiple differential motion vector precisions. For the former, the differential motion vector precision can be an example determined based on the reference picture interpolation precision.
[0654] When multiple differential motion vector precisions are supported (e.g., adaptive_mvd_precision_enabled_flag. If it is 0, the preset differential motion vector precision is used; if it is 1, one of the multiple differential motion vector precisions is used), differential motion vector precision selection information (e.g., mvd_precision_flag) can be generated. In this case, the precision candidate group can be configured with precisions in units of 4, 2, 1, 1 / 2, 1 / 4, 1 / 8, or 1 / 16. In other words, the number of candidate groups can be an integer such as 2, 3, 4, 5, or greater.
[0655] Although the motion vector precision and the differential motion vector precision in the examples in the above-mentioned embodiments are classified and described according to the motion information coding settings, they can have the same or similar configurations because: for the motion information to be encoded, they have fixed or adaptive precisions.
[0656] In the present disclosure, the description is made under the following assumption: the motion vector to be encoded is configured with a differential motion vector and sign information about the differential motion vector.
[0657] In the coding mode, the skip mode and the merge mode obtain prediction information and immediately use the prediction information as the motion information of the corresponding block without generating difference information, so adaptive differential motion vector precision is not supported. When at least one differential motion vector is not 0 in the competing mode, adaptive differential motion vector precision can be supported.
[0658] The reason for storing the motion information is as follows: because the motion information of neighboring blocks is used to encode the motion information of the target block, and after the target block is encoded, the target block can also be referred to when encoding the motion information of subsequent blocks. The description is made under the following setting: the precision of the motion vector storage unit is determined according to the reference picture interpolation precision. In other words, the description is made under the assumption that the representation unit of the motion vector is set to be different from the storage unit, and the examples mentioned later assume the case where adaptive differential motion vector precision setting is supported.
[0659] As described above, the precision of representing the motion vector (assuming it is 1 / 8 unit) can be the same as the precision of storing the motion vector. The case where the precision of representing the motion vector is different from the precision of storing the motion vector is also described. However, for the sake of easy description, in this example, the description is made under the assumption that the precision of representing the motion vector is the same as the precision of storing the motion vector, and the content about it can be extended and applied to the case where the precision of representing the motion vector is different from the precision of storing the motion vector.
[0660] In the example, since differential components are not generated in the skip mode (or merge mode), motion vectors can be stored according to a preset motion vector storage unit. In other words, since the accuracy of the predicted motion vector (because prediction is performed using the motion vectors stored in neighboring blocks) is the same as the preset motion vector accuracy, the predicted motion vector can be considered as a reconstructed motion vector and stored without change (as the motion vector of the target block).
[0661] In the example, differential components are generated in the competition mode, and the differential motion vector accuracy can be the same as or different from the preset motion vector storage unit. Therefore, motion vectors can be stored according to the preset motion vector storage unit. In other words, since the accuracy of the predicted motion vector can be the same as or different from the accuracy of the differential motion vector, a process of matching the accuracy of the differential motion vector with the accuracy of the predicted motion vector can be performed before storing the predicted motion vector in the memory.
[0662] The differential motion vector accuracy is equal to or lower than the motion vector accuracy. Therefore, in this case, a process of matching the accuracy of the differential motion vector with the accuracy of the predicted motion vector may be required (e.g., multiplication, division, rounding / down-rounding / up-rounding, shift operations, etc. In this example, a left shift operation). In this example, it is assumed that integer (1) unit, 1 / 4 unit, and 1 / 8 unit are included as candidates for the differential motion vector accuracy.
[0663] Assume that each candidate has the following settings.
[0664] mvd_pres_idx = 0 (1 / 1 unit) -> mvd_shift = 3 (number of shifts)
[0665] mvd_pres_idx = 1 (1 / 2 unit) -> mvd_shift = 2 (number of shifts)
[0666] mvd_pres_idx = 2 (1 / 8 unit) -> mvd_shift = 0 (number of shifts)
[0667] Motion vectors can be stored in the memory through the following process.
[0668] if(adaptive_mvd_precision_enabled_flag == 0)
[0669] {
[0670] mv_x = mvp_x + mvd_x
[0671] mv_y = mvp_y + mvd_y
[0672] }
[0673] else
[0674] {
[0675] mv_x = mvp_x+(mvd_x << mvd_shift_x)
[0676] mv_y = mvp_y+(mvd_y << mvd_shift_y)
[0677] }
[0678] In the example, mv represents the reconstructed motion vector, mvp represents the predicted motion vector, and mvd represents the differential motion vector. _x and _y following mvd_shift are used for classification under the precision settings applied to the differential motion vector configuration components respectively, and they can be removed and represented with the precision settings usually applied to the differential motion vector. Additionally, for a fixed differential motion vector precision setting, the reconstructed motion vector can be obtained by adding the predicted motion vector and the differential motion vector. This example can be the following situation: the motion estimation and compensation processing, the representation unit (precision) and storage unit of the motion vector are fixed to one precision (for example, they are unified to 1 / 4 unit and used, and only the differential motion vector precision is adaptively determined).
[0679] In summary, when having adaptive precision, the process of matching the precision can be performed before storing the motion vector.
[0680] Additionally, it is described above that the motion vector precision can be different from the precision to be stored, and in this case, the motion vector can be stored in the memory through the following process. In this example, the description is made by assuming the case where the motion estimation and compensation processing, the representation unit and storage unit of the motion vector are determined to be multiple precisions (since it is set differently according to the motion information coding mode, this example describes the case where the skip mode, merge mode and competition mode are mixed and used).
[0681] if (cu_skip_flag || merge_flag)
[0682] {
[0683] mv_x = mvp_x
[0684] mv_y = mvp_y
[0685] }
[0686] else
[0687] {
[0688] if(adaptive_mvd_precision_enabled_flag==0)
[0689] {
[0690] mv_x=mvp_x+(mvd_x<<diff_prec)
[0691] mv_y=mvp_y+(mvd_y<<diff_prec)
[0692] }
[0693] else
[0694] {
[0695] mv_x=mvp_x+{(mvd_x<<mvd_shift_x)<<diff_prec}
[0696] mv_y=mvp_y+{(mvd_y<<mvd_shift_y)<<diff_prec}
[0697] }
[0698] }
[0699] In this example, for the skip mode (in this case, cu_skip_flag = 1) and the merge mode (in this case, merge_flag = 1), the following is assumed: the precision in motion estimation and compensation processing is 1 / 16 unit, the motion vector precision is 1 / 4 unit, and the storage unit is 1 / 16 unit; for the competing mode, the following is assumed: the precision in motion estimation and compensation processing is 1 / 4 unit, the motion vector precision is 1 / 4 unit, and the storage unit is 1 / 16 unit. As in this example, the processing for matching the motion vector storage unit can be additionally performed.
[0700] In this example, mvp is the predicted motion vector obtained from the neighboring block, and mvp refers to the stored motion vector information of the neighboring block. Therefore, mvp has been set to the motion vector storage unit (1 / 16). For mvd, although the processing for matching the motion vector precision of the target block (1 / 4 in this example) with the differential motion vector precision (assumed to be an integer unit in this example) is performed, the precision adjustment processing for mvp (1 / 16 in this example) can be retained. In other words, the processing for matching the precision regarding the motion vector storage unit (in this example, by the left shift operation of diff_prec) can be performed.
[0701] As described in some assumptions (e.g., the accuracy of the motion information encoding mode is set differently, etc.), the above example can be described by obtaining information that is still applied equally or similarly by the example even when changed by various assumptions.
[0702] Alternatively, a candidate group of differential motion vector precisions can be determined according to the interpolation precision (or motion vector precision). For example, the differential motion vectors that can be generated when the interpolation precision is an integer unit can configure the candidate group in integer units of 1 or greater (e.g., 1, 2, 4 units); in the case of a fractional unit, the differential motion vectors that can be generated when the interpolation precision is an integer unit can configure the candidate group in units equal to or greater than the interpolation precision (e.g., 2, 1, 1 / 4, 1 / 8 units. This example assumes that 1 / 8 is the interpolation precision).
[0703] The method according to the present disclosure can be recorded in a computer-readable medium after being implemented in the form of program instructions executable by various computer devices. The computer-readable medium can include program instructions, data files, data structures, etc. alone or in combination. The program instructions recorded in the computer-readable medium can be specifically designed and configured for the present disclosure, or can be obtained after notifying those skilled in the field of computer software.
[0704] Examples of computer-readable media can include hardware devices such as ROM, RAM, flash memory, etc., which are specifically configured to store and execute program instructions. In addition to machine language code generated by a compiler, examples of program instructions can also include high-level language code that can be run by a computer with an interpreter. The above-mentioned hardware devices can be configured to operate as at least one software module to perform the actions of the present disclosure, and vice versa.
[0705] In addition, the above-mentioned method or device can be implemented after all or part of such configurations or functions are combined or separated.
[0706] Although described above by referring to the desired embodiments of the present disclosure, those skilled in the relevant technical field can understand that various modifications and changes can be made to the present disclosure without departing from the spirit and scope of the present disclosure set forth in the appended claims.
[0707] Industrial Applicability
[0708] The present disclosure can be used for encoding / decoding video signals. < / abt> < / sbt> < / qt>
Claims
1. An inter - frame prediction method performed by an image decoding device, comprising: Determining a merge candidate list for a current block including merge candidates; Selecting a merge candidate from the merge candidate list; And Performing inter - frame prediction based on the motion vector and reference picture of the selected merge candidate, and Wherein, the merge candidate list includes spatial merge candidates and temporal merge candidates, and When the number of the spatial merge candidates and the temporal merge candidates is less than the maximum number of the merge candidate list, the merge candidate list further includes statistical merge candidates derived from previously decoded blocks that belong to the current picture and are not adjacent to the current block, and When the number of the spatial merge candidates, the temporal merge candidates, and the statistical merge candidates is less than the maximum number of the merge candidate list, the merge candidate list further includes combined merge candidates derived from two pre - included merge candidates.
2. The method according to claim 1, wherein The statistical merge list includes the motion information of a predetermined number of previously decoded blocks, and the statistical merge list is managed using the FIFO (First In First Out) method.
3. The method according to claim 2, wherein The statistical merge list includes the motion information of previously decoded blocks and does not include the motion information of previously decoded sub - blocks.
4. The method according to claim 1, wherein the motion vector of the combined merge candidate is derived by averaging the motion vectors of a first merge candidate and a second merge candidate in the merge candidate list.
5. An inter - frame prediction method performed by an image encoding device, comprising: Determining a merge candidate list for a current block including merge candidates; Selecting a merge candidate from the merge candidate list; And Performing inter - frame prediction based on the motion vector and reference picture of the selected merge candidate, and Wherein, the merge candidate list includes spatial merge candidates and temporal merge candidates, and When the number of the spatial merge candidates and the temporal merge candidates is less than the maximum number of the merge candidate list, the merge candidate list further includes statistical merge candidates derived from previously decoded blocks that belong to the current picture and are not adjacent to the current block, and When the number of the spatial merge candidates, the temporal merge candidates, and the statistical merge candidates is less than the maximum number of the merge candidate list, the merge candidate list further includes combined merge candidates derived from two pre - included merge candidates.
6. A non - transitory computer - readable recording medium that stores a bitstream generated by an inter - frame prediction method performed by an image encoding device, the method comprising: Determining a merge candidate list for a current block including merge candidates; Selecting a merge candidate from the merge candidate list; And Performing inter - frame prediction based on the motion vector and reference picture of the selected merge candidate, and Wherein, the merge candidate list includes spatial merge candidates and temporal merge candidates, and When the number of the spatial merge candidates and the temporal merge candidates is less than the maximum number of the merge candidate list, the merge candidate list further includes statistical merge candidates derived from previously decoded blocks that belong to the current picture and are not adjacent to the current block, and In the case where the number of the spatial merge candidates, the temporal merge candidates, and the statistical merge candidates is less than the maximum number of the merge candidate list, the merge candidate list further includes combined merge candidates derived from two pre-included merge candidates.