Method for encoding / decoding image signal, and device for the same
By utilizing an inter-region motion information table to derive merge candidates, the method addresses inefficiencies in existing video compression technologies, enhancing inter prediction efficiency and improving video encoding/decoding performance.
Patent Information
- Application Number
- JP2025135519
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-11-27
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-05
AI Technical Summary
Existing video compression technologies, such as HEVC, are reaching their performance limits as they struggle to efficiently derive merge candidates for video encoding/decoding, leading to inefficiencies in inter prediction.
A method is introduced to derive merge candidates using an inter-region motion information table, which includes adding inter-region merge candidates to the merge candidate list based on spatial and temporal merge candidates, and updating these candidates dynamically during the encoding/decoding process.
This approach enhances inter prediction efficiency by providing alternative merge candidates beyond those derived from neighboring blocks, improving overall video encoding/decoding performance.
Smart Images

Figure 2025166201000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a video signal encoding / decoding method and apparatus therefor. [Background technology]
[0002] As display panels continue to grow larger, the demand for higher-quality video services is increasing. The biggest problem with high-quality video services is the significant increase in data volume. To address this issue, active research is being conducted to improve video compression rates. As a representative example, in 2009, the Motion Picture Experts Group (MPEG) and the Video Coding Experts Group (VCEG) of the International Telecommunication Union-Telecommunication (ITU-T) formed the Joint Collaborative Team on Video Coding (JCT-VC). JCT-VC proposed High Efficiency Video Coding (HEVC), a compression standard with approximately twice the compression performance of H.264 / AVC, and the standard was approved on January 25, 2013. With the rapid development of high-quality video services, HEVC's performance is gradually reaching its limits. Summary of the Invention [Problem to be solved by the invention]
[0003] The present invention aims to provide a method for deriving a merging candidate other than a merging candidate derived from a candidate block adjacent to a current block in encoding / decoding a video signal, and an apparatus for performing the method.
[0004] SUMMARY OF THE INVENTION The present invention aims to provide a method for deriving merge candidates using an inter-region motion information table in encoding / decoding a video signal, and an apparatus for performing the method.
[0005] The present invention aims to provide a method for deriving merge candidates for blocks included in a merge processing area in encoding / decoding a video signal, and an apparatus for performing the method.
[0006] The technical problems to be solved by the present invention are not limited to the technical problems mentioned above, and from the following description, those skilled in the art will be able to clearly understand other technical problems not mentioned above. [Means for solving the problem]
[0007] A video signal decoding / encoding method according to the present invention includes generating a merge candidate list for a first block, selecting one of the merge candidates included in the merge candidate list, and performing motion compensation for the first block based on motion information of the selected merge candidate. In this case, inter-region merge candidates included in an inter-region motion information table may be added to the merge candidate list based on the number of spatial merge candidates and temporal merge candidates included in the merge candidate list.
[0008] In the video signal decoding / encoding method according to the present invention, the inter region motion information table may include inter region merge candidates derived based on motion information of a block decoded before the first block, and in this case, the inter region motion information table may not be updated based on motion information of a second block included in the same merge processing region as the first block.
[0009] In the video signal decoding / encoding method according to the present invention, when the first block is included in a merge processing area, a temporary merge candidate derived based on the motion information of the first block is added to a temporary motion information table, and when decoding of all blocks included in the merge processing area is completed, the temporary merge candidate can be updated to the inter-area motion information table.
[0010] In the video signal decoding / encoding method according to the present invention, it is possible to determine whether to add a first inter region merge candidate included in the inter region motion information table to the merge candidate list based on a determination result as to whether the first inter region merge candidate is the same as at least one merge candidate included in the merge candidate list.
[0011] In the video signal decoding / encoding method according to the present invention, the determination can be performed by comparing the first inter region merging candidate with at least one merging candidate having an index value less than or equal to a threshold.
[0012] In the video signal decoding / encoding method according to the present invention, when it is determined that a merge candidate identical to the first inter region merge candidate exists, the first inter region merge candidate is not added to the merge candidate list, and whether to add the second inter region merge candidate to the merge candidate list may be determined based on a determination result of whether the first inter region merge candidate included in the inter region motion information table is identical to at least one merge candidate included in the merge candidate list. In this case, it is possible to omit determining whether the second inter region merge candidate is identical to the first inter region merge candidate.
[0013] The above briefly described features of the present invention are merely exemplary embodiments of the detailed description of the present invention that follows, and are not intended to limit the scope of the present invention. [Effects of the Invention]
[0014] According to the present invention, inter prediction efficiency can be improved by providing a method for deriving merge candidates other than those derived from candidate blocks neighboring the current block.
[0015] According to the present invention, by providing a method for deriving merge candidates using an inter-region motion information table, inter prediction efficiency can be improved.
[0016] According to the present invention, a method for deriving merge candidates for blocks included in a merge processing region is provided, thereby improving inter prediction efficiency.
[0017] The technical effects obtained by the present invention are not limited to the effects mentioned above, and those skilled in the art will be able to clearly understand other technical effects not mentioned from the following description. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a block diagram of a video encoder according to an embodiment of the present invention; [Figure 2] FIG. 1 is a block diagram of a video decoder according to an embodiment of the present invention. [Figure 3] 1 is a diagram illustrating a basic coding tree unit according to an embodiment of the present invention. [Figure 4] 1 is a diagram showing various division forms of a coding block; [Figure 5] 10 is a diagram showing a division mode of a coding tree unit. [Figure 6] 1 is a flowchart of an inter prediction method according to an embodiment of the present invention. [Figure 7] 1 is a diagram illustrating non-linear movement of an object. [Figure 8] 1 is a flowchart of an affine motion-based inter prediction method according to an embodiment of the present invention. [Figure 9] 10 is a diagram illustrating affine seed vectors for different affine motion models; [Figure 10] 10 is a diagram showing affine vectors of sub-blocks under a four-parameter motion model. [Figure 11] 10 is a flowchart of a process for deriving motion information for a current block under merge mode. [Figure 12] 10 is a diagram showing candidate blocks for deriving merge candidates; [Figure 13] 1 is a diagram showing the position of a reference sample. [Figure 14] 10 is a diagram showing candidate blocks for deriving merge candidates; [Figure 15] 10 is a diagram showing an example in which the position of a reference sample is changed; [Figure 16] 10 is a diagram showing an example in which the position of a reference sample is changed; [Figure 17] 10 is a flowchart illustrating an update state of the inter-region motion information table. [Figure 18] 10 is a diagram illustrating an example of updating an inter-region merge candidate list. [Figure 19] 10 is a diagram illustrating an example in which the indexes of stored inter-region merge candidates are updated. [Figure 20] 10 is a diagram showing the position of a representative sub-block; [Figure 21] 10 is a diagram illustrating an example in which an inter-region motion information table is generated for each inter-prediction mode. [Figure 22] 10 is a diagram illustrating an example in which inter-region merge candidates included in a long-term motion information table are added to a merge candidate list. [Figure 23] 10 is a diagram illustrating an example in which redundancy detection is performed on only some of the merge candidates. [Figure 24] 10 is a diagram illustrating an example in which redundant detection with a specific merge candidate is omitted. [Figure 25]10 is a diagram illustrating an example in which a candidate block included in the same merge processing area as the current block is set as unavailable as a merge candidate; [Figure 26] 10 is a diagram illustrating a temporary motion information table. [Figure 27] 10 is a diagram illustrating an example of merging an inter-region motion information table and a temporary motion information table. [Figure 28] 1 is a flowchart of an intra prediction method according to an embodiment of the present invention. [Figure 29] 1 is a diagram of the reference samples included in each reference sample line. [Figure 30] 1 is a diagram illustrating intra-prediction modes. [Figure 31] 1 is a diagram showing an example of a one-dimensional array in which reference samples are arranged in a line. [Figure 32] 1 is a diagram showing an example of a one-dimensional array in which reference samples are arranged in a line. [Figure 33] 10 is a diagram illustrating angles formed by directional intra prediction modes with a line parallel to the x-axis. [Figure 34] 10 is a diagram illustrating how prediction samples are obtained when a current block is non-square. [Figure 35] 10 is a diagram illustrating a wide-angle intra prediction mode. [Figure 36] 1 is a diagram showing an embodiment in which a PDPC is applied. [Figure 37] 10 is a diagram illustrating an example in which a second merging candidate is identified in consideration of the search order of candidate blocks. [Figure 38] 10 is a diagram illustrating an example in which a first merging candidate and a second merging candidate are selected from merging candidates derived from non-adjacent blocks. [Figure 39] 10 is a diagram illustrating an example in which a weight applied to a prediction block is determined based on the type of a candidate block; [Figure 40] 10 is a diagram illustrating an example in which a non-affine merge candidate is set as a second merge candidate instead of an affine merge candidate. [Figure 41]10 is a diagram illustrating an example in which merge candidates are replaced. [Figure 42] 10 is a flowchart illustrating a process for determining the strength of a block. [Figure 43] 1 is a diagram illustrating predefined filter candidates. DETAILED DESCRIPTION OF THE INVENTION
[0019] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0020] Image encoding and decoding is performed on a block-by-block basis. For example, encoding / decoding processes such as transform, quantization, prediction, in-loop filtering, or reconstruction may be performed on a coding block, a transform block, or a prediction block.
[0021] Hereinafter, the block to be coded / decoded is referred to as a “current block.” For example, the current block may refer to a coding block, a transformation block, or a prediction block according to the current coding / decoding process step.
[0022] Furthermore, as used herein, the term "unit" refers to a basic unit for performing a specific encoding / decoding process, and "block" can be understood to refer to a sample array of a predetermined size. Unless otherwise specified, "block" and "unit" can be used interchangeably. For example, in the embodiments described below, a coding block and a coding unit can be understood to have the same meaning.
[0023] FIG. 1 is a block diagram of a video encoder according to an embodiment of the present invention.
[0024] Referring to FIG. 1, the video encoding device 100 may include an image division unit 110, prediction units 120 and 125, a transformation unit 130, a quantization unit 135, a rearrangement unit 160, an entropy encoding unit 165, an inverse quantization unit 140, an inverse transform unit 145, a filter unit 150, and a memory 155.
[0025] 1 are shown independently to indicate different characteristic functions in the video encoding device, and do not mean that each component is composed of separate hardware or a single software unit. That is, for the sake of convenience, each component may be arranged such that at least two of the components are combined into one component, or one component may be divided into multiple components to perform its function. The scope of protection of the present invention includes both an embodiment in which these components are integrated and an embodiment in which the components are separated, provided that the embodiment does not deviate from the essence of the present invention.
[0026] In addition, some components are not necessary to perform the essential functions of the present invention, but are merely optional components for improving performance. The present invention can be realized by including only the components necessary to embody the essence of the present invention, excluding components used to improve performance, and a structure including only the necessary components, excluding optional components for improving performance, is also included in the scope of protection of the present invention.
[0027] The image division unit 110 can divide an input image into at least one processing unit. In this case, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The image division unit 110 can divide one image into a plurality of coding units, combinations of prediction units and transform units, and select one coding unit, combination of prediction units and transform units based on a predetermined criterion (e.g., a cost function) to code the image.
[0028] For example, an image can be divided into multiple coding units. A recursive tree structure, such as a quad tree structure, can be used to divide the coding units into an image. A coding unit that is divided into other coding units using a single image or largest coding unit as the root can have as many child nodes as the number of divided coding units. Depending on certain restrictions, a coding unit that is not further divided becomes a leaf node. In other words, assuming that only square division is possible for a coding unit, one coding unit can be divided into a maximum of four different coding units.
[0029] Hereinafter, in the embodiments of the present invention, a coding unit may be used to mean a unit for performing encoding or a unit for performing decoding.
[0030] The prediction units may be divided into at least one square or rectangular shape of the same size within one coding unit, or may be divided such that one of the divided prediction units within one coding unit has a different shape and / or size from the other prediction units.
[0031] When generating a prediction unit for performing intra prediction based on a coding unit, if the coding unit is not the smallest coding unit, intra prediction can be performed without dividing the coding unit into a plurality of NxN prediction units.
[0032] The prediction units 120 and 125 may include an inter prediction unit 120 that performs inter prediction and an intra prediction unit 125 that performs intra prediction. It is possible to determine whether to use inter prediction or intra prediction for a prediction unit, and to determine specific information (e.g., intra prediction mode, motion vector, reference image, etc.) related to each prediction method. In this case, the processing unit for performing prediction may differ from the processing unit for determining the prediction method and specific details. For example, the prediction method and prediction mode may be determined for each prediction unit, and prediction may be performed for each transform unit. Residual values (residual blocks) between the generated prediction block and the original block may be input to the transform unit 130. In addition, prediction mode information, motion vector information, etc. used for prediction may be coded in the entropy coding unit 165 together with the residual values and transmitted to the decoder. When a specific coding mode is used, it is also possible to code the original block as is and transmit it to the decoder without generating a prediction block via the prediction units 120 and 125.
[0033] The inter prediction unit 120 may predict a prediction unit based on information of at least one image between an image preceding the current image and an image following the current image, and in some cases, may predict a prediction unit based on information of a part of an area in the current image for which coding has been completed. The inter prediction unit 120 may include a reference image interpolation unit, a motion prediction unit, and a motion compensation unit.
[0034] The reference image interpolator receives reference image information from memory 155 and generates sub-integer pixel information from the reference image. In the case of luminance pixels, a DCT-based 8-tab interpolation filter with different filter coefficients may be used to generate sub-integer pixel information in 1 / 4 pixel units. In the case of color difference signals, a DCT-based 4-tab interpolation filter with different filter coefficients may be used to generate sub-integer pixel information in 1 / 8 pixel units.
[0035] The motion prediction unit may perform motion prediction based on the reference image interpolated by the reference image interpolator. Various methods may be used to calculate a motion vector, such as a full search-based block matching algorithm (FBMA), a three-step search (TSS), or a new three-step search algorithm (NTS). The motion vector may have a motion vector value in half or quarter pixel units based on the interpolated pixel. The motion prediction unit may predict the current prediction unit using different motion prediction methods. Various motion prediction methods may be used, such as a skip method, a merge method, an advanced motion vector prediction (AMVP) method, and an intra block copy method.
[0036] The intra prediction unit 125 can generate a prediction unit based on reference pixel information surrounding the current block, which is pixel information within the current image. When the blocks surrounding the current prediction unit are inter-predicted blocks and the reference pixels are inter-predicted pixels, the reference pixels included in the inter-predicted blocks can be replaced with reference pixel information of the surrounding blocks that have been intra-predicted. That is, when reference pixels are unavailable, the unavailable reference pixel information can be replaced with at least one of the available reference pixels.
[0037] In intra prediction, prediction modes may include a directional prediction mode that uses reference pixel information according to a prediction direction, and a non-directional mode that does not use directional information when predicting. A mode for predicting luma information and a mode for predicting chroma information may be different, and intra prediction mode information used to predict luma information or predicted luma signal information may be applied to predict chroma information.
[0038] When performing intra prediction, if the size of the prediction unit and the size of the transform unit are the same, intra prediction for the prediction unit can be performed based on the pixel on the left, the pixel on the upper left, and the pixel on the top of the prediction unit. However, when performing intra prediction, if the size of the prediction unit and the size of the transform unit are different, intra prediction can be performed using one reference pixel based on the transform unit. Also, intra prediction using NxN division can be used only for the smallest coding unit.
[0039] The intra prediction method may generate a predicted block after applying an adaptive intra smoothing (AIS) filter to reference pixels according to a prediction mode. The type of AIS filter applied to the reference pixels may vary. To perform the intra prediction method, the intra prediction mode of a current prediction unit may be predicted from the intra prediction modes of prediction units surrounding the current prediction unit. When predicting the prediction mode of the current prediction unit using mode information predicted from surrounding prediction units, if the intra prediction modes of the current prediction unit and the surrounding prediction units are the same, information indicating that the prediction modes of the current prediction unit and the surrounding prediction units are the same may be transmitted using predetermined flag information. If the prediction modes of the current prediction unit and the surrounding prediction units are different, the prediction mode information of the current block may be encoded by performing entropy encoding.
[0040] Furthermore, a residual block including residual value information, which is a difference value between the prediction unit that performed the prediction and the original block of the prediction unit, can be generated based on the prediction unit generated by the prediction units 120 and 125. The generated residual block can be input to the conversion unit 130.
[0041] The transform unit 130 may use a transform method such as a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), or a KL Transform (KLT) to transform a residual block including residual value information of the original block and the prediction unit generated by the predictor 120 or 125. Whether to apply the DCT, the DST, or the KLT to transform the residual block may be determined based on intra prediction mode information of the prediction unit used to generate the residual block.
[0042] The quantization unit 135 quantizes the values transformed into the frequency domain by the transform unit 130. The quantization coefficients may vary depending on the block or the importance of the image. The values calculated by the quantization unit 135 may be provided to the inverse quantization unit 140 and the rearrangement unit 160.
[0043] The rearrangement unit 160 can perform rearrangement of coefficient values on the quantized residual values.
[0044] The rearrangement unit 160 may convert two-dimensional block shape coefficients into one-dimensional vector format through a coefficient scanning method. For example, the rearrangement unit 160 may convert two-dimensional block shape coefficients into one-dimensional vector format by scanning from DC coefficients to high-frequency region coefficients using a zig-zag scan method. Instead of zig-zag scan, vertical scan, in which two-dimensional block shape coefficients are scanned along the column direction, or horizontal scan, in which two-dimensional block shape coefficients are scanned along the row direction, may be used depending on the size of the transform unit and the intra prediction mode. That is, it may be determined which scanning method to use among zig-zag scan, vertical scan, or horizontal scan depending on the size of the transform unit and the intra prediction mode.
[0045] The entropy encoding unit 165 may perform entropy encoding based on the value calculated by the rearrangement unit 160. The entropy encoding may use various encoding methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC).
[0046] The entropy coding unit 165 can encode various information such as residual value coefficient information and block type information of the coding unit, prediction mode information, division unit information, prediction unit information and transmission unit information, motion vector information, reference frame information, block interpolation information, filtering information, etc. from the rearrangement unit 160 and the prediction units 120, 125.
[0047] The entropy encoding unit 165 can entropy encode the coefficient values of the coding unit input to the rearrangement unit 160 .
[0048] The inverse quantization unit 140 and the inverse transform unit 145 inversely quantize the values quantized by the quantization unit 135 and inversely transform the values transformed by the transform unit 130. The residual values generated by the inverse quantization unit 140 and the inverse transform unit 145 can be combined with prediction units predicted by the motion estimation unit, motion compensation unit, and intra prediction unit included in the prediction units 120 and 125 to generate reconstructed blocks.
[0049] The filter unit 150 may include at least one of a deblocking filter, an offset correction unit, and an adaptive loop filter (ALF).
[0050] A deblocking filter can remove block artifacts caused by boundaries between blocks in a reconstructed image. To determine whether to deblock, it can be determined whether to apply a deblocking filter to a current block based on the pixels contained in several columns or rows of the block. When a deblocking filter is applied to a block, a strong filter or a weak filter can be applied depending on the strength of the deblocking filtering required. In addition, applying a deblocking filter can enable horizontal filtering and vertical filtering to be processed in parallel when performing vertical filtering and horizontal filtering.
[0051] The offset correction unit may correct the offset between the deblocked image and the original image on a pixel-by-pixel basis. To perform offset correction on a specific image, pixels included in the image may be divided into a specific number of regions, and regions to be offset may be determined, and an offset may be applied to the corresponding regions. Alternatively, an offset may be applied by considering edge information of each pixel.
[0052] Adaptive Loop Filtering (ALF) can be performed based on a comparison between a filtered reconstructed image and the original image. Pixels included in an image can be divided into predetermined groups, and a filter to be applied to each corresponding group can be determined to perform differential filtering for each group. Information regarding whether to apply ALF and a luminance signal can be transmitted for each ALF coding unit (CU), and the shape and filter coefficients of the ALF filter to be applied vary for each block. It is also possible to apply the same type (fixed type) of ALF filter regardless of the characteristics of the block to which it is applied.
[0053] The memory 155 can store the reconstructed blocks or images calculated by the filter unit 150, and the stored reconstructed blocks or images can be provided to the prediction units 120 and 125 when performing inter-prediction.
[0054] FIG. 2 is a block diagram of a video decoder according to an embodiment of the present invention.
[0055] Referring to FIG. 2, the video decoder 200 may include an entropy decoding unit 210, a rearrangement unit 215, an inverse quantization unit 220, an inverse transform unit 225, prediction units 230 and 235, a filter unit 240, and a memory 245.
[0056] When a video bitstream is input from a video encoder, the input bitstream can be decoded by following the reverse steps of the video encoder.
[0057] The entropy decoding unit 210 may perform entropy decoding in a step opposite to the step of entropy encoding performed by the entropy encoding unit of the video encoder. For example, various methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) may be applied in accordance with the method performed by the video encoder.
[0058] The entropy decoding unit 210 can decode information related to intra-prediction and inter-prediction performed by the encoder.
[0059] The rearrangement unit 215 may rearrange the bitstream entropy decoded by the entropy decoding unit 210 in the encoding unit based on a rearrangement method. The rearrangement unit 215 may rearrange coefficients represented in the form of one-dimensional vectors by reconstructing them into coefficients in the form of two-dimensional blocks. The rearrangement unit 215 may receive information related to coefficient scanning performed by the encoding unit and rearrange the coefficients using a method of scanning in reverse based on the scanning order performed by the corresponding encoding unit.
[0060] The inverse quantization unit 220 can perform inverse quantization based on the quantization parameter provided by the encoder and the coefficient values of the rearranged block.
[0061] The inverse transform unit 225 may perform inverse transforms, i.e., inverse DCT, inverse DST, and inverse KLT, on the transforms, i.e., DCT, DST, and KLT, performed by the transform unit on the quantization result performed by the video encoder. The inverse transform may be performed based on a transmission unit determined by the video encoder. The inverse transform unit 225 of the video decoder may selectively perform a transform method (e.g., DCT, DST, or KLT) according to multiple information such as a prediction method, a size of a current block, and a prediction direction.
[0062] The prediction units 230, 235 can generate the prediction blocks based on relevant information related to the generation of the prediction blocks provided by the entropy decoding unit 210 and previously decoded block or image information provided by the memory 245.
[0063] As described above, when intra prediction is performed, as in the operation of a video encoder, if the size of the prediction unit and the size of the transform unit are the same, intra prediction for the prediction unit is performed based on the pixel on the left, the pixel on the top left, and the pixel on the top. However, if the size of the prediction unit and the size of the transform unit are different, intra prediction can be performed using one reference pixel based on the transform unit. Also, intra prediction using NxN division can be used only for the smallest coding unit.
[0064] The prediction units 230 and 235 may include a prediction unit determination unit, an inter prediction unit, and an intra prediction unit. The prediction unit determination unit receives various information, such as prediction unit information input from the entropy decoding unit 210, prediction mode information for the intra prediction method, and motion prediction-related information for the inter prediction method, to classify prediction units by the current coding unit and determine whether inter prediction or intra prediction is to be performed for the prediction unit. The inter prediction unit 230 may perform inter prediction for the current prediction unit based on information included in at least one image, including a previous image or a subsequent image, of the current image including the current prediction unit, using information necessary for inter prediction of the current prediction unit provided from the video encoder. Alternatively, the inter prediction unit may perform inter prediction based on information of a subregion reconstructed in the current image including the current prediction unit.
[0065] To perform inter prediction, it is possible to determine, based on the coding unit, which of the following methods is the motion prediction method of the prediction unit included in the corresponding coding unit: skip mode, merge mode, motion vector prediction mode (AMVP mode), or intra block copy mode.
[0066] The intra prediction unit 235 may generate a prediction block based on pixel information within the current image. If the prediction unit is a prediction unit for which intra prediction is performed, the intra prediction may be performed based on intra prediction mode information of the prediction unit provided from the video encoder. The intra prediction unit 235 may include an adaptive intra smoothing (AIS) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is a part that filters reference pixels of the current block and may determine whether to apply a filter according to the prediction mode of the current prediction unit. The AIS filtering may be performed on reference pixels of the current block using the prediction mode of the prediction unit and AIS filter information provided from the video encoder. If the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.
[0067] The reference pixel interpolation unit can interpolate the reference pixels to generate reference pixels of integer or smaller pixel units when the prediction mode of the prediction unit is a prediction unit that performs intra prediction based on pixel values to interpolate the reference pixels. When the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating reference pixels, the reference pixels may not be interpolated. When the prediction mode of the current block is DC mode, the DC filter can generate a prediction block by filtering.
[0068] The reconstructed block or image may be provided to a filter unit 240. The filter unit 240 may include a deblocking filter, an offset correction unit, and an ALF.
[0069] The video decoder may receive, from the video encoder, information on whether a deblocking filter has been applied to a corresponding block or image, and, if a deblocking filter has been applied, information on whether a strong filter or a weak filter has been applied. The deblocking filter of the video decoder may receive the deblocking filter-related information provided from the video encoder and perform deblocking filtering on the corresponding block by the video decoder.
[0070] The offset correction unit may perform offset correction on the reconstructed image based on the type of offset correction and offset value information applied to the image when it was encoded.
[0071] The ALF can be applied to a coding unit based on information related to whether to apply the ALF, ALF coefficient information, etc., provided by the encoder. Such ALF information can be provided by being included in a specific parameter set.
[0072] The memory 245 can store the reconstructed image or block so that it can be used as a reference image or block, and can provide the reconstructed image to an output unit.
[0073] FIG. 3 is a diagram illustrating a basic coding tree unit according to an embodiment of the present invention.
[0074] A coding block of the largest size can be defined as a coding tree block. An image is divided into multiple coding tree units (CTUs). A coding tree unit is a coding unit of the largest size and is also called a largest coding unit (LCU). Figure 3 shows an example of an image divided into multiple coding tree units.
[0075] The size of a coding tree unit can be defined at the picture level or the sequence level, and therefore, information representing the size of a coding tree unit can be signaled via a picture parameter set or a sequence parameter set.
[0076] For example, the coding tree unit size for an image in a sequence can be set to 128x128. Alternatively, the coding tree unit size can be determined to be either 128x128 or 256x256 at the image level. For example, the coding tree unit size for a first image can be set to 128x128, and the coding tree unit size for a second image can be set to 256x256.
[0077] Coding blocks may be generated by dividing the coding tree unit. The coding block represents a basic unit for encoding / decoding processing. For example, prediction or transformation may be performed for each coding block, or a predictive coding mode may be determined for each coding block. Here, the predictive coding mode represents a method for generating a predicted image. For example, the predictive coding mode may include intra-frame prediction (intra prediction), inter-frame prediction (inter prediction), current picture referencing (CPR) or intra block copy (IBC), or combined prediction. For a coding block, a predictive block associated with the coding block may be generated using at least one predictive coding mode from intra prediction, inter prediction, current picture referencing, and combined prediction.
[0078] Information indicating the predictive coding mode of the current block may be transmitted via a bitstream signal. For example, the information may indicate whether the predictive coding mode is an intra mode or a one-bit flag indicating an inter mode. Only when the predictive coding mode of the current block is determined to be an inter mode, can current picture reference or hybrid prediction be used.
[0079] The current image reference is used to set the current image as a reference image and obtain a prediction block for the current block from an encoded / decoded region in the current image. Here, the current image refers to an image including the current block. Information indicating whether the current image reference is applied to the current block may be transmitted via a bitstream signal. For example, the information may be a 1-bit flag. If the flag is true, the prediction coding mode of the current block may be determined as the current image reference, and if the flag is false, the prediction mode of the current block may be determined as inter prediction.
[0080] Alternatively, the predictive coding mode of the current block may be determined based on a reference image index. For example, if the reference image index points to the current image, the predictive coding mode of the current block may be determined as current image reference. If the reference image index points to another image other than the current image, the predictive coding mode of the current block may be determined as inter prediction. That is, current image reference is a prediction method using information on an encoded / decoded region in the current image, and inter prediction is a prediction method using information on another encoded / decoded image.
[0081] Hybrid prediction is a coding mode that combines two or more of intra prediction, inter prediction, and current image reference. For example, when hybrid prediction is applied, a first predicted block may be generated based on one of intra prediction, inter prediction, or current image reference, and a second predicted block may be generated based on another one of them. After the first predicted block and the second predicted block are generated, a final predicted block may be generated by averaging or weighted summing the first predicted block and the second predicted block. Information indicating whether hybrid prediction is applied may be transmitted via a bitstream signal. The information may be a 1-bit flag.
[0082] FIG. 4 is a diagram showing various division forms of a coding block.
[0083] A coding block partition can be divided into multiple coding blocks based on a quad-tree partition, a binary tree partition, or a triple-tree partition. A divided coding block can also be divided into multiple coding blocks again based on a quad-tree partition, a binary tree partition, or a triple-tree partition.
[0084] Quadtree partitioning refers to a partitioning technique that divides the current block partition into four blocks. As a result of quadtree partitioning, the current block partition can be divided into four square partitions ("SPLIT_QT" in Figure 4(a)).
[0085] Binary tree splitting refers to a splitting technique that splits the current block into two blocks. Splitting the current block into two blocks along the vertical direction (i.e., using a vertical line across the current block) can be referred to as vertical binary tree splitting, and splitting the current block into two blocks along the horizontal direction (i.e., using a horizontal line across the current block) can be referred to as horizontal binary tree splitting. As a result of the binary tree splitting, the current block can be split into two non-square partitions. Figure 4(b) "SPLIT_BT_VER" represents the vertical binary tree split result, and Figure 4(c) "SPLIT_BT_HOR" represents the horizontal binary tree split result.
[0086] Triple tree splitting refers to a splitting technique that splits the current block into three blocks. Splitting the current block into three blocks along the vertical direction (i.e., using two vertical lines crossing the current block) can be referred to as vertical triple tree splitting, and splitting the current block into three blocks along the horizontal direction (i.e., using two horizontal lines crossing the current block) can be referred to as horizontal triple tree splitting. As a result of the triple tree splitting, the current block can be split into three non-square partitions. In this case, the width / height of the partition located in the center of the current block may be twice the width / height of the other partitions. (d) "SPLIT_TT_VER" in Figure 4 represents the vertical triple tree split result, and (e) "SPLIT_TT_HOR" in Figure 4 represents the horizontal triple tree split result.
[0087] The number of times a coding tree unit is divided can be defined as the partitioning depth. The maximum partitioning depth of a coding tree unit can be determined depending on the sequence or image level. This means that the maximum partitioning depth of a coding tree unit may differ for each sequence or image.
[0088] Alternatively, the maximum split depth for each of the splitting techniques can be determined separately. As an example, the maximum split depth allowed for quad-tree splitting may be different from the maximum split depth allowed for binary tree and / or triple-tree splitting.
[0089] The encoder may signal information indicating at least one of the partitioning type and the partitioning depth of the current block via the bitstream, and the decoder may determine the partitioning type and the partitioning depth of the coding tree unit based on the information parsed from the bitstream.
[0090] FIG. 5 is a diagram illustrating a division mode of a coding tree unit.
[0091] Partitioning a coding block using partitioning techniques such as quad-tree partitioning, binary tree partitioning, and / or triple-tree partitioning can be referred to as multi-tree partitioning.
[0092] A coding block generated by applying multi-tree partitioning to a coding block can be called a sub-coding block. If the partitioning depth of a coding block is k, the partitioning depth of the sub-coding block is set to k+1.
[0093] Conversely, for a coding block with a division depth of k+1, a coding block with a division depth of k can be called a higher-order coding block.
[0094] The division type of the current coding block may be determined based on at least one of the division form of the upper coding block or the division type of the adjacent coding block. Here, the adjacent coding block is adjacent to the current coding block, and the adjacent coding block may include at least one of the top adjacent block, the left adjacent block, or the adjacent block adjacent to the upper left corner of the current coding block. Here, the division type may include at least one of whether to perform quad-tree division, whether to perform binary-tree division, the binary-tree division direction, whether to perform triple-tree division, or the triple-tree division direction.
[0095] To determine the split form of a coding block, information indicating whether the coding block is split can be signaled via the bitstream. The information is a 1-bit flag "split_cu_flag," which indicates that the coding block is split using the multi-tree splitting technique if the flag is true.
[0096] If split_cu_flag is true, information indicating whether the coding block is quad-tree split can be signaled via the bitstream. The information is a 1-bit flag "split_qt_flag", and if the flag is true, the coding block can be split into four blocks.
[0097] 5 shows that a coding tree unit is quad-tree divided into four coding blocks with a division depth of 1. It also shows that quad-tree division is applied again to the first and fourth coding blocks among the four coding blocks generated as a result of the quad-tree division. As a result, four coding blocks with a division depth of 2 can be generated.
[0098] Furthermore, by applying quadtree division to a coding block with a division depth of 2 again, a coding block with a division depth of 3 can be generated.
[0099] If quad-tree partitioning is not applied to a coding block, it may be determined whether to perform binary tree partitioning or triple tree partitioning on the coding block, taking into account at least one of the size of the coding block, whether the coding block is located on an image boundary, the maximum partition depth, or the partition shape of adjacent blocks. If it is determined that binary tree partitioning or triple tree partitioning is applied to the coding block, information indicating the partitioning direction may be signaled via a bitstream. The information may be a 1-bit flag "mtt_split_cu_vertical_flag." Based on the flag, it may be determined whether the partitioning direction is vertical or horizontal. In addition, information indicating whether binary tree partitioning or triple tree partitioning is applied to the coding block may be signaled via a bitstream. The information may be a 1-bit flag "mtt_split_cu_binary_flag." Based on the flag, it may be determined whether binary tree partitioning or triple tree partitioning is applied to the coding block.
[0100] As an example, in the example shown in Figure 5, vertical binary tree partitioning is applied to a coding block with a partitioning depth of 1, and vertical triple tree partitioning is applied to the left coding block of the coding blocks generated as a result of the partitioning, and vertical binary tree partitioning is applied to the right coding block.
[0101] Inter-prediction is a predictive coding mode that predicts a current block using information of a previous image. For example, a block located at the same position as the current block in the previous image (hereinafter referred to as a collocated block) may be set as a prediction block for the current block. Hereinafter, a prediction block generated based on a block located at the same position as the current block will be referred to as a collocated prediction block.
[0102] On the other hand, if an object that existed in a previous image moves to another location in a current image, the current block can be effectively predicted using the object's motion. For example, by comparing the previous image with the current image, the direction and size of the object's movement can be known, and a predicted block (or predicted image) of the current block can be generated taking into account the object's motion information. Hereinafter, the predicted block generated using the motion information can be referred to as a motion predicted block.
[0103] A residual block can be generated by subtracting a predicted block from a current block. In this case, if there is object motion, the energy of the residual block can be reduced by using a motion prediction block instead of a co-located prediction block, thereby improving the compression performance of the residual block.
[0104] As described above, generating a prediction block using motion information can be referred to as motion compensated prediction. In most inter predictions, a prediction block can be generated based on motion compensated prediction.
[0105] The motion information may include at least one of a motion vector, a reference image index, a prediction direction, or a bidirectional weight value index. The motion vector represents the movement direction and size of an object. The reference image index identifies a reference image of the current block among reference images included in a reference image list. The prediction direction refers to one of unidirectional L0 prediction, unidirectional L1 prediction, or bidirectional prediction (L0 prediction and L1 prediction). According to the prediction direction of the current block, at least one of the L0 direction motion information or the L1 direction motion information can be used. The bidirectional weight value index identifies a weight value to be applied to the L0 prediction block and a weight value to be applied to the L1 prediction block.
[0106] FIG. 6 is a flowchart of an inter prediction method according to an embodiment of the present invention.
[0107] Referring to FIG. 6, the inter prediction method includes: determining an inter prediction mode of a current block (step S601); obtaining motion information of the current block according to the determined inter prediction mode (step S602); and performing motion compensation prediction for the current block based on the obtained motion information (step S603).
[0108] Here, the inter prediction mode represents various techniques for determining motion information of a current block, and the inter prediction mode may include an inter prediction mode using translation motion information and an inter prediction mode using affine motion information. For example, the inter prediction mode using translation motion information may include a merge mode and a motion vector prediction mode, and the inter prediction mode using affine motion information may include an affine merge mode and an affine motion vector prediction mode. Depending on the inter prediction mode, the motion information of the current block may be determined based on information analyzed from a neighboring block or a bitstream adjacent to the current block.
[0109] The inter prediction method using affine motion information will be described in detail below.
[0110] FIG. 7 is a diagram illustrating the nonlinear movement of an object.
[0111] The movement of an object in an image may be nonlinear. For example, as shown in FIG. 7, nonlinear movement of an object may occur, such as zoom-in, zoom-out, rotation, or affine transformation. When nonlinear movement of an object occurs, the object movement cannot be effectively represented by a translational motion vector. Therefore, in areas where nonlinear movement of an object occurs, affine motion can be used instead of translational motion to improve coding efficiency.
[0112] FIG. 8 is a flowchart of an affine motion-based inter prediction method according to an embodiment of the present invention.
[0113] Whether to apply an affine motion-based inter prediction technique to a current block may be determined based on information analyzed from the bitstream. Specifically, whether to apply an affine motion-based inter prediction technique to a current block may be determined based on at least one of a flag indicating whether an affine merge mode is applied to the current block or a flag indicating whether an affine motion vector prediction mode is applied to the current block.
[0114] When an affine motion-based inter prediction technique is applied to a current block, an affine motion model of the current block can be determined (step S801). The affine motion model can be determined by at least one of a six-parameter affine motion model or a four-parameter affine motion model. The six-parameter affine motion model represents affine motion using six parameters, and the four-parameter affine motion model represents affine motion using four parameters.
[0115] Equation 1 describes affine motion using six parameters: Affine motion describes translational movement relative to a given region determined by an affine seed vector.
[0116]
number
[0117] When affine motion is represented using six parameters, it is possible to represent complex motion, but the number of bits required to encode each parameter increases, which may reduce coding efficiency. Therefore, affine motion can also be represented using four parameters. Equation 2 represents affine motion using four parameters.
[0118]
number
[0119] Information for determining an affine motion model of a current block may be coded and signaled via a bitstream. For example, the information may be a 1-bit flag "affine_type_flag." A value of 0 of the flag may indicate that a 4-parameter affine motion model is applied, and a value of 1 of the flag may indicate that a 6-parameter affine motion model is applied. The flag may be coded in units of slices, tiles, or blocks (e.g., coding blocks or coding tree units). When a flag is signaled at the slice level, the affine motion model determined by the slice level may be applied to all blocks belonging to the slice.
[0120] Alternatively, the affine motion model of the current block may be determined based on the affine inter prediction mode of the current block. For example, when an affine merge mode is applied, the affine motion model of the current block may be determined as a four-parameter motion model. On the other hand, when an affine motion vector prediction mode is applied, information for determining the affine motion model of the current block may be encoded and signaled via a bitstream. For example, when an affine motion vector prediction mode is applied to the current block, the affine motion model of the current block may be determined based on a 1-bit flag "affine_type_flag."
[0121] Next, an affine seed vector for the current block can be derived (step S802). If a four-parameter affine motion model is selected, motion vectors at two control points of the current block can be derived. On the other hand, if a six-parameter affine motion model is selected, motion vectors at three control points of the current block can be derived. The motion vectors at the control points can be referred to as affine seed vectors. The control points can include at least one of the top left corner, the top right corner, or the bottom left corner of the current block.
[0122] FIG. 9 is a diagram showing affine seed vectors for each affine motion model.
[0123] In a four-parameter affine motion model, affine seed vectors for two of the upper-left corner, the upper-right corner, or the lower-left corner can be derived. As an example, as shown in (a) of FIG. 9, when a four-parameter affine motion model is selected, an affine vector can be derived using an affine seed vector sv0 for the upper-left corner of the current block (e.g., the upper-left sample (x1, y1)) and an affine seed vector sv1 for the upper-right corner of the current block (e.g., the upper-right sample (x1, y1)). An affine seed vector for the lower-left corner can be used instead of the affine seed vector for the upper-left corner, or an affine seed vector for the lower-left corner can be used instead of the affine seed vector for the upper-right corner.
[0124] In a six-parameter affine motion model, affine seed vectors for the upper left corner, the upper right corner, and the lower left corner can be derived. As an example, as shown in (b) of FIG. 9, when a six-parameter affine motion model is selected, affine vectors can be derived using an affine seed vector sv0 for the upper left corner of the current block (e.g., the upper left sample (x1, y1)), an affine seed vector sv1 for the upper right corner of the current block (e.g., the upper right sample (x1, y1)), and an affine seed vector sv2 for the upper left corner of the current block (e.g., the upper left sample (x2, y2)).
[0125] In the embodiment described below, under the four-parameter affine motion model, the affine seed vectors of the upper-left control point and the upper-right control point will be referred to as the first affine seed vector and the second affine seed vector, respectively. In the embodiment described below that uses the first affine seed vector and the second affine seed vector, at least one of the first affine seed vector and the second affine seed vector can be replaced with the affine seed vector of the lower-left control point (third affine seed vector) or the affine seed vector of the lower-right control point (fourth affine seed vector).
[0126] Furthermore, under the six-parameter affine motion model, the affine seed vectors of the upper-left control point, the upper-right control point, and the lower-left control point are referred to as the first affine seed vector, the second affine seed vector, and the third affine seed vector, respectively. In an embodiment using the first affine seed vector, the second affine seed vector, and the third affine seed vector, which will be described later, at least one of the first affine seed vector, the second affine seed vector, and the third affine seed vector can be replaced with the affine seed vector of the lower-right control point (fourth affine seed vector).
[0127] Using the affine seed vector, an affine vector can be derived for each sub-block (step S803). Here, the affine vector represents a translational motion vector derived based on the affine seed vector. The affine vector of a sub-block can be referred to as an affine sub-block motion vector or a sub-block motion vector.
[0128] FIG. 10 is a diagram showing affine vectors of sub-blocks under a four-parameter motion model.
[0129] Based on the control point positions, the sub-block positions, and the affine seed vector, the affine vector of the sub-block can be derived. As an example, Equation 3 shows an example of deriving an affine sub-block vector.
[0130]
number
[0131] In the above equation 3, (x, y) represents the position of the sub-block. Here, the position of the sub-block represents the position of a reference sample included in the sub-block. The reference sample may be a sample located in the upper left corner of the sub-block, or a sample whose x-axis or y-axis coordinate is at the center. (x0, y0) represents the position of the first control point, (sv0x, sv0y) represents the first affine seed vector, (x1, y1) represents the position of the second control point, and (sv1x, sv1y) represents the second affine seed vector.
[0132] If the first and second control points correspond to the upper left and upper right corners of the current block, respectively, x1-x0 can be set to the same value as the width of the current block.
[0133] Then, using the affine vector of each sub-block, motion compensation prediction can be performed for each sub-block (step S804). As a result of the motion compensation prediction, a prediction block for each sub-block can be generated. The prediction block of the sub-block can be set as the prediction block of the current block.
[0134] The inter prediction method using translational motion information will be described in detail below.
[0135] The motion information of the current block may be derived from the motion information of another block of the current block. Here, the other block may be a block that has been coded / decoded using inter prediction before the current block. Setting the motion information of the current block to be the same as the motion information of another block may be defined as a merge mode. Also, setting the motion vector of another block as a predicted value of the motion vector of the current block may be defined as a motion vector prediction mode.
[0136] FIG. 11 is a flowchart of the process of deriving motion information for the current block under merge mode.
[0137] Merge candidates for the current block can be derived (step S1101). The merging candidates for the current block can be derived from blocks that have been coded / decoded using inter prediction prior to the current block.
[0138] FIG. 12 is a diagram showing candidate blocks for deriving merge candidates.
[0139] The candidate block may include at least one of a neighboring block including samples neighboring the current block or a non-neighboring block including samples not neighboring the current block. Hereinafter, samples determining the candidate block are defined as reference samples. Also, reference samples neighboring the current block are referred to as neighboring reference samples, and reference samples not neighboring the current block are referred to as non-neighboring reference samples.
[0140] The neighboring reference samples may include a column adjacent to the leftmost column of the current block or a row adjacent to the top row of the current block. For example, if the coordinates of the top left sample of the current block are (0,0), at least one of the block including the reference sample at (-1,H-1), the block including the reference sample at (W-1,-1), the block including the reference sample at (W,-1), the block including the reference sample at (-1,H), or the block including the reference sample at (-1,-1) can be used as a candidate block. Referring to the drawing, neighboring blocks with indexes 0 to 4 can be used as candidate blocks.
[0141] A non-adjacent reference sample refers to a sample in which at least one of the x-axis distance or y-axis distance from a reference sample adjacent to the current block has a predefined value. For example, at least one of a block including a reference sample whose x-axis distance from the left reference sample is a predefined value, a block including a non-adjacent sample whose y-axis distance from the top reference sample is a predefined value, or a block including a non-adjacent sample whose x-axis distance and y-axis distance from the top-left reference sample are predefined values can be used as a candidate block. The predefined value may be a natural number such as 4, 8, 12, or 16. Referring to the drawing, at least one of the blocks with indexes 5 through 26 can be used as a candidate block.
[0142] A sample that is not located on the same vertical, horizontal, or diagonal line as an adjacent reference sample can be set as a non-adjacent reference sample.
[0143] FIG. 13 is a diagram of the location of the reference sample.
[0144] 13, the x-coordinate of the upper non-adjacent reference sample may be set to be different from the x-coordinate of the upper adjacent reference sample. As an example, if the position of the upper adjacent reference sample is (W-1, -1), the position of the upper non-adjacent reference sample that is N away from the upper adjacent reference sample on the y-axis may be set to ((W / 2)-1, -1-N), and the position of the upper non-adjacent reference sample that is 2N away from the upper adjacent reference sample on the y-axis may be set to (0, -1-2N). In other words, the positions of the non-adjacent reference samples may be determined based on the distance between the adjacent reference sample and the adjacent reference sample.
[0145] Hereinafter, among the candidate blocks, a candidate block containing adjacent reference samples will be referred to as an adjacent block, and a block containing non-adjacent reference samples will be referred to as a non-adjacent block.
[0146] If the distance between the current block and the candidate block is greater than or equal to a threshold, the candidate block can be set as unavailable as a merge candidate. The threshold can be determined based on the size of the coding tree unit. For example, the threshold can be set to the height of the coding tree unit (ctu_height) or a value obtained by adding or subtracting an offset from the height of the coding tree unit (e.g., ctu_height±N). The offset N is a value predefined by the encoder and decoder and can be set to 4, 8, 16, 32, or ctu_height.
[0147] If the difference between the y-axis coordinate of the current block and the y-axis coordinate of the sample included in the candidate block is greater than a threshold, it can be determined that the candidate block is not available as a merging candidate.
[0148] Alternatively, a candidate block that does not belong to the same coding tree unit as the current block may be set as unavailable as a merging candidate. For example, if a reference sample exceeds the upper boundary of the coding tree unit to which the current block belongs, a candidate block containing the reference sample may be set as unavailable as a merging candidate.
[0149] If the upper boundary of the current block is adjacent to the upper boundary of the coding tree unit, many candidate blocks may be determined to be unavailable as merge candidates, which may reduce the encoding / decoding efficiency of the current block. To solve the above problem, the candidate blocks may be set so that the number of candidate blocks located to the left of the current block is greater than the number of candidate blocks located above the current block.
[0150] FIG. 14 is a diagram showing candidate blocks for deriving merge candidates.
[0151] As shown in the example of Figure 14, the top blocks belonging to the row of N blocks above the current block and the left blocks belonging to the row of M blocks to the left of the current block can be set as candidate blocks. In this case, M can be set to be larger than N, so that the number of left candidate blocks can be set to be larger than the number of top candidate blocks.
[0152] For example, the difference between the y-axis coordinate of a reference sample in the current block and the y-axis coordinate of an upper block that can be used as a candidate block can be set to not exceed N times the height of the current block, and the difference between the x-axis coordinate of a reference sample in the current block and the x-axis coordinate of a left block that can be used as a candidate block can be set to not exceed M times the width of the current block.
[0153] As an example, in the example shown in Figure 14, it is shown that the blocks belonging to the two block columns above the current block and the blocks belonging to the five block columns to the left of the current block are set as candidate blocks.
[0154] As another example, if a candidate block does not belong to the same coding tree unit as the current block, a merging candidate can be derived instead using a block that belongs to the same coding tree unit as the current block or a block containing a reference sample adjacent to the boundary of the coding tree unit.
[0155] FIG. 14 is a diagram showing candidate blocks for deriving merge candidates.
[0156] If the reference sample is included in a different coding tree unit from the current block and the reference sample is not adjacent to the boundary of the coding tree unit, a reference sample adjacent to the boundary of the coding tree unit can be used instead of the reference sample to determine the candidate block.
[0157] 15(a) and 15(b), when the top boundary of the current block and the top boundary of the coding tree unit are adjacent, the reference sample at the top of the current block belongs to a coding tree unit different from the current block. Of the reference samples belonging to a coding tree unit different from the current block, the reference samples that are not adjacent to the top boundary of the coding tree unit can be replaced with samples that are adjacent to the top boundary of the coding tree unit.
[0158] As an example, as shown in (a) of Figure 15, the reference sample at position 6 can be replaced with a sample at position 6' located at the upper boundary of the coding tree unit, and as shown in (b) of Figure 15, the reference sample at position 15 can be replaced with a sample at position 15' located at the upper boundary of the coding tree unit. In this case, the y coordinate of the replacement sample can be biased to an adjacent position in the coding tree unit, and the x coordinate of the replacement sample can be set to the same as the reference sample. As an example, the sample at position 6' can have the same x coordinate as the sample at position 6, and the sample at position 15' can have the same x coordinate as the sample at position 15.
[0159] Alternatively, the x-coordinate of the substitute sample may be set to the x-coordinate of the reference sample plus or minus an offset. As an example, if the x-coordinates of an adjacent reference sample and a non-adjacent reference sample located at the top of the current block are the same, the x-coordinate of the substitute sample may be set to the x-coordinate of the reference sample plus or minus an offset. This is to prevent the substitute sample that replaces the non-adjacent reference sample from being at the same position as other non-adjacent or adjacent reference samples.
[0160] FIG. 16 is a diagram showing an example in which the position of the reference sample is changed.
[0161] When replacing a reference sample that is included in a coding tree unit different from the current block and is not adjacent to the boundary of the coding tree unit with a sample located on the boundary of the coding tree unit, the x coordinate of the replacement sample can be set to a value obtained by adding or subtracting an offset from the x coordinate of the reference sample.
[0162] 16, the reference samples at positions 6 and 15 may be replaced with samples at positions 6′ and 15′, respectively, that have the same y-coordinate as the row adjacent to the top boundary of the coding tree unit. In this case, the x-coordinate of the sample at position 6′ may be set to the x-coordinate of the reference sample at position 6 minus W / 2, and the x-coordinate of the sample at position 15′ may be set to the x-coordinate of the reference sample at position 15 minus W−1.
[0163] Unlike the examples shown in Figures 15 and 16, the y coordinate of the replacement sample can also be set to the y coordinate of the row located above the top row of the current block or the y coordinate of the top boundary of the coding tree unit.
[0164] Although not shown, a sample to replace the reference sample can also be determined based on the left boundary of the coding tree unit. As an example, if the reference sample is not included in the same coding tree unit as the current block and is not adjacent to the left boundary of the coding tree unit, the reference sample can be replaced with a sample adjacent to the left boundary of the coding tree unit. In this case, the replacement sample can have the same y-coordinate as the reference sample, or a y-coordinate obtained by adding or subtracting an offset from the y-coordinate of the reference sample.
[0165] Then, the block containing the replacement sample can be set as a candidate block, and a merging candidate for the current block can be derived based on the candidate block.
[0166] Merging candidates can also be derived from temporally adjacent blocks contained in a different image from the current block. For example, merging candidates can be derived from co-located blocks contained in a co-located image.
[0167] The motion information of the merging candidate may be set to the same as the motion information of the candidate block. For example, at least one of the motion vector, reference image index, and weight index of the prediction direction or both directions of the candidate block may be set to the motion information of the merging candidate.
[0168] A merge candidate list including merge candidates may be generated (step S1102). The merge candidates may be classified into adjacent merge candidates derived from adjacent blocks adjacent to the current block and non-adjacent merge candidates derived from non-adjacent blocks.
[0169] Indices for multiple merge candidates in the merge candidate list can be assigned according to a predetermined order. For example, indices assigned to adjacent merge candidates have lower values than indices assigned to non-adjacent merge candidates. Alternatively, indices can be assigned to each merge candidate based on the index of each block shown in FIG. 12 or FIG. 14.
[0170] If the merge candidate set includes multiple merge candidates, at least one of the multiple merge candidates may be selected (step S1103). In this case, information indicating whether motion information of the current block is derived from a neighboring merge candidate may be signaled via the bitstream. The information may be a 1-bit flag. As an example, a syntax element isAdjancentMergeFlag indicating whether motion information of the current block is derived from a neighboring merge candidate may be signaled via the bitstream. If the value of the syntax element isAdjancentMergeFlag is 1, motion information of the current block may be derived based on a neighboring merge candidate. On the other hand, if the value of the syntax element isAdjancentMergeFlag is 0, motion information of the current block may be derived based on a non-neighboring merge candidate.
[0171] Table 1 shows a syntax table including the syntax element isAdjancentMergeFlag.
[0172] [Table 1] JPEG2025166201000006.jpg121161
[0173] Information for identifying one of the multiple merge candidates may be signaled via the bitstream. For example, information indicating the index of one of the merge candidates included in the merge candidate list may be signaled via the bitstream.
[0174] A syntax element merge_idx for identifying any of the adjacent merge candidates can be signaled when isAdjacentMergeflag is 1. The maximum value of the syntax element merge_idx can be set to the number of adjacent merge candidates minus 1.
[0175] If isAdjacentMergeflag is 0, the syntax element NA_merge_idx can be signaled to identify one of the non-adjacent merge candidates. The syntax element NA_merge_idx represents the index of the non-adjacent merge candidate minus the number of adjacent merge candidates. The decoder can select a non-adjacent merge candidate by adding the number of adjacent merge candidates to the index identified by NA_merge_idx.
[0176] If the number of merge candidates included in the merge candidate list is less than a threshold, merge candidates included in the inter-region motion information table may be added to the merge candidate list. Here, the threshold may be the maximum number of merge candidates that the merge candidate list can include or the maximum number of merge candidates minus an offset. The offset may be a natural number such as 1 or 2. The inter-region motion information table may include merge candidates derived based on blocks encoded / decoded before the current block.
[0177] The inter-region motion information table includes merge candidates derived from blocks encoded / decoded based on inter prediction in the current image. For example, the motion information of the merge candidates included in the inter-region motion information table may be set to the same as the motion information of blocks encoded / decoded based on inter prediction. Here, the motion information may include at least one of a motion vector, a reference image index, and a weight index in the prediction direction or both directions.
[0178] For convenience of explanation, the merge candidates included in the inter-region motion information table will be referred to as inter-region merge candidates.
[0179] The maximum number of merge candidates that the inter region motion information table can contain may be predefined by the encoder and decoder. As an example, the maximum number of merge candidates that the inter region motion information table can contain may be 1, 2, 3, 4, 5, 6, 7, 8, or more (e.g., 16).
[0180] Alternatively, information representing the maximum number of merge candidates in the inter-region motion information table can be signaled via the bitstream, which can be signaled at the sequence, picture, or slice level.
[0181] Or, the maximum number of merge candidates of the inter-region motion information table can be determined according to the size of the image, the size of the slice, or the size of the coding tree unit.
[0182] The inter-region motion information table may be initialized in units of an image, slice, tile, brick, coding tree unit, or coding tree unit line (row or column). For example, when a slice is initialized, the inter-region motion information table is also initialized, and the inter-region motion information table may not include merge candidates.
[0183] Alternatively, information indicating whether to initialize the inter-region motion information table can be signaled via the bitstream. The information can be signaled at the slice, tile, brick, or block level. Before the information indicates that the inter-region motion information table should be initialized, a pre-configured inter-region motion information table can be used.
[0184] Alternatively, information about the initial inter region merge candidate may be signaled via a picture parameter set or a slice header. Even when a slice is initialized, the inter region motion information table may include the initial inter region merge candidate. This allows the inter region merge candidate to be used even for the first coded / decoded block in the slice.
[0185] Blocks are coded / decoded in the coding / decoding order, and blocks coded / decoded based on inter prediction can be set as inter region merge candidates in sequence in the coding / decoding order.
[0186] FIG. 17 is a flowchart for explaining the update state of the inter-region motion information table.
[0187] When inter prediction is performed on the current block (step S1701), inter region merging candidates can be derived based on the current block (step S1702). The motion information of the inter region merging candidates can be set to the same as the motion information of the current block.
[0188] If the inter region motion information table is empty (step S1703), the inter region merge candidates derived based on the current block can be added to the inter region motion information table (step S1704).
[0189] If the inter region motion information table already includes inter region merge candidates (step S1703), redundancy detection can be performed on the motion information of the current block (or on inter region merge candidates derived based on the motion information) (step S1705). Redundancy detection is used to determine whether the motion information of the inter region merge candidates stored in the inter region motion information table is the same as the motion information of the current block. Redundancy detection can be performed on all inter region merge candidates stored in the inter region motion information table. Alternatively, redundancy detection can be performed on inter region merge candidates stored in the inter region motion information table whose indexes are greater than or equal to a threshold value.
[0190] If the inter-prediction merge candidate having the same motion information as the current block is not included, an inter-region merge candidate derived based on the current block may be added to the inter-region motion information table (step S1708). Whether the inter-prediction merge candidates are the same may be determined based on whether their motion information (e.g., motion vectors and / or reference image indexes) is the same.
[0191] In this case, if the maximum number of inter region merge candidates is stored in the inter region motion information table (step S1706), the oldest inter region merge candidate can be deleted (step S1707), and an inter region merge candidate derived based on the current block can be added to the inter region motion information table (step S1708).
[0192] Inter region merge candidates may be identified by their respective indexes. When an inter region merge candidate derived from the current block is added to the inter region motion information table, the inter region merge candidate may be assigned the lowest index (e.g., 0), and the indexes of the stored inter region merge candidates may be incremented by 1. In this case, if the maximum number of inter prediction merge candidates is already stored in the inter region motion information table, the inter region merge candidate with the highest index is removed.
[0193] Alternatively, when an inter region merge candidate derived from the current block is added to the inter region motion information table, the inter region merge candidate may be assigned the largest index. For example, if the number of inter prediction merge candidates stored in the inter region motion information table is less than the maximum value, the inter region merge candidate may be assigned an index with a value equal to the number of stored inter prediction merge candidates. Alternatively, if the number of inter prediction merge candidates stored in the inter region motion information table is the same as the maximum value, the inter region merge candidate may be assigned an index obtained by subtracting 1 from the maximum value. In addition, the inter region merge candidate with the smallest index is removed, and the indexes of the remaining stored inter region merge candidates are decremented by 1.
[0194] FIG. 18 is a diagram illustrating an example of updating an inter-region merge candidate list.
[0195] Assume that an inter region merge candidate derived from the current block is added to an inter region merge candidate table, the inter region merge candidate is assigned the highest index, and the maximum number of inter region merge candidates is already stored in the inter region merge candidate table.
[0196] When adding an inter region merge candidate HmvpCand[n+1] derived from the current block to the inter region merge candidate table HmvpCandList, the inter region merge candidate HmvpCand[0] with the smallest index among the stored inter region merge candidates may be deleted, and the indexes of the remaining inter region merge candidates may be decremented by 1. In addition, the index of the inter region merge candidate HmvpCand[n+1] derived from the current block may be set to the maximum value (n in the example shown in FIG. 18).
[0197] If an inter region merge candidate that is the same as the inter region merge candidate derived based on the current block is stored (step S1705), the inter region merge candidate derived based on the current block may not be added to the inter region motion information table (step S1709).
[0198] Alternatively, an inter region merge candidate derived based on the current block may be added to the inter region motion information table, while a stored inter region merge candidate that is the same as the inter region merge candidate may be removed, which has the same effect as updating the index of the stored inter region merge candidate.
[0199] FIG. 19 is a diagram showing an example in which the indexes of stored inter region merge candidates are updated.
[0200] If the index of a stored inter-prediction merge candidate that is the same as the inter-region merge candidate mvCand derived based on the current block is hIdx, the stored inter-prediction merge candidate may be deleted, and the indexes of inter-prediction merge candidates whose indexes are greater than hIdx may be decremented by 1. For example, in the example shown in Figure 19, HmvpCand[2] that is the same as mvCand is deleted from the inter-region motion information table HvmpCandList, and the indices from HmvpCand[3] to HmvpCand[n] are decremented by 1.
[0201] Then, the inter region merge candidate mvCand derived based on the current block can be added to the end of the inter region motion information table.
[0202] Alternatively, the index assigned to a stored inter region merge candidate that is the same as the inter region merge candidate derived based on the current block can be updated, for example, the index of the stored inter region merge candidate can be changed to a minimum or maximum value.
[0203] The motion information of blocks included in a predetermined region may be set not to be added to the inter region motion information table. For example, inter region merge candidates derived based on the motion information of blocks included in the merge processing region may not be added to the inter region motion information table. Because the encoding / decoding order of blocks included in the merge processing region is not defined, it is inappropriate to use the motion information of any of these blocks for inter prediction of other blocks. Therefore, inter region merge candidates derived based on blocks included in the merge processing region may not be added to the inter region motion information table.
[0204] When motion compensation prediction is performed on a sub-block basis, inter region merging candidates may be derived based on motion information of a representative sub-block among a plurality of sub-blocks included in the current block. For example, when sub-block merging candidates are used for the current block, inter region merging candidates may be derived based on motion information of a representative sub-block among the sub-blocks.
[0205] The motion vector of a sub-block may be derived in the following order: First, one of the merge candidates included in the merge candidate list of the current block is selected, and an initial shift vector (shVector) may be derived based on the motion vector of the selected merge candidate. Then, the initial shift vector may be added to the position (xSb, ySb) of a reference sample (e.g., the upper left sample or the sample at the center) of each sub-block of the coding block to derive a shifted sub-block whose reference sample position is (xColSb, yColSb). The following Equation 4 represents an equation for deriving the shifted sub-block.
[0206]
number
[0207] Then, the motion vector of the block located at the same position corresponding to the center position of the sub-block containing (xColSb, yColSb) can be set as the motion vector of the sub-block containing (xSb, ySb).
[0208] The representative sub-block may refer to the sub-block including the top left sample or the center sample of the current block.
[0209] FIG. 20 is a diagram showing the position of the representative sub-block.
[0210] 20(a) shows an example in which a sub-block located at the top left of a current block is set as a representative sub-block, and FIG. 20(b) shows an example in which a sub-block located at the center of the current block is set as a representative sub-block. When motion compensation prediction is performed in units of sub-blocks, an inter region merge candidate for the current block can be derived based on the motion vector of the sub-block including the top left sample of the current block or the sub-block including the center sample of the current block.
[0211] Whether to use the current block as an inter region merging candidate may also be determined based on the inter prediction mode of the current block. For example, a block encoded / decoded based on an affine motion model may be set as unavailable as an inter region merging candidate. Thus, even if the current block is encoded / decoded using inter prediction, if the inter prediction mode of the current block is an affine prediction mode, the inter prediction motion information table may not be updated based on the current block.
[0212] Alternatively, an inter region merging candidate may be derived based on at least one subblock vector of subblocks included in a block encoded / decoded based on an affine motion model. For example, an inter region merging candidate may be derived using a subblock located at the upper left, center, or upper right of the current block. Alternatively, an average value of subblock vectors of multiple subblocks may be set as the motion vector of the inter region merging candidate.
[0213] Alternatively, the inter region merging candidate may be derived based on the average value of the affine seed vectors of the blocks encoded / decoded based on the affine motion model. For example, the average value of at least one of the first affine seed vector, the second affine seed vector, and the third affine seed vector of the current block may be set as the motion vector of the inter region merging candidate.
[0214] Alternatively, an inter-region motion information table may be configured for each inter-prediction mode. For example, at least one of an inter-region motion information table for a block coded / decoded using intra block copying, an inter-region motion information table for a block coded / decoded based on a translational motion model, or an inter-region motion information table for a block coded / decoded based on an affine motion model may be defined. One of the inter-region motion information tables may be selected according to the inter-prediction mode of the current block.
[0215] FIG. 21 illustrates an example in which an inter-region motion information table is generated for each inter-prediction mode.
[0216] If a block is encoded / decoded based on a non-affine motion model, the inter region merge candidate mvCand derived based on the block can be added to the inter region non-affine motion information table HmvpCandList. On the other hand, if a block is encoded / decoded based on an affine motion model, the inter region merge candidate mvAfCand derived based on the block can be added to the inter region affine motion information table HmvpAfCandList.
[0217] An inter-region merge candidate derived from a block encoded / decoded based on an affine motion model can store the affine seed vector of the block, allowing the inter-region merge candidate to be used as a merge candidate for deriving the affine seed vector of the current block.
[0218] In addition to the inter-region motion information table described above, another inter-region motion information table can also be defined. In addition to the inter-region motion information table described above (hereinafter referred to as the first inter-region motion information table), a long-term motion information table (hereinafter referred to as the second inter-region motion information table) can be defined. Here, the long-term motion information table includes long-term merge candidates.
[0219] If the first inter-region motion information table and the second inter-region motion information table are empty, inter-region merge candidates can be added to the second inter-region motion information table first, and only after the number of inter-region merge candidates available in the second inter-region motion information table reaches the maximum number can inter-region merge candidates be added to the first inter-region motion information table.
[0220] Alternatively, one inter prediction merge candidate can be added to both the second inter region motion information table and the first inter region motion information table.
[0221] In this case, the second inter region motion information table whose construction has been completed may not be updated any more, or may be updated if the decoded region is equal to or greater than a predetermined ratio of the slice, or may be updated every N coding tree unit lines.
[0222] On the other hand, the first inter region motion information table may be updated each time a block encoded / decoded by inter prediction occurs, but the inter region merge candidates added to the second inter region motion information table may be set not to be used to update the first inter region motion information table.
[0223] Information for selecting either the first inter-region motion information table or the second inter-region motion information table may be signaled via the bitstream. If the number of merge candidates included in the merge candidate list is less than a threshold, merge candidates included in the inter-region motion information table indicated by the information may be added to the merge candidate list.
[0224] Alternatively, the inter-region motion information table may be selected based on the size, shape, inter-prediction mode, whether bidirectional prediction is performed, whether motion vector refinement is performed, or whether triangulation is performed.
[0225] Alternatively, if the number of merge candidates included in the merge candidate list is less than the maximum number of merges even after adding an inter region merge candidate included in the first inter region motion information table, an inter region merge candidate included in the second inter region motion information table can be added to the merge candidate list.
[0226] FIG. 22 is a diagram illustrating an example in which inter-region merge candidates included in the long-term motion information list are added to the merge candidate list.
[0227] If the number of merge candidates included in the merge candidate list is less than the maximum number, the inter region merge candidates included in the first inter region motion information table HmvpCandList can be added to the merge candidate list. If the number of merge candidates included in the merge candidate list is less than the maximum number even after the inter region merge candidates included in the first inter region motion information table have been added to the merge candidate list, the inter region merge candidates included in the long term motion information table HmvpLTCandList can be added to the merge candidate list.
[0228] Table 2 shows the process of adding inter-region merge candidates included in the long-term motion information table to the merge candidate list.
[0229] [Table 2]
[0230] Inter region merge candidates may be configured to include additional information in addition to motion information. For example, block size, shape, or block partition information of the inter region merge candidates may be additionally stored. When constructing a merge candidate list for the current block, only inter prediction merge candidates having the same or similar size, shape, or partition information as the current block may be used, or inter prediction merge candidates having the same or similar size, shape, or partition information as the current block may be first added to the merge candidate list.
[0231] Alternatively, an inter-region motion information table may be generated for each size, shape, or partition information of a block. A merge candidate list for the current block may be generated using an inter-region motion information table that matches the shape, size, or partition information of the current block among the multiple inter-region motion information tables.
[0232] If the number of merge candidates included in the merge candidate list of the current block is less than a threshold, the inter-region merge candidates included in the inter-region motion information table may be added to the merge candidate list. The addition process may be performed in ascending or descending order based on index. For example, the inter-region merge candidate with the largest index may be added to the merge candidate list first.
[0233] When an inter-region merge candidate included in the inter-region motion information table is to be added to the merge candidate list, redundancy detection may be performed between the inter-region merge candidate and the merge candidates stored in the merge candidate list.
[0234] As an example, Table 3 shows the process by which inter-region merge candidates are added to the merge candidate list.
[0235] [Table 3]
[0236] Redundancy detection may be performed only for a portion of the inter region merge candidates included in the inter region motion information table. For example, redundancy detection may be performed only for inter region merge candidates whose indexes are greater than or equal to a threshold value, or only for the N merge candidates with the largest index or the N merge candidates with the smallest index.
[0237] Alternatively, redundancy detection may be performed only on a portion of the merge candidates stored in the merge candidate list. For example, redundancy detection may be performed only on merge candidates whose indexes are greater than or less than a threshold or merge candidates derived from blocks at a specific location. Here, the specific location may include at least one of the left neighboring block, the top neighboring block, the top-right neighboring block, or the bottom-left neighboring block of the current block.
[0238] FIG. 23 is a diagram illustrating an example in which redundancy detection is performed on only some of the merge candidates.
[0239] When an inter-region merge candidate HmvpCand[j] is to be added to the merge candidate list, redundancy detection can be performed for the inter-region merge candidate with the two merge candidates with the highest indices, mergeCandList[NumMerge-2] and mergeCandList[NumMerge-1], where NumMerge represents the number of available spatial and temporal merge candidates.
[0240] Unlike the illustrated example, when adding an inter-region merge candidate HmvpCand[j] to the merge candidate list, redundancy detection can be performed on the inter-region merge candidate with up to two merge candidates with the smallest indexes. For example, it can be checked whether mergeCandList[0] and mergeCandList[1] are the same as HmvpCand[j]. Alternatively, redundancy detection can be performed only on merge candidates derived from a specific position. For example, redundancy detection can be performed on at least one of merge candidates derived from neighboring blocks located to the left of the current block or merge candidates derived from neighboring blocks located above the current block. If there is no merge candidate derived from a specific position in the merge candidate list, the inter-region merge candidate can be added to the merge candidate list without redundancy detection.
[0241] If a merge candidate identical to the first inter region merge candidate is found, redundancy detection with the same merge candidate as the first inter region merge candidate can be omitted when redundancy detection is performed for the second inter region merge candidate.
[0242] FIG. 24 is a diagram illustrating an example in which redundant detection with a specific merge candidate is omitted.
[0243] When an inter-region merge candidate HmvpCand[i] with index i is to be added to the merge candidate list, redundancy detection is performed between the inter-region merge candidate and the merge candidates stored in the merge candidate list. In this case, if a merge candidate mergeCandList[j] identical to the inter-region merge candidate HmvpCand[i] is found, redundancy detection can be performed between the inter-region merge candidate HmvpCand[i-1] with index i-1 and the merge candidate without adding the inter-region merge candidate HmvpCand[i-1] to the merge candidate list. In this case, redundancy detection between the inter-region merge candidate HmvpCand[i-1] and the merge candidate mergeCandList[j] can be omitted.
[0244] 24, it is determined that HmvpCand[i] and mergeCandList[2] are the same. Therefore, HmvpCand[i] is not added to the merge candidate list, and redundancy detection can be performed on HmvpCand[i-1]. In this case, redundancy detection between HmvpCand[i-1] and mergeCandList[2] can be omitted.
[0245] If the number of merge candidates included in the merge candidate list of the current block is less than a threshold, the list may further include at least one of pairwise merge candidates or zero merge candidates in addition to inter-region merge candidates. A pairwise merge candidate is a merge candidate that uses the average value of the motion vectors of two or more merge candidates as a motion vector, and a zero merge candidate is a merge candidate whose motion vector is 0.
[0246] The current block's merge candidate list is populated with merge candidates in the following order:
[0247] Spatial merge candidates - Temporal merge candidates - Inter-region merge candidates - (Inter-region affine merge candidates) - Pairwise merge candidates - Zero merge candidates The spatial merge candidates refer to merge candidates derived from at least one of adjacent or non-adjacent blocks, and the temporal merge candidates refer to merge candidates derived from previous reference images. The inter-region affine merge candidates represent inter-region merge candidates derived from blocks coded / decoded with an affine motion model.
[0248] The inter-region motion information table can also be used in the motion vector prediction mode. For example, if the number of motion vector prediction candidates included in the motion vector prediction candidate list of the current block is less than a threshold, an inter-region merge candidate included in the inter-region motion information table can be set as a motion vector prediction candidate for the current block. Specifically, the motion vector of the inter-region merge candidate can be set as a motion vector prediction candidate.
[0249] When one of the motion vector prediction candidates included in the motion vector prediction candidate list for the current block is selected, the selected candidate can be set as the motion vector prediction value for the current block. Then, after decoding the motion vector residual value for the current block, the motion vector prediction value and the motion vector residual value can be added to obtain the motion vector for the current block.
[0250] The motion vector prediction candidate list for the current block can be constructed according to the following order:
[0251] Spatial motion vector prediction candidate - Temporal motion vector prediction candidate - Inter-decoded region merge candidate - (Inter-decoded region affine merge candidate) - Zero motion vector prediction candidate The spatial motion vector prediction candidate refers to a motion vector prediction candidate derived from at least one of adjacent blocks or non-adjacent blocks, and the temporal motion vector prediction candidate refers to a motion vector prediction candidate derived from a previous reference image. The inter-region affine merge candidate refers to an inter-region motion vector prediction candidate derived from a block coded / decoded using an affine motion model. The zero motion vector prediction candidate refers to a candidate whose motion vector value is 0.
[0252] A merge processing region larger than the coding block can be defined. The coding blocks included in the merge processing region can be processed in parallel rather than sequentially encoded / decoded. Here, not being sequentially encoded / decoded means that the encoding / decoding order is not defined. This allows the encoding / decoding processes of the blocks included in the merge processing region to be processed independently. Alternatively, the blocks included in the merge processing region can share merge candidates. Here, the merge candidates can be derived based on the merge processing region.
[0253] According to the above characteristics, the merge processing region can also be referred to as a parallel processing region, a shared merge region (SMR), or a merge estimation region (MER).
[0254] Merge candidates for the current block may be derived based on the coded blocks, but if the current block is included in a merge processing area that is larger than the current block, candidate blocks included in the same merge processing area as the current block may be set as unavailable as merge candidates.
[0255] FIG. 25 is a diagram showing an example in which a candidate block included in the same merge processing area as the current block is set as unavailable as a merge candidate.
[0256] In the example shown in (a) of Figure 25, when encoding / decoding CU5, blocks including reference samples adjacent to CU5 can be set as candidate blocks. In this case, candidate blocks X3 and X4 included in the same merging processing area as CU5 can be set as unavailable as merge candidates for CU5. On the other hand, candidate blocks X0, X1, and X2 not included in the same merging processing area as CU5 can be set as available as merge candidates.
[0257] In the example shown in (b) of Figure 25, when encoding / decoding CU8, blocks including reference samples adjacent to CU8 can be set as candidate blocks. In this case, candidate blocks X6, X7, and X8 included in the same merge processing region as CU8 can be set as unavailable as merge candidates. On the other hand, candidate blocks X5 and X9 not included in the same merge region as CU8 can be set as available as merge candidates.
[0258] The merge processing region may be square or non-square. Information for determining the merge processing region may be signaled via the bitstream. The information may include at least one of information representing the shape of the merge processing region or information representing the size of the merge processing region. If the merge processing region is non-square, at least one of information representing the size of the merge processing region, information representing the width and / or height of the merge processing region, or information representing the ratio between the width and height of the merge processing region may be signaled via the bitstream.
[0259] The size of the merge processing region can be determined based on at least one of information signaled by a bitstream, an image resolution, a slice size, or a tile size.
[0260] When motion compensation prediction is performed on a block included in the merge processing region, an inter-region merge candidate derived based on the motion information of the block on which motion compensation prediction is performed can be added to the inter-region motion information table.
[0261] However, when an inter region merge candidate derived from a block included in the merge processing region is added to the inter region motion information table, the inter region merge candidate derived from the block may be used when encoding / decoding another block in the merge processing region that is actually slower to encode / decode than the block. That is, when encoding / decoding a block included in the merge processing region, motion prediction compensation may be performed using motion information of another block included in the merge processing region, even though inter-block dependency should be eliminated. To solve this problem, even if encoding / decoding of a block included in the merge processing region is completed, the motion information of the block whose encoding / decoding is completed may not be added to the inter region motion information table.
[0262] Alternatively, when motion compensation prediction is performed on blocks included in a merge processing region, inter-region merge candidates derived from the blocks may be added to the inter-region motion information table in a predefined order. Here, the predefined order may be determined according to a scan order of coding blocks in a merge processing region or a coding tree unit. The scan order may be at least one of raster scan, horizontal scan, vertical scan, and zigzag scan. Alternatively, the predefined order may be determined based on the motion information of each block or the number of blocks having the same motion information.
[0263] Alternatively, inter region merge candidates containing unidirectional motion information can be added to the inter region merge list before inter region merge candidates containing bidirectional motion information, and conversely, inter region merge candidates containing bidirectional motion information can be added to the inter region merge candidate list before inter region merge candidates containing unidirectional motion information.
[0264] Alternatively, the inter region merge candidates can be added to the inter region motion information table in order of most frequently used or least frequently used within the merge processing region or coding tree unit.
[0265] When a current block is added to a merge processing region and the number of merge candidates included in the merge candidate list of the current block is less than the maximum number, inter region merge candidates included in the inter region motion information table may be added to the merge candidate list. In this case, inter region merge candidates derived from blocks included in the same merge processing region as the current block may be set not to be added to the merge candidate list of the current block.
[0266] Alternatively, if the current block is included in the merge processing region, the inter region merge candidates included in the inter region motion information table may be set not to be used. That is, even if the number of merge candidates included in the merge candidate list of the current block is less than the maximum number, the inter region merge candidates included in the inter region motion information table may not be added to the merge candidate list.
[0267] An inter-region motion information table for a merge processing region or a coding tree unit can be configured. The inter-region motion information table serves to temporarily store motion information of blocks included in the merge processing region. To distinguish between a general inter-region motion information table and an inter-region motion information table for a merge processing region or a coding tree unit, the inter-region motion information table for a merge processing region or a coding tree unit will be referred to as a temporary motion information table. Furthermore, the inter-region merge candidates stored in the temporary motion information table will be referred to as temporary merge candidates.
[0268] FIG. 26 is a diagram showing a temporary motion information table.
[0269] A temporary motion information table may be configured for a coding tree unit or a merge processing region. If motion compensation prediction is performed on a current block included in a coding tree unit or a merge processing region, the motion information of the block may not be added to the inter prediction motion information table HmvpCandList. Instead, temporary merge candidates derived from the block may be added to the temporary motion information table HmvpMERCandList. That is, temporary merge candidates added to the temporary motion information table may not be added to the inter region motion information table. Thus, the inter region motion information table may not include inter region merge candidates derived based on the motion information of blocks included in the coding tree unit or merge processing region including the current block.
[0270] The maximum number of merge candidates that the temporary motion information table can include can be set in the same way as the inter-region motion information table, or can be determined according to the size of the coding tree unit or the merge processing region.
[0271] A current block included in a coding tree unit or a merge processing region may be set not to use a temporary motion information table for the corresponding coding tree unit or the corresponding merge processing region. That is, if the number of merge candidates included in the merge candidate list of the current block is less than a threshold, inter-region merge candidates included in the inter-region motion information table may be added to the merge candidate list, and temporary merge candidates included in the temporary motion information table may not be added to the merge candidate list. This may prevent motion information of other blocks included in the same coding tree unit or the same merge processing region as the current block from being used for motion compensation prediction of the current block.
[0272] When the coding / decoding of all blocks included in the coding tree unit or merge processing region is completed, the inter-region motion information table and the temporary motion information table can be merged.
[0273] FIG. 27 is a diagram showing an example of merging an inter-region motion information table and a temporary motion information table.
[0274] Once the encoding / decoding of all blocks included in the coding tree unit or merge processing region is completed, the inter-region motion information table can be updated with the temporary merge candidates included in the temporary motion information table, as shown in the example of FIG. 27.
[0275] In this case, the temporary merge candidates included in the temporary motion information table can be added to the inter-region motion information table according to the order in which they were inserted into the temporary motion information table (i.e., ascending or descending order of index values).
[0276] As another example, the temporal merge candidates included in the temporal motion information table can be added to the inter-region motion information table according to a predefined order.
[0277] Here, the predefined order may be determined according to a scan order of coding blocks in a merge processing region or a coding tree unit. The scan order may be at least one of a raster scan, a horizontal scan, a vertical scan, or a zigzag scan. Alternatively, the predefined order may be determined based on motion information of each block or the number of blocks having the same motion information.
[0278] Alternatively, temporary merge candidates containing unidirectional motion information can be added to the inter-region merge list before temporary merge candidates containing bidirectional motion information, and conversely, temporary merge candidates containing bidirectional motion information can be added to the inter-region merge candidate list before temporary merge candidates containing unidirectional motion information.
[0279] Alternatively, temporary merge candidates can be added to the inter-region motion information table according to the order of most frequently used or least frequently used within the merge processing region or coding tree unit.
[0280] When adding a temporary merge candidate included in the temporary motion information table to the inter region motion information table, redundancy detection may be performed on the temporary merge candidate. For example, if an inter region merge candidate identical to a temporary merge candidate included in the temporary motion information table is stored in the inter region motion information table, the temporary merge candidate may not be added to the inter region motion information table. In this case, redundancy detection may be performed on some of the inter region merge candidates included in the inter region motion information table. For example, redundancy detection may be performed on inter prediction merge candidates whose indexes are greater than or equal to a threshold. For example, if a temporary merge candidate is identical to an inter region merge candidate whose index is greater than or equal to a predetermined value, the temporary merge candidate may not be added to the inter region motion information table.
[0281] Intra prediction is a method of predicting a current block using reconstructed samples that have been completely coded / decoded around the current block. In this case, the intra prediction of the current block can use reconstructed samples before the in-loop filter is applied.
[0282] Intra prediction techniques include matrix-based intra prediction and general intra prediction that takes into account the directionality of surrounding reconstructed samples. Information indicating the intra prediction technique for the current block may be signaled via a bitstream. The information may be a one-bit flag. Alternatively, the intra prediction technique for the current block may be determined based on at least one of the position, size, and shape of the current block or the intra prediction techniques of neighboring blocks. For example, if the current block exists across images, it may be set so that matrix-based intra prediction is not applied to the current block.
[0283] Matrix-based intra prediction is a method of obtaining a prediction block for a current block based on a matrix multiplication between a matrix stored in the encoder and decoder and reconstructed samples surrounding the current block. Information for identifying one of a plurality of stored matrices can be signaled via the bitstream. The decoder can determine a matrix for intra prediction of the current block based on the information and the size of the current block.
[0284] General intra prediction is a method of obtaining a prediction block for a current block based on a non-directional intra prediction mode or a directional intra prediction mode. Hereinafter, the process of performing intra prediction based on general intra prediction will be described in more detail with reference to the accompanying drawings.
[0285] FIG. 28 is a flowchart of an intra prediction method according to an embodiment of the present invention.
[0286] A reference sample line of the current block may be determined (step S2801). The reference sample line refers to a set of reference samples included in the line k-th away from the top and / or left side of the current block. The reference sample may be derived from reconstructed samples that have been coded / decoded around the current block.
[0287] Index information identifying a reference sample line of the current block from among multiple reference sample lines can be signaled via the bitstream. The multiple reference sample lines can include at least one of the first, second, third, or fourth row / column from the top and / or left side of the current block. Table 4 shows the indexes assigned to each of the reference sample lines. In Table 4, it is assumed that the first, second, and fourth rows / columns use reference sample line candidates.
[0288] [Table 4]
[0289] The reference sample line of the current block may also be determined based on at least one of the position, size, shape, or predictive coding mode of the neighboring block of the current block. For example, if the current block is at the boundary of an image, a tile, a slice, or a coding tree unit, the first reference sample line may be determined as the reference sample line of the current block.
[0290] The reference sample line may include a top reference sample located at the top of the current block and a left reference sample located to the left of the current block. The top reference sample and the left reference sample may be derived from reconstructed samples around the current block. The reconstructed samples may be in a state before an in-loop filter is applied.
[0291] FIG. 29 is a drawing of the reference samples included in each reference sample line.
[0292] According to the intra prediction mode of the current block, at least one of the reference samples belonging to the reference sample line can be used to obtain a prediction sample.
[0293] Next, an intra-prediction mode of the current block may be determined (step S2802). The intra-prediction mode of the current block may be determined as at least one of a non-directional intra-prediction mode and a directional intra-prediction mode. The non-directional intra-prediction modes include Planar and DC, and the directional intra-prediction modes include 33 or 65 modes from the lower-left diagonal to the upper-right diagonal.
[0294] FIG. 30 is a diagram showing intra prediction modes.
[0295] (a) of FIG. 30 shows 35 intra prediction modes, and (b) of FIG. 30 shows 67 intra prediction modes.
[0296] A greater or lesser number of intra-prediction modes than those shown in FIG. 30 may also be defined.
[0297] A most probable mode (MPM) may be set based on the intra-prediction modes of neighboring blocks adjacent to the current block. Here, the neighboring blocks may include a left neighboring block adjacent to the left of the current block and an upper neighboring block adjacent to the top of the current block. If the coordinates of the top-left sample of the current block are (0,0), the left neighboring block may include a sample at (-1,0), (-1,H-1), or (-1,(H-1) / 2), where H represents the height of the current block. The upper neighboring block may include a sample at (0,-1), (W-1,-1), or ((W-1) / 2,-1), where W represents the width of the current block.
[0298] If the neighboring blocks are coded using general intra prediction, the MPM can be derived based on the intra prediction modes of the neighboring blocks. Specifically, the intra prediction mode of the left neighboring block can be set to a variable candIntraPredModeA, and the intra prediction mode of the upper neighboring block can be set to a variable candIntraPredModeB.
[0299] In this case, if the neighboring block is unavailable (e.g., if the neighboring block has not yet been coded / decoded or if the neighboring block is located beyond the image boundary), if the neighboring block is coded using matrix-based intra prediction, if the neighboring block is coded using inter prediction, or if the neighboring block is included in a different coding tree unit from the current block, the derived variable candIntraPredModeX (where X is A or B) can be set to a default mode based on the intra prediction mode of the neighboring block. Here, the default mode may include at least one of planar, DC, vertical mode, or horizontal mode.
[0300] Alternatively, if a neighboring block is coded using intra prediction based on a matrix, an intra prediction mode corresponding to an index value for specifying one of the matrices may be set to candIntraPredModeX. Therefore, a lookup table representing a mapping relationship between index values for specifying a matrix and intra prediction modes may be stored in the encoder and decoder.
[0301] The MPM can be derived based on the variables candIntraPredModeA and candIntraPredModeB. The number of MPMs included in the MPM list can be set by the encoder and decoder. As an example, the number of MPMs may be three, four, five, or six. Alternatively, information indicating the number of MPMs can be signaled via the bitstream. Alternatively, the number of MPMs can be determined based on at least one of the predictive coding modes of neighboring blocks, the size or shape of the current block.
[0302] In the embodiment described below, it is assumed that the number of MPMs is three, and the three MPMs are referred to as MPM[0], MPM[1], and MPM[2]. If the number of MPMs is more than three, the MPM is configured to include three MPMs as described in the embodiment described below.
[0303] If candIntraPredA and candIntraPredB are the same and candIntraPredA is planar or DC mode, MPM[0] and MPM[1] can be set to planar and DC mode, respectively. MPM[2] can be set to vertical intra prediction mode, horizontal intra prediction mode, or diagonal intra prediction mode. The diagonal intra prediction mode may be bottom-left diagonal intra prediction mode, top-left intra prediction mode, or top-right intra prediction mode.
[0304] If candIntraPredA and candIntraPredB are the same and candIntraPredA is a directional intra prediction mode, MPM[0] can be set to the same as candIntraPredA. MPM[1] and MPM[2] can set candIntraPredA to a similar intra prediction mode. An intra prediction mode similar to candIntraPredA may be an intra prediction mode whose index difference value from candIntraPredA is ±1 or ±2. A modular operation (%) and an offset can be used to derive an intra prediction mode similar to candIntraPredA.
[0305] If candIntraPredA and candIntraPredB are different, MPM[0] can be set to the same as candIntraPredA, and MPM[1] can be set to the same as candIntraPredB. In this case, if candIntraPredA and candIntraPredB are both non-directional intra prediction modes, MPM[2] can be set to a vertical intra prediction mode, a horizontal intra prediction mode, or a diagonal intra prediction mode. Alternatively, if at least one of candIntraPredA and candIntraPredB is a directional intra prediction mode, MPM[2] can be set to an intra prediction mode derived by adding or subtracting an offset to the maximum of planar, DC, or candIntraPredA or candIntraPredB. Here, the offset may be 1 or 2.
[0306] An MPM list including multiple MPMs may be generated, and information indicating whether the MPM list includes the same MPM as the intra prediction mode of the current block may be signaled via a bitstream. The information may be a 1-bit flag referred to as an MPM flag. If the MPM flag indicates that the MPM list includes the same MPM as the current block, index information identifying one of the MPMs may be signaled via the bitstream. The MPM identified by the index information may be set as the intra prediction mode of the current block. If the MPM flag indicates that the MPM list does not include the same MPM as the current block, remaining mode information indicating one of the remaining intra prediction modes excluding the MPM may be signaled via the bitstream. The remaining mode information indicates an index value corresponding to the intra prediction mode of the current block when indexes are reallocated to the remaining intra prediction modes excluding the MPM. The decoder may determine the intra prediction mode of the current block by arranging the MPMs in ascending order and comparing the remaining mode information with the MPM. As an example, if the residual mode information is less than or equal to the MPM, 1 may be added to the residual mode information to derive the intra-prediction mode of the current block.
[0307] Instead of setting the default mode to an MPM, information indicating whether the intra prediction mode of the current block is the default mode may be signaled via a bitstream. The information may be a 1-bit flag, and the flag may be referred to as a default mode flag. The default mode flag may be signaled only if the MPM flag indicates that the same MPM as that of the current block is included in the MPM list. As described above, the default mode may include at least one of planner, DC, vertical mode, or horizontal mode. For example, if planner is set as the default mode, the default mode flag may indicate whether the intra prediction mode of the current block is planner. If the default mode flag indicates that the intra prediction mode of the current block is not the default mode, one of the MPMs indicated by the index information may be set as the intra prediction mode of the current block.
[0308] If multiple intra prediction modes are set to a default mode, index information indicating one of the default modes may be further signaled, and the intra prediction mode of the current block may be set to the default mode indicated by the index information.
[0309] The default mode can be set not to be used if the index of the reference sample line of the current block is not 0. Thus, if the index of the reference sample line is not 0, the default mode flag can be set to a predefined value (i.e., false) without signaling the default mode flag.
[0310] Once the intra prediction mode of the current block is determined, prediction samples for the current block can be obtained based on the determined intra prediction mode (step S2803).
[0311] When the DC mode is selected, a predicted sample for the current block is generated based on the average value of the reference samples. Specifically, values of all samples in the predicted block can be generated based on the average value of the reference samples. The average value can be derived using at least one of an upper reference sample located at the top of the current block and a left reference sample located to the left of the current block.
[0312] The number or range of reference samples used to derive the average value may vary depending on the shape of the current block. For example, if the current block is a non-square block whose width is greater than its height, only the top reference sample may be used to calculate the average value. On the other hand, if the current block is a non-square block whose width is less than its height, only the left reference sample may be used to calculate the average value. That is, if the width and height of the current block are different, only the reference sample adjacent to the longer side may be used to calculate the average value. Alternatively, based on the ratio of the width and height of the current block, it may be determined whether to calculate the average value using only the top reference sample or only the left reference sample.
[0313] When the planar mode is selected, a prediction sample may be obtained using a horizontal prediction sample and a vertical prediction sample. Here, the horizontal prediction sample is obtained based on a left reference sample and a right reference sample located on the same horizontal line as the prediction sample, and the vertical prediction sample is obtained based on an upper reference sample and a lower reference sample located on the same vertical line as the prediction sample. Here, the right reference sample may be generated by duplicating a reference sample adjacent to the upper right corner of the current block, and the lower reference sample may be generated by duplicating a reference sample adjacent to the lower left corner of the current block. The horizontal prediction sample may be obtained based on a weighted sum operation of the left reference sample and the right reference sample, and the vertical prediction sample may be obtained based on a weighted sum operation of the upper reference sample and the lower reference sample. In this case, a weight value assigned to each reference sample may be determined according to the position of the prediction sample. The prediction sample may be obtained based on an average operation or a weighted sum operation of the horizontal prediction sample and the vertical prediction sample. When a weighted sum operation is performed, the weight values assigned to the horizontal prediction sample and the vertical prediction sample may be determined based on the position of the prediction sample.
[0314] When a directional prediction mode is selected, a parameter indicating the prediction direction (or prediction angle) of the selected directional prediction mode can be determined. Table 5 below shows the intra direction parameter intraPredAng for each intra prediction mode.
[0315] [Table 5]
[0316] When 35 intra prediction modes are defined, Table 5 shows the intra direction parameters of each intra prediction mode having an index of any one of 2 to 34. When more than 33 directional intra prediction modes are defined, Table 5 can be further subdivided to set the intra direction parameters of each directional intra prediction mode.
[0317] After aligning the top reference sample and the left reference sample of the current block, a predicted sample can be obtained based on the value of the intra direction parameter. In this case, if the value of the intra direction parameter is negative, the left reference sample and the top reference sample can be aligned.
[0318] 31 and 32 are diagrams showing examples of one-dimensional arrays in which reference samples are arranged in a line.
[0319] Figure 31 shows an example of a one-dimensional vertical array in which reference samples are arranged vertically, and Figure 32 shows an example of a one-dimensional horizontal array in which reference samples are arranged horizontally. The examples of Figures 31 and 32 will be described assuming that 35 intra prediction modes are defined.
[0320] If the intra prediction mode index is one of 11 to 18, a horizontal one-dimensional array in which the top reference sample is rotated counterclockwise is applied, and if the intra prediction mode index is one of 19 to 25, a vertical one-dimensional array in which the left reference sample is rotated clockwise is applied. When the reference samples are arranged in a row, the intra prediction mode angle can be taken into consideration.
[0321] Based on the intra direction parameters, reference sample decision parameters can be determined, which may include a reference sample index for identifying the reference sample and a weight value parameter for determining a weight value to be applied to the reference sample.
[0322] The reference sample index iIdx and the weight parameter ifact can be obtained by the following Equations 5 and 6, respectively.
[0323]
number
[0324]
number
[0325] In Equations 5 and 6, Pang represents an intra-direction parameter. The reference sample identified by the reference sample index iIdx corresponds to an integer pel.
[0326] At least one reference sample may be identified to derive a prediction sample. Specifically, the position of the reference sample used to derive the prediction sample may be identified taking into account the gradient of the prediction mode. For example, the reference sample index iIdx may be used to identify the reference sample used to derive the prediction sample.
[0327] In this case, if the slope of the intra prediction mode is not represented by one reference sample, a prediction sample can be generated by interpolating multiple reference samples. For example, if the slope of the intra prediction mode is a value between the slope between the prediction sample and a first reference sample and the slope between the prediction sample and a second reference sample, the prediction sample can be obtained by interpolating the first reference sample and the second reference sample. That is, if an angular line along the intra prediction angle does not pass through a reference sample located at an integer pel, the prediction sample can be obtained by interpolating reference samples located adjacent to the left, right, top, or bottom of the position where the angular line passes.
[0328] Below, Equation 7 shows an example of obtaining a predicted sample based on a reference sample.
[0329]
number
[0330] In Equation 7, P represents a predicted sample, and Ref_1D represents one of the one-dimensionally arranged reference samples. In this case, the position of the reference sample can be determined according to the position (x, y) of the predicted sample and the reference sample index iIdx.
[0331] If the gradient of the intra prediction mode can be represented by one reference sample, the weight parameter ifact is set to 0. This allows Equation 7 to be simplified to Equation 8 below.
[0332]
number
[0333] Intra prediction for the current block may also be performed based on multiple intra prediction modes. For example, an intra prediction mode may be derived for each prediction sample, and the prediction sample may be derived based on the intra prediction mode assigned to each prediction sample.
[0334] Alternatively, an intra prediction mode may be derived for each region, and intra prediction for each region may be performed based on the intra prediction mode assigned to the region. Here, the region may include at least one sample. At least one of the size or shape of the region may be adaptively determined based on at least one of the size, shape, and intra prediction mode of the current block. Alternatively, at least one of the size or shape of the region may be predefined in the encoder and decoder, regardless of the size or shape of the current block.
[0335] Alternatively, intra prediction may be performed based on each of multiple intra predictions, and a final predicted sample may be derived based on an average or weighted sum of multiple predicted samples obtained by the multiple intra predictions. For example, intra prediction may be performed based on a first intra prediction mode to obtain a first predicted sample, and intra prediction may be performed based on a second intra prediction mode to obtain a second predicted sample. Then, a final predicted sample may be obtained based on an average or weighted sum of the first and second predicted samples. In this case, weights assigned to the first and second predicted samples may be determined based on at least one of whether the first intra prediction mode is a non-directional / directional prediction mode, whether the second intra prediction mode is a non-directional / directional prediction mode, or the intra prediction mode of a neighboring block.
[0336] The multiple intra-prediction modes may be a combination of a non-directional intra-prediction mode and a directional prediction mode, a combination of directional prediction modes, or a combination of non-directional prediction modes.
[0337] FIG. 33 is a diagram illustrating angles formed by directional intra prediction modes with a line parallel to the x-axis.
[0338] As shown in the example of Figure 33, the directional prediction modes can exist between the bottom left diagonal and the top right diagonal. In terms of the angle formed by the x-axis and the directional prediction modes, the directional prediction modes can exist between 45 degrees (bottom left diagonal) and -135 degrees (top right diagonal).
[0339] If the current block is non-square, it may occur that, according to the intra prediction mode of the current block, a prediction sample is derived using a reference sample located on an angular line according to the intra prediction angle that is farther from the prediction sample instead of a reference sample that is closer to the prediction sample.
[0340] FIG. 34 is a diagram illustrating how prediction samples are obtained when the current block is non-square.
[0341] For example, assume that the current block is a non-square block whose width is greater than its height, as shown in (a) of Figure 34, and the intra prediction mode of the current block is a directional intra prediction mode having an angle between 0 and 45 degrees. In this case, when deriving a predicted sample A near the right column of the current block, it may be necessary to use the left reference sample L far from the predicted sample among the reference samples located in the angular mode according to the angle, instead of the top reference sample T close to the predicted sample.
[0342] As another example, assume that the current block is a non-square block with its height greater than its width, as shown in (b) of Figure 34, and the intra prediction mode of the current block is a directional intra prediction mode between -90 degrees and -135 degrees. In this case, when deriving a prediction sample A near the bottom row of the current block, it may be necessary to use the top reference sample T far from the prediction sample among the reference samples positioned in the angular mode according to the angle, instead of the left reference sample L close to the prediction sample.
[0343] To solve the above problem, if the current block is non-square, the intra prediction mode of the current block may be replaced with an intra prediction mode in the opposite direction. This allows a directional prediction mode with a larger or smaller angle than the directional prediction modes shown in FIG. 24 to be used for non-square blocks. Thus, the directional intra prediction mode may be defined as a wide-angle intra prediction mode. The wide-angle intra prediction mode refers to a directional intra prediction mode that does not fall within the range of 45 degrees to −135 degrees.
[0344] FIG. 35 is a diagram illustrating a wide-angle intra prediction mode.
[0345] In the example shown in FIG. 35, the intra prediction modes with indexes from −1 to −14 and the intra prediction modes with indexes from 67 to 80 represent wide-angle intra prediction modes.
[0346] Figure 35 shows 14 wide-angle intra-prediction modes with angles greater than 45 degrees (-1 to -14) and 14 wide-angle intra-prediction modes with angles less than -135 degrees (67 to 80), but more or fewer wide-angle intra-prediction modes can be defined.
[0347] When a wide-angle intra prediction mode is used, the length of the top reference sample may be set to 2W+1, and the length of the left reference sample may be set to 2H+1.
[0348] By using the wide-angle intra prediction mode, the reference sample T can be used to predict the sample A shown in (a) of Figure 34, and the reference sample L can be used to predict the sample A shown in (b) of Figure 34.
[0349] In addition to the existing intra prediction modes and N wide-angle intra prediction modes, a total of 67+N intra prediction modes can be used. As an example, Table 6 shows intra direction parameters of the intra prediction modes when 20 wide-angle intra prediction modes are defined.
[0350] [Table 6]
[0351] If the current block is non-square and the intra prediction mode of the current block obtained in step S2802 belongs to a transform range, the intra prediction mode of the current block may be transformed to a wide-angle intra prediction mode. The transform range may be determined based on at least one of the size, shape, or ratio of the current block. Here, the ratio may represent the ratio between the width and height of the current block.
[0352] If the current block is non-square, with its width greater than its height, the transform range may be set from the upper right diagonal intra prediction mode index (e.g., 66) to (the upper right diagonal intra prediction mode index - N), where N may be determined based on the ratio of the current block. If the intra prediction mode of the current block belongs to the transform range, the intra prediction mode may be converted to a wide-angle intra prediction mode. The conversion may be performed by subtracting a predefined value from the intra prediction mode, where the predefined value may be the total number of intra prediction modes excluding the wide-angle intra prediction mode (e.g., 67).
[0353] According to the above embodiment, the 66th to 53rd intra prediction modes can be converted to the −1st to −14th wide-angle intra prediction modes, respectively.
[0354] If the current block is non-square, with its height greater than its width, the transform range may be set from the lower-left diagonal intra-prediction mode index (e.g., 2) to (lower-left diagonal intra-prediction mode index + M), where M may be determined based on the ratio of the current block. If the intra-prediction mode of the current block belongs to the transform range, the intra-prediction mode may be converted to a wide-angle intra-prediction mode. The conversion may be performed by adding a predefined value to the intra-prediction mode, and the predefined value may be the total number of directional intra-prediction modes excluding the wide-angle intra-prediction mode (e.g., 65).
[0355] According to the above embodiment, each of the 2nd to 15th intra prediction modes can be converted into a 67th to 80th wide-angle intra prediction mode.
[0356] Hereinafter, the intra prediction modes that belong to the transform range will be referred to as wide-angle intra alternative prediction modes.
[0357] The transform range may be determined based on the ratio of the current block. As an example, Tables 7 and 8 show transform ranges when 35 intra prediction modes excluding the wide-angle intra prediction mode are defined and when 67 intra prediction modes are defined, respectively.
[0358] [Table 7]
[0359] [Table 8]
[0360] As shown in the examples of Tables 7 and 8, the number of wide-angle intra alternative prediction modes included in the transform range may vary according to the ratio of the current block.
[0361] As the wide-angle intra-prediction mode is used in addition to the existing intra-prediction mode, the resources required to encode the wide-angle intra-prediction mode may increase, which may result in a decrease in encoding efficiency. Therefore, instead of directly encoding the wide-angle intra-prediction mode, encoding an alternative intra-prediction mode to the wide-angle intra-prediction mode may improve encoding efficiency.
[0362] For example, if the current block is encoded using the 67th wide-angle intra prediction mode, the 67th wide-angle alternative intra prediction mode, No. 2, may be encoded as the intra prediction mode of the current block. Also, if the current block is encoded in the −1th wide-angle intra prediction mode, the −1st wide-angle alternative intra prediction mode, No. 66, may be encoded as the intra prediction mode of the current block.
[0363] The decoder may decode the intra-prediction mode of the current block and determine whether the decoded intra-prediction mode is included in the transform range. If the decoded intra-prediction mode is a wide-angle alternative intra-prediction mode, the decoder may convert the intra-prediction mode to a wide-angle intra-prediction mode.
[0364] Alternatively, if the current block is coded in a wide-angle intra-prediction mode, the wide-angle intra-prediction mode may be coded as is.
[0365] The encoding of the intra prediction mode may be implemented based on the above MPM list. Specifically, if a neighboring block is encoded in a wide-angle intra prediction mode, the MPM may be set based on a wide-angle alternative intra prediction mode corresponding to the wide-angle intra prediction mode. For example, if a neighboring block is encoded in a wide-angle intra prediction mode, the variable candIntraPredX (X is A or B) may be set to the wide-angle alternative intra prediction mode.
[0366] When a prediction block is generated through the execution result of intra prediction, prediction samples may be updated based on the positions of the prediction samples included in the prediction block. This updating method may be referred to as a position dependent prediction combination (PDPC) method.
[0367] Whether to use PDPC may be determined taking into account the intra prediction mode of the current block, the reference sample line of the current block, the size of the current block, or a color component. For example, PDPC may be used if the intra prediction mode of the current block is at least one of Planar, DC, vertical, horizontal, a mode with a smaller index value than the vertical, or a mode with a larger index value than the horizontal. Alternatively, PDPC may be used only if at least one of the width or height of the current block is greater than 4. Alternatively, PDPC may be used only if the index of the reference image line of the current block is 0. Alternatively, PDPC may be used only if the index of the reference image line of the current block is greater than or equal to a predefined value. Alternatively, PDPC may be used for only the luma component. Alternatively, whether to use PDPC may be determined depending on whether two or more of the above listed conditions are satisfied.
[0368] As another example, information indicating whether PDPC applies can be signaled via the bitstream.
[0369] When a prediction sample is obtained through an intra-prediction sample, a reference sample used to correct the prediction sample can be determined based on the position of the obtained prediction sample. For convenience of explanation, in the embodiments described below, the reference sample used to correct the prediction sample will be referred to as a PDPC reference sample. Furthermore, a prediction sample obtained through intra-prediction will be referred to as a first prediction sample, and a prediction sample obtained by correcting the first prediction sample will be referred to as a second prediction sample.
[0370] FIG. 36 is a diagram showing a mode in which PDPC is applied.
[0371] At least one PDPC reference sample may be used to correct the first predicted sample, and the PDPC reference sample may include at least one of a reference sample adjacent to the upper left corner of the current block, an upper reference sample located at the top of the current block, or a left reference sample located to the left of the current block.
[0372] At least one of the reference samples belonging to the reference sample line of the current block may be set as the PDPC reference sample. Alternatively, regardless of the reference sample line of the current block, at least one of the reference samples belonging to the reference sample line with index 0 may be set as the PDPC reference sample. For example, even if a first predicted sample is obtained using a reference sample included in a reference sample line with index 1 or 2, a second predicted sample may be obtained using a reference sample included in the reference sample line with index 0.
[0373] The number or positions of PDPC reference samples used to correct the first predicted sample may be determined taking into account at least one of the intra prediction mode of the current block, the size of the current block, the shape of the current block, or the positions of the first predicted sample.
[0374] For example, if the intra prediction mode of the current block is Planar or DC mode, a top reference sample and a left reference sample may be used to obtain a second predicted sample. In this case, the top reference sample may be a reference sample perpendicular to the first predicted sample (e.g., a reference sample having the same x-coordinate), and the left reference sample may be a reference sample horizontal to the first predicted sample (e.g., a reference sample having the same y-coordinate).
[0375] If the intra prediction mode of the current block is a horizontal intra prediction mode, the top reference sample may be used to obtain the second prediction sample, where the top reference sample may be a reference sample perpendicular to the first prediction sample.
[0376] If the intra prediction mode of the current block is a vertical intra prediction mode, the left reference sample may be used to obtain the second predicted sample, in which case the left reference sample may be a reference sample horizontal to the first predicted sample.
[0377] If the intra prediction mode of the current block is a diagonal-bottom or diagonal-top intra prediction mode, the second predicted sample may be obtained based on an upper-left reference sample, a top reference sample, and a left reference sample. The upper-left reference sample may be a reference sample adjacent to the upper-left corner of the current block (e.g., a reference sample at position (-1, -1)). The top reference sample may be a reference sample located diagonally to the upper right of the first predicted sample, and the left reference sample may be a reference sample located diagonally to the lower left of the first predicted sample.
[0378] In summary, if the position of the first predicted sample is (x, y), R(-1, -1) can be set as the top-left reference sample, R(x+y+1, -1) or R(x, -1) can be set as the top reference sample, and R(-1, x+y+1) or R(-1, y) can be set as the left reference sample.
[0379] As another example, the position of the left reference sample or the top reference sample may be determined taking into account at least one of the shape of the current block or whether a wide-angle intra mode is applied.
[0380] Specifically, when the intra prediction mode of the current block is a wide-angle intra prediction mode, reference samples spaced apart by an offset from reference samples diagonally opposite the first predicted sample may be set as PDPC reference samples. For example, the top reference sample R(x+y+k+1,-1) and the left reference sample R(-1,x+y-k+1) may be set as PDPC reference samples.
[0381] In this case, the offset k can be determined based on the wide-angle intra prediction mode. Equations 9 and 10 show examples of deriving the offset based on the wide-angle intra prediction mode.
[0382]
number
[0383]
number
[0384] The second predicted sample may be determined based on a weighted sum operation between the first predicted sample and the PDPC reference sample. For example, the second predicted sample may be obtained based on the following Equation 11:
[0385]
number
[0386] In Equation 11, RL represents the left reference sample, RT represents the top reference sample, and RTL represents the top-left reference sample. Pred(x,y) represents a predicted sample at the (x,y) position. wL represents a weight assigned to the left reference sample, wT represents a weight assigned to the top reference sample, and wTL represents a weight assigned to the top-left reference sample. The weight assigned to the first predicted sample can be derived by subtracting the weight assigned to the reference sample from the maximum value. For convenience of explanation, the weight assigned to the PDPC reference sample will be referred to as the PDPC weight.
[0387] The weight value assigned to each reference sample may be determined based on at least one of the intra prediction mode of the current block or the position of the first prediction sample.
[0388] For example, at least one of wL, wT, or wTL may be proportional or inversely proportional to at least one of the x-axis coordinate value or the y-axis coordinate value of the predicted sample, or at least one of wL, wT, or wTL may be proportional or inversely proportional to at least one of the width or the height of the current block.
[0389] If the intra prediction mode of the current block is DC, the PDPC weight value may be determined by Equation 12 below.
[0390]
number
[0391] In Equation 12, x and y represent the position of the first predicted sample.
[0392] In Equation 12, the variable shift used in the bit shift operation can be derived based on the width or height of the current block. As an example, the variable shift can be derived based on Equation 13 or Equation 14 below.
[0393]
number
[0394]
number
[0395] Alternatively, the variable shift can be derived by taking into account the intra direction parameters of the current block.
[0396] The number or type of parameters used to derive the variable "shift" may be determined differently depending on the intra prediction mode of the current block. As an example, if the intra prediction mode of the current block is Planner, DC, vertical, or horizontal, the variable "shift" may be derived using the width and height of the current block, as shown in Equation 13 or 14. If the intra prediction mode of the current block is an intra prediction mode with a larger index than the vertical intra prediction mode, the variable "shift" may be derived using the height and intra direction parameters of the current block. If the intra prediction mode of the current block is an intra prediction mode with a smaller index than the horizontal intra prediction mode, the variable "shift" may be derived using the width and intra direction parameters of the current block.
[0397] If the intra prediction mode of the current block is planar, the value of wTL may be set to 0. wL and wT may be derived based on Equation 15 below.
[0398]
number
[0399] If the intra prediction mode of the current block is a horizontal intra prediction mode, wT may be set to 0, and wTL and wL may be set to the same value. On the other hand, if the intra prediction mode of the current block is a vertical intra prediction mode, wL may be set to 0, and wTL and wT may be set to the same value.
[0400] If the intra prediction mode of the current block is an intra prediction mode pointing in the upper right direction and having an index value greater than the vertical intra prediction mode, the PDPC weight value can be derived as shown in Equation 16 below.
[0401]
number
[0402] On the other hand, if the intra prediction mode of the current block is an intra prediction mode facing the bottom left direction and having an index value smaller than the horizontal intra prediction mode, the PDPC weight value can be derived as shown in Equation 17 below.
[0403]
number
[0404] As in the above example, the PDPC weight values can be determined based on the x and y positions of the predicted samples.
[0405] As another example, weights assigned to the PDPC reference samples may be determined on a sub-block basis, and the prediction samples included in the sub-block may share the same PDPC weights.
[0406] The size of the sub-block, which is the basic unit for determining the weight value, may be predefined by the encoder and decoder. For example, a weight value may be determined for each of the sub-blocks having a size of 2x2 or 4x4.
[0407] Alternatively, the size, shape, or number of sub-blocks may be determined according to the size or shape of the current block. For example, a coding block may be divided into four sub-blocks regardless of the size of the coding block. Alternatively, a coding block may be divided into four or sixteen sub-blocks according to the size of the coding block.
[0408] Alternatively, the size, shape, or number of sub-blocks may be determined based on the intra-prediction mode of the current block. For example, if the intra-prediction mode of the current block is horizontal, N columns (or N rows) may be set as one sub-block, while if the intra-prediction mode of the current block is vertical, N rows (or N columns) may be set as one sub-block.
[0409] Equations 18 to 20 show an example of determining the PDPC weight value for a 2x2 size sub-block. Equation 18 shows the case where the intra prediction mode of the current block is the DC mode.
[0410]
number
[0411] In Equation 18, K can be determined based on the size of the sub-block.
[0412] Equation 19 indicates the case where the intra prediction mode of the current block is an intra prediction mode pointing in the upper right direction and having an index value greater than the vertical intra prediction mode.
[0413]
number
[0414] Equation 20 indicates the case where the intra prediction mode of the current block is an intra prediction mode that faces the bottom left direction and has an index value smaller than the horizontal intra prediction mode.
[0415]
number
[0416] In Equations 18 to 20, x and y represent the position of the reference sample within the sub-block, which may be one of the sample located at the top left, center, or bottom right of the sub-block.
[0417] Equations 21 to 23 show an example of determining the PDPC weight value for a 4x4 size sub-block. Equation 21 shows the case where the intra prediction mode of the current block is the DC mode.
[0418]
number
[0419] Equation 22 indicates the case where the intra prediction mode of the current block is an intra prediction mode pointing in the upper right direction and having an index value greater than the vertical intra prediction mode.
[0420]
number
[0421] Equation 23 indicates the case where the intra prediction mode of the current block is an intra prediction mode that faces the bottom left direction and has an index value smaller than the horizontal intra prediction mode.
[0422]
number
[0423] In the above embodiment, the PDPC weights are determined taking into consideration the positions of the first predicted sample or the predicted samples included in the sub-block. However, the PDPC weights may be determined taking into consideration the shape of the current block.
[0424] As an example, for DC mode, the PDPC weight values may be derived differently depending on whether the current block is a non-square block with width greater than height or height greater than width.
[0425] Equation 24 shows an example of deriving PDPC weight values when the current block is non-square, with its width greater than its height, and Equation 25 shows an example of deriving PDPC weight values when the current block is non-square, with its height greater than its width.
[0426]
number
[0427]
number
[0428] If the current block is non-square, it can be predicted using a wide-angle intra prediction mode. Thus, even when the wide-angle intra prediction mode is applied, the first prediction sample can be updated by applying PDPC.
[0429] If wide-angle intra prediction is applied to the current block, the PDPC weight value can be determined taking into account the shape of the coding block.
[0430] For example, if the current block is non-square and its width is greater than its height, the upper reference sample located to the upper right of the first predicted sample may be closer to the first predicted sample than the left reference sample located to the lower left of the first predicted sample, depending on the position of the first predicted sample. Therefore, when correcting the first predicted sample, the weight applied to the upper reference sample may be set to be greater than the weight applied to the left reference sample.
[0431] On the other hand, if the current block is non-square with its height greater than its width, the left reference sample located at the lower left of the first predicted sample may be closer to the first predicted sample than the top reference sample located at the upper right of the first predicted sample, depending on the position of the first predicted sample. Therefore, when correcting the first predicted sample, the weight applied to the left reference sample may be set to be greater than the weight applied to the top reference sample.
[0432] Equation 26 shows an example of deriving the PDPC weight value when the intra prediction mode of the current block is a wide-angle intra prediction mode with an index greater than 66.
[0433]
number
[0434] Equation 27 shows an example of deriving the PDPC weight value when the intra prediction mode of the current block is a wide-angle intra prediction mode with an index less than 0.
[0435]
number
[0436] The PDPC weight value can also be determined based on the ratio of the current block. The ratio of the current block indicates the ratio of the width to the height of the current block and can be defined as follows:
[0437]
number
[0438] The method for deriving the PDPC weight values can be variably determined according to the intra prediction mode of the current block.
[0439] For example, Equation 29 and Equation 30 show an example of deriving a PDPC weight value when the intra prediction mode of the current block is DC. Specifically, Equation 29 is an example when the current block is a non-square block whose width is greater than its height, and Equation 30 is an example when the current block is a non-square block whose height is greater than its width.
[0440]
number
[0441]
number
[0442] Equation 31 shows an example of deriving the PDPC weight value when the intra prediction mode of the current block is a wide-angle intra prediction mode with an index greater than 66.
[0443]
number
[0444] Equation 32 shows an example of deriving the PDPC weight value when the intra prediction mode of the current block is a wide-angle intra prediction mode with an index less than 0.
[0445]
number
[0446] A single prediction mode can be applied multiple times to the current block, or multiple prediction modes can be applied overlappingly. In this way, a prediction method using the same or different types of prediction modes can be called a combined prediction mode (or a multi-hypothesis prediction mode).
[0447] Information indicating whether the current block applies a combined prediction mode may be signaled via a bitstream. For example, the information may be a 1-bit flag.
[0448] The combined prediction mode can generate a first predicted block based on a first prediction mode, generate a second predicted block based on a second prediction mode, and generate a third predicted block based on a weighted sum of the first and second predicted blocks. The third predicted block can be set as the final predicted block for the current block.
[0449] The combined prediction modes may include at least one of a merge mode and a combined merge mode, a combined inter prediction and intra prediction mode, a combined merge mode and a motion vector prediction mode, or a combined merge mode and an intra prediction mode.
[0450] The merge mode and the combined merge mode can perform motion compensation prediction using multiple merge candidates. Specifically, a first predicted block can be generated using a first merge candidate, and a second predicted block can be generated using a second merge candidate. A third predicted block can be generated based on a weighted sum of the first predicted block and the second predicted block.
[0451] Information for identifying the first merge candidate and the second merge candidate may be signaled via the bitstream. For example, index information merge_idx for identifying the first merge candidate and index information merge_2nd_idx for identifying the second merge candidate may be signaled via the bitstream. The second merge candidate may be determined based on the index information merge_2nd_idx and the index information merge_idx.
[0452] The index information merge_idx identifies one of the merge candidates included in the merge candidate list.
[0453] The index information merge_2nd_idx can identify any one of the remaining merge candidates other than the merge candidate identified by merge_idx. As a result, if the value of merge_2nd_idx is smaller than merge_idx, the merge candidate whose index is the value of merge_2nd_idx can be set as the second merge candidate. If the value of merge_2nd_idx is greater than or equal to the value of merge_idx, the merge candidate whose index is the value of merge_2nd_idx plus 1 can be set as the second merge candidate.
[0454] Alternatively, the second merging candidate can be identified by taking into consideration the search order of the candidate blocks.
[0455] FIG. 37 shows an example of identifying second merging candidates taking into consideration the search order of candidate blocks.
[0456] In the example shown in Figure 37, the indexes written in the adjacent and non-adjacent samples indicate the search order of the candidate blocks. For example, the candidate blocks can be searched sequentially from position A0 to position A14.
[0457] If block A4 is selected as the first merge candidate, a merge candidate derived from the candidate block next in search order to A4 can be identified as the second merge candidate. For example, a merge candidate derived from A5 can be selected as the second merge candidate. If the candidate block at position A5 is unavailable as a merge candidate, a merge candidate derived from the next candidate block can be selected as the second merge candidate.
[0458] The first and second merging candidates may also be selected from among merging candidates derived from non-adjacent blocks.
[0459] FIG. 38 illustrates an example in which first and second merging candidates are selected from merging candidates derived from non-adjacent blocks.
[0460] As shown in the example of Figure 38, merge candidates derived from a first candidate block and a second candidate block that are not adjacent to the current block may be selected as the first merge candidate and the second merge candidate, respectively. In this case, the block line to which the first candidate block belongs and the block line to which the second candidate block belongs may be different. As an example, the first merge candidate may be derived from any one of candidate blocks A5 to A10, and the second merge candidate may be derived from any one of candidate blocks A11 to A15.
[0461] Alternatively, the first and second candidate blocks can be set so as not to be included in the same line (for example, row or column).
[0462] As another example, a second merge candidate may be identified based on a first merge candidate. In this case, the first merge candidate may be identified by index information merge_idx signaled from the bitstream. For example, a merge candidate adjacent to the first merge candidate may be identified as the second merge candidate. Here, a merge candidate adjacent to the first merge candidate may refer to a merge candidate whose index difference value from the first merge candidate is 1. For example, a merge candidate with an index value of merge_idx+1 may be set as the second merge candidate. In this case, if the value of merge_idx+1 is greater than the maximum index value (or if the index value of the first merge candidate is the maximum index), a merge candidate with an index value of merge_idx-1 or a merge candidate with an index value of a predefined value (e.g., 0) may be set as the second merge candidate.
[0463] Alternatively, a merge candidate adjacent to the first merge candidate may refer to a merge candidate derived from a candidate block spatially adjacent to the candidate block used to derive the first merge candidate, where a neighboring candidate block of a candidate block may refer to a block adjacent to the left, right, top, bottom, or diagonal of the candidate block.
[0464] As another example, the second merging candidate can be identified based on motion information of the first merging candidate. For example, a merging candidate that shares the same reference image as the first merging candidate can be selected as the second merging candidate. If there are multiple merging candidates that share the same reference image as the first merging candidate, the merging candidate with the smallest index or the smallest index difference from the first merging candidate can be selected as the second merging candidate. Alternatively, the second merging candidate can be selected based on index information that identifies one of the multiple merging candidates.
[0465] Alternatively, if the first merge candidate is unidirectional prediction in the first direction, a merge candidate including motion information for the second direction may be set as the second merge candidate. For example, if the first merge candidate has motion information in the L0 direction, a merge candidate having motion information in the L1 direction may be selected as the second merge candidate. If there are multiple merge candidates with motion information in the L1 direction, the merge candidate with the smallest index among the multiple merge candidates or the merge candidate with the smallest index difference from the first merge candidate may be set as the second merge candidate. Alternatively, the second merge candidate may be selected based on index information identifying any one of the multiple merge candidates.
[0466] As another example, one of the merge candidates derived from adjacent blocks adjacent to the current block can be set as the first merge candidate, and one of the merge candidates derived from non-adjacent blocks not adjacent to the current block can be set as the second merge candidate.
[0467] As another example, one of the merge candidates derived from the candidate block located above the current block can be set as the first merge candidate, and one of the merge candidates derived from the candidate block located to the left can be set as the second merge candidate.
[0468] A combined predicted block may be obtained by performing a weighted sum operation on a first predicted block derived from a first merging candidate and a second predicted block derived based on a second merging candidate. In this case, a weight applied to the first predicted block may be set to a value greater than a weight applied to the second predicted block.
[0469] Alternatively, the weighting value may be determined based on the motion information of the first merging candidate and the motion information of the second merging candidate. For example, the weighting value to be applied to the first predicted block and the second predicted block may be determined based on the difference in output order between the reference image and the current image. Specifically, the weighting value to be applied to the predicted block may be set to a smaller value as the difference in output order between the reference image and the current image increases.
[0470] Alternatively, the weights to be applied to the first and second predicted blocks may be determined taking into consideration the sizes or shapes of a candidate block used to derive the first merge candidate (hereinafter referred to as the first candidate block) and a candidate block used to derive the second merge candidate (hereinafter referred to as the second candidate block). For example, the weight to be applied to the predicted block derived from the first and second candidate blocks because it has a shape similar to that of the current block may be set to a large value. On the other hand, the weight to be applied to the predicted block derived from the shape dissimilar to that of the current block may be set to a small value.
[0471] FIG. 39 illustrates an example in which a weight value to be applied to a prediction block is determined based on the type of a candidate block.
[0472] Assume the current block is non-square, with width greater than height.
[0473] A first prediction block and a second prediction block may be derived based on the first merge candidate and the second merge candidate, and a combined prediction block may be generated based on a weighted sum operation of the first prediction block and the second prediction block. In this case, weights to be applied to the first prediction block and the second prediction block may be determined based on the shapes of the first candidate block and the second candidate block.
[0474] For example, in the example shown in Figure 39, the first candidate block is square, and the second candidate block is non-square, with its width greater than its height. Because the shape of the second candidate block is the same as the current block, the weight value applied to the second predicted block may be set to be greater than the weight value applied to the first predicted block. For example, a weight value of 5 / 8 may be applied to the second predicted block, and a weight value of 3 / 8 may be applied to the first predicted block. Equation 33 shows an example of deriving a combined predicted block based on a weighted sum operation of the first predicted block and the second predicted block.
[0475]
number
[0476] P(x,y) represents the combined prediction block, P1(x,y) represents the first prediction block, and P2(x,y) represents the second prediction block.
[0477] As another example, weight values to be applied to the first and second predicted blocks may be determined based on the shape of the current block. For example, if the current block is non-square in which the width is greater than the height, a larger weight value may be applied to the predicted block generated based on one of the first and second merge candidates, which is derived based on the candidate block located at the top of the current block. If both the first and second merge candidates are derived from the candidate block located at the top, the weight values to be applied to the first and second predicted blocks may be set to the same. On the other hand, if the current block is non-square in which the height is greater than the width, a larger weight value may be applied to the predicted block generated based on one of the first and second merge candidates, which is derived based on the candidate block located at the left of the current block. If both the first and second merge candidates are derived from the candidate block located at the left, the weight values to be applied to the first and second predicted blocks may be set to the same. If the current block is square, the weight values to be applied to the first and second predicted blocks may be set to the same.
[0478] As another example, a weight value to be applied to each prediction block may be determined based on the distance between the current block and the candidate block. Here, the distance may be derived based on the x-axis coordinate difference, the y-axis coordinate difference, or the minimum value thereof. The weight value to be applied to a prediction block derived from a merge candidate having a small distance from the current block may be set to be greater than the weight value to be applied to a prediction block derived from a merge candidate having a large distance from the current block. As an example, in the example shown in FIG. 37, the first merge candidate is derived from a neighboring block adjacent to the current block, and the second merge candidate is derived from a non-neighboring block not adjacent to the current block. In this case, since the x-axis distance between the first candidate block and the current block is smaller than the x-axis distance between the second candidate block and the current block, the weight value to be applied to the first prediction block may be set to be greater than the weight value to be applied to the second prediction block.
[0479] Alternatively, if both the first and second merging candidates are derived from non-adjacent blocks, a larger weight value may be assigned to the predicted block derived from the non-adjacent block that is closer to the current block. For example, in the example shown in FIG. 38, since the y-axis distance between the first candidate block and the current block is smaller than the y-axis distance between the second candidate block and the current block, the weight value applied to the first predicted block may be set to be larger than the weight value applied to the second predicted block.
[0480] In the combined prediction mode obtained by combining the merge modes, the merge mode may refer to a merge mode based on a translational motion model (hereinafter referred to as a translational merge mode) or a merge mode based on an affine motion model (hereinafter referred to as an affine merge mode). That is, motion compensation prediction may be performed by combining a translational merge mode with a translational merge mode or by combining an affine merge mode with an affine merge mode.
[0481] For example, if the first merge candidate is an affine merge candidate, the second merge candidate may also be set as an affine merge candidate. Here, the affine merge candidate refers to a case where the motion vector of the block including the reference candidate is an affine motion vector. The second merge candidate may be identified using the various embodiments described above. For example, the second merge candidate may be set as a neighboring merge candidate of the first merge candidate. In this case, if a merge candidate neighboring the first merge candidate is not coded using an affine motion model, a merge candidate coded using an affine motion model may be set as the second merge candidate instead of the first merge candidate.
[0482] Conversely, if the first merge candidate is a non-affine merge candidate, the second merge candidate can also be set as a non-affine merge candidate. In this case, if a merge candidate adjacent to the first merge candidate is coded using an affine motion model, a merge candidate coded using a translational motion model can be set as the second merge candidate instead of the first merge candidate.
[0483] FIG. 40 is a diagram illustrating an example in which a non-affine merge candidate is set as the second merge candidate instead of an affine merge candidate.
[0484] When a merge candidate at position A1 is identified as the first merge candidate via merge_idx, merge candidate A2, which has an index value one greater than the first merge candidate, can be selected as the second merge candidate. In this case, if the first merge candidate is a non-affine merge candidate but the second merge candidate is an affine merge candidate, the second merge candidate can be reset. As an example, among merge candidates with an index greater than merge_idx+1, the non-affine merge candidate with a smaller difference value from merge_idx+1 can be reset as the second merge candidate. As an example, the example shown in FIG. 18 indicates that merge candidate A3, whose index is merge_idx+2, is set as the second merge candidate.
[0485] As another example, the translational merge mode and the affine merge mode can be combined to perform motion compensated prediction, i.e., one of the first merge candidate or the second merge candidate can be an affine merge candidate, and the other can be a non-affine merge candidate.
[0486] It is also possible to derive integrated motion information based on the first and second merge candidates, and perform motion compensation prediction for the current block based on the integrated motion information. For example, the motion vector of the current block may be derived based on an average or weighted sum of the motion vectors of the second merge candidate among the motion vectors of the first merge candidate. In this case, the weights applied to the motion vectors of the first merge candidate and the motion vectors of the second merge candidate may be determined according to the above-described embodiment.
[0487] If the first merge candidate is a non-affine merge candidate and the second affine merge candidate is an affine merge candidate, the motion vector of the current block can be derived by scaling the motion vector of the second merge candidate. Equation 34 shows an example of deriving the motion vector of the current block.
[0488]
number
[0489] In Equation 34, (mvX, mvY) represents the motion vector of the current block, (mv0x, mv0y) represents the motion vector of the first merging candidate, and (mv1x, mv1y) represents the motion vector of the second merging candidate. M represents a scaling parameter. M may be predefined in the encoder and decoder. Alternatively, the value of the scaling parameter M may be determined according to the size of the current block or the candidate block. For example, if the width or height of the second candidate block is greater than 32, M may be set to 3; otherwise, M may be set to 2.
[0490] In a prediction mode that combines merge mode and motion vector prediction mode, a first prediction block can be generated using motion information derived from a merge candidate, and a second prediction block can be generated using a motion vector derived from a motion vector prediction candidate.
[0491] In the motion vector prediction mode, a motion vector prediction candidate can be derived from a neighboring block adjacent to the current block or a block at the same position in the same image. Then, one of the motion vector prediction candidates can be identified and set as the motion vector prediction value of the current block. Then, the motion vector of the current block can be derived by adding the motion vector prediction value of the current block and the motion vector difference value.
[0492] In a prediction mode that combines a merge mode and a motion vector prediction mode, a merge candidate and a motion vector prediction candidate may be derived from the same candidate block. For example, when a merge candidate is identified via merge_idx, the motion vector of the candidate block used to derive the identified merge candidate may be set as the motion vector predictor. Alternatively, when a motion vector prediction candidate is identified via mvp_flag, a merge candidate derived from the candidate block used to derive the identified merge candidate may be selected.
[0493] Alternatively, the candidate block used to derive the merge candidate may be different from the candidate block used to derive the motion vector prediction candidate. For example, when a merge candidate derived from a candidate block located above the current block is selected, a setting may be made to select a motion vector prediction candidate derived from a candidate block located to the left of the current block.
[0494] Alternatively, if the merge candidate selected by the index information and the motion vector prediction candidate selected by the index information are derived from the same candidate block, the motion vector prediction candidate can be replaced with a motion vector prediction candidate derived from a candidate block adjacent to the candidate block, or the merge candidate can be replaced with a merge candidate derived from a candidate block adjacent to the candidate block.
[0495] FIG. 41 is a diagram showing an example in which merging candidates are replaced.
[0496] In the example shown in (a) of Figure 41, a merge candidate and a motion vector prediction candidate derived from a candidate block located at position A2 are selected. As shown, when a merge candidate and a motion vector prediction candidate are derived from the same candidate block, a merge candidate or a motion vector prediction candidate derived from a candidate block adjacent to the candidate block can be used instead of the merge candidate or the motion vector prediction candidate. As an example, as shown in (b) of Figure 41, a merge candidate located at position A1 can be used instead of the merge candidate located at position A2.
[0497] A first predicted block may be derived based on a merge candidate for the current block, and a second predicted block may be derived based on a motion vector prediction candidate. Then, a combined predicted block may be derived by performing a weighted sum operation on the first predicted block and the second predicted block. In this case, a weight applied to the second predicted block generated in the motion vector prediction mode may be set to be greater than a weight applied to the first predicted block generated in the merge mode.
[0498] A residual image can be derived by subtracting a predicted image from an original image. In this case, when the residual image is converted into the frequency domain, removing high-frequency components from the frequency components does not significantly reduce the subjective image quality of the image. Therefore, by reducing the values of the high-frequency components or setting them to zero, compression efficiency can be improved without significant visual distortion. Reflecting the above characteristics, the current block can be transformed to decompose the residual image into two-dimensional frequency components. The transformation can be performed using a transform technique such as a discrete cosine transform (DCT) or a discrete sine transform (DST).
[0499] DCT decomposes (or transforms) a residual image into two-dimensional frequency components using a cosine transform, while DST decomposes (or transforms) a residual image into two-dimensional frequency components using a sine transform. The frequency components resulting from the transformation of the residual image can be expressed as basic images. For example, when a DCT transform is performed on an NxN block, N2 basic pattern components can be obtained. The size of each of the basic pattern components included in the NxN block can be obtained through the transform. Depending on the transform technique used, the size of the basic pattern components can be referred to as DCT coefficients or DST coefficients.
[0500] The DCT transform technique is mainly used to transform images with a large distribution of non-zero low-frequency components, while the DST transform technique is mainly used for images with a large distribution of high-frequency components.
[0501] Transformation techniques other than DCT or DST can also be used to transform the residual image.
[0502] Hereinafter, converting a residual image into two-dimensional frequency components may be referred to as two-dimensional image conversion. Furthermore, the size of a basic pattern component obtained by the conversion result will be referred to as a transform coefficient. For example, the transform coefficient may refer to a DCT coefficient or a DST coefficient. When both a first transform and a second transform (to be described later) are applied, the transform coefficient may refer to the size of a basic pattern component generated by the second transform result.
[0503] The transform technique may be determined on a block-by-block basis. The transform technique may be determined based on at least one of the predictive coding mode of the current block, the size of the current block, or the size of the current block. For example, if the current block is coded in intra prediction mode and the size of the current block is smaller than NxN, the transform technique DST may be used to perform the transform. On the other hand, if the above conditions are not met, the transform technique DCT may be used to perform the transform.
[0504] For some blocks of the residual image, 2D image transform may not be performed. Not performing 2D image transform may be referred to as a transform skip. When a transform skip is applied, quantization may be applied to residual values that have not been transformed.
[0505] After transforming the current block using a DCT or DST, the transformed current block may be retransformed. In this case, the transformation based on the DCT or DST may be defined as a first transformation, and retransforming the block to which the first transformation was applied may be defined as a second transformation.
[0506] The first transform may be performed using any one of a number of candidate transform cores. As an example, the first transform may be performed using any one of a DCT2, a DCT8, or a DCT7.
[0507] Different transform cores can be used for the horizontal and vertical directions. Information indicating the combination of horizontal and vertical transform cores can also be signaled via the bitstream.
[0508] The first and second transforms may be performed in different units. For example, the first transform may be performed on an 8x8 block, and the second transform may be performed on a 4x4 sub-block of the transformed 8x8 block. In this case, the transform coefficients of the remaining area where the second transform has not been performed may be set to 0.
[0509] Alternatively, a first transform can be performed on a 4x4 block, and a second transform can be performed on an 8x8 sized region containing the transformed 4x4 block.
[0510] Information indicating whether or not a second transformation is performed can be signaled via the bitstream.
[0511] The decoder may perform an inverse transform (second inverse transform) of the second transform, and then perform an inverse transform (first inverse transform) of the first transform on the result of the second inverse transform. A residual signal for the current block may be obtained as a result of the second inverse transform and the first inverse transform.
[0512] Quantization is used to reduce the energy of a block, and the quantization process involves dividing the transform coefficients by a certain constant, which may be derived from a quantization parameter, which can be defined as a value between 1 and 63.
[0513] Once the encoder performs the transform and quantization, the decoder can obtain the residual block through inverse quantization and inverse transform, and can add the prediction block and the residual block to obtain a reconstructed block for the current block.
[0514] Once a reconstructed block for the current block is obtained, information loss that occurs during quantization and encoding can be reduced through in-loop filtering. The in-loop filter may include at least one of a deblocking filter, a sample adaptive offset filter (SAO), or an adaptive loop filter (ALF). Hereinafter, the reconstructed block before the in-loop filter is applied will be referred to as a first reconstructed block, and the reconstructed block after the in-loop filter is applied will be referred to as a second reconstructed block.
[0515] At least one of a deblocking filter, SAO, or ALF may be applied to the first reconstructed block to obtain a second reconstructed block, where SAO or ALF may be applied after the deblocking filter is applied.
[0516] The deblocking filter is used to mitigate block boundary degradation (blocking artifacts) that occur due to block-by-block quantization. To apply the deblocking filter, the blocking strength (BS) between a first reconstructed block and an adjacent reconstructed block can be determined.
[0517] FIG. 42 is a flow chart illustrating the process of determining the strength of a block.
[0518] In the example shown in Figure 42, P represents the first reconstructed block and Q represents the adjacent reconstructed block, where the adjacent reconstructed block can be adjacent to the left or top of the current block.
[0519] The example shown in Figure 42 shows that the block strength is determined taking into account the predictive coding modes of P and Q, whether they contain non-zero transform coefficients, whether they are inter-predicted using the same reference image, or whether the motion vector difference value is greater than or equal to a threshold.
[0520] Whether a deblocking filter is applied can be determined based on the block strength. For example, if the block strength is 0, no filtering can be performed.
[0521] SAO is intended to mitigate ringing artifacts caused by quantization in the frequency domain. SAO can be performed by adding or subtracting an offset determined by taking into account the pattern of the first reconstructed image. Methods for determining the offset include edge offset (EO) and band offset. EO refers to a method for determining the offset of a current sample according to the pattern of surrounding pixels. BO refers to a method for applying a common offset to a group of pixels having similar brightness values within a region. Specifically, pixel brightness can be divided into 32 equal intervals, and pixels having similar brightness can be grouped together. For example, four adjacent bands out of 32 bands can be grouped together, and the same offset value can be applied to samples belonging to the four bands.
[0522] ALF is a method of applying a filter of a predefined size / shape to a first reconstructed image or a reconstructed image to which a deblocking filter has been applied to generate a second reconstructed image. Equation 35 shows an example of application of ALF.
[0523]
number
[0524] One of predefined filter candidates can be selected for each image, coding tree unit, coding block, prediction block, or transform block, and each filter candidate may differ in either size or shape.
[0525] FIG. 43 shows predefined filter candidates.
[0526] As an example shown in FIG. 43, at least one of diamond shapes of 5x5, 7x7 or 9x9 size can be selected.
[0527] For the chroma components, only a diamond shape of size 5x5 can be used.
[0528] It is within the scope of the present invention to apply embodiments described in the context of a decoding or encoding process to an encoding or decoding process, and it is also within the scope of the present invention to rearrange embodiments described in a given order into an order different from that described.
[0529] Although the above embodiments have been described based on a series of steps or flowcharts, this does not limit the chronological sequence of the invention, and steps may be executed simultaneously or in a different order as needed. Furthermore, in the above embodiments, each of the components (e.g., units, modules, etc.) constituting the block diagrams may be embodied as a hardware device or software, or multiple components may be combined and embodied as a single hardware device or software. The above embodiments may be embodied in the form of program instructions that can be executed by various computer components and stored on a computer-readable storage medium. The computer-readable storage medium may include program instructions, data files, data structures, etc., independently or in combination. Examples of computer-readable storage media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical storage media such as CD-ROMs and DVDs; magneto-optical media such as optical disks; and hardware devices specially configured to store and execute program instructions, such as ROM, RAM, and flash memory. The hardware devices may be configured to execute one or more software modules to perform the processes of the present invention, and vice versa. [Industrial Applicability]
[0530] The present invention can be applied to electronic devices that encode / decode video.
Claims
1. 1. A video decoding method, comprising: generating a merge candidate list for the first block; selecting one of the merge candidates included in the merge candidate list; performing motion compensation for the first block based on motion information of the selected merging candidate; performing motion compensation on the first block includes performing motion compensated prediction of the first block using a plurality of merging candidates; the plurality of merging candidates include a first merging candidate and a second merging candidate, the first merging candidate and the second merging candidate being included in a merging candidate list for the first block; a first prediction block is generated using the first merging candidate, a second prediction block is generated using the second merging candidate, a third prediction block is generated based on the first prediction block and the second prediction block, and a reconstructed block of the first block is determined based on the third prediction block and a residual block of the first block; The index information merge_idx of the first merge candidate and the index information merge_2nd_idx of the second merge candidate are obtained by analyzing a bitstream, and if the value of the index information merge_2nd_idx is greater than or equal to the value of the index information merge_idx, the value of the index of the second merge candidate is obtained by adding 1 to the value of the index information merge_2nd_idx. The video decoding method.
2. the inter-region motion information table includes inter-region merge candidates derived based on motion information of blocks decoded before the first block.
2. The video decoding method of claim 1.
3. If the first block is included in the merge processing area, adding a temporary merge candidate derived based on the motion information of the first block to a temporary motion information table; When decoding of all blocks included in the merge processing region is completed, the temporary merge candidate is updated in the inter-region motion information table.
2. The video decoding method of claim 1.
4. If it is determined that a merge candidate identical to the one inter-region merge candidate exists, the one inter-region merge candidate is not added to the merge candidate list; determining whether to add another inter-region merge candidate included in the inter-region motion information table to the merge candidate list based on a determination result as to whether the other inter-region merge candidate is the same as at least one merge candidate included in the merge candidate list; a determination is not made as to whether the other inter-region merge candidate and the same merge candidate as the one inter-region merge candidate are identical; 2. The video decoding method of claim 1.
5. If an inter region merge candidate having the same motion information as the first block exists in the merge candidate list, an index assigned to the inter region merge candidate in the inter region motion information table is updated to a maximum value.
2. The video decoding method of claim 1.
6. performing the determination by comparing the one inter-region merge candidate with at least one merge candidate having an index value greater than a threshold.
2. The video decoding method of claim 1.
7. the determining is performed by comparing one inter-region merge candidate with a merge candidate derived from a block at a specific location, the specific location including at least one of an upper right neighboring block or a lower left neighboring block of the first block.
2. The video decoding method of claim 1.
8. The third predicted block is generated based on a weighted sum operation of the first predicted block and the second predicted block.
2. The video decoding method of claim 1.
9. 1. A video encoding method, comprising: generating a merge candidate list for the first block; selecting one of the merge candidates included in the merge candidate list; performing motion compensation for the first block based on motion information of the selected merging candidate; performing motion compensation on the first block includes performing motion compensated prediction of the first block using a plurality of merging candidates; the plurality of merging candidates include a first merging candidate and a second merging candidate, the first merging candidate and the second merging candidate being included in a merging candidate list for the first block; a first prediction block is generated using the first merging candidate, a second prediction block is generated using the second merging candidate, a third prediction block is generated based on the first prediction block and the second prediction block, and a residual block of the first block is determined based on the third prediction block; information for identifying the index information merge_idx of the first merging candidate and the index information merge_2nd_idx of the second merging candidate are signaled via a bitstream, If the value of the index information merge_2nd_idx is greater than or equal to the value of the index information merge_idx, the value of the index of the second merging candidate is obtained by adding 1 to the value of the index information merge_2nd_idx; The video encoding method.
10. the inter-region motion information table includes inter-region merge candidates derived based on motion information of blocks coded before the first block. The video encoding method according to claim 9.
11. If the first block is included in the merge processing area, adding a temporary merge candidate derived based on the motion information of the first block to a temporary motion information table; When encoding of all blocks included in the merge processing region is completed, the temporary merge candidate is updated in the inter-region motion information table. The video encoding method according to claim 9.
12. If it is determined that a merge candidate identical to the one inter-region merge candidate exists, the one inter-region merge candidate is not added to the merge candidate list; determining whether to add another inter-region merge candidate included in the inter-region motion information table to the merge candidate list based on a determination result as to whether the other inter-region merge candidate is the same as at least one merge candidate included in the merge candidate list; a determination is not made as to whether the other inter-region merge candidate and the same merge candidate as the one inter-region merge candidate are identical; The video encoding method according to claim 9.
13. If an inter region merge candidate having the same motion information as the first block exists in the merge candidate list, an index assigned to the inter region merge candidate in the inter region motion information table is updated to a maximum value. The video encoding method according to claim 9.
14. performing the determination by comparing the one inter-region merge candidate with at least one merge candidate having an index value greater than a threshold. The video encoding method according to claim 9.
15. the determining is performed by comparing one inter-region merge candidate with a merge candidate derived from a block at a specific location, the specific location including at least one of an upper right neighboring block or a lower left neighboring block of the first block. The video encoding method according to claim 9.
16. The third predicted block is generated based on a weighted sum operation of the first predicted block and the second predicted block. The video encoding method according to claim 9.
17. A video decoding device comprising a memory and a processor, the memory is configured to store a processor-executable computer program; The video decoding device, wherein the processor is configured to execute a computer program to perform the video decoding method of any one of claims 1 to 8.
18. A video encoding device comprising a memory and a processor, the memory is configured to store a processor-executable computer program; The video encoding device, wherein the processor is configured to execute a computer program to perform the video encoding method of any one of claims 9 to 16.
19. A bitstream transmission method, comprising: executing the video encoding method according to claim 9 to generate a bitstream; and transmitting the bitstream.
Citation Information
Patent Citations
Multiple History-Based Non-Adjacent MVP for Wavefront Processing in Video Coding
JP2021530904A
Inter prediction method and device using history-based motion vectors
JP2021533681A
MULTIPLE HISTORY BASED NON-ADJACENT MVPs FOR WAVEFRONT PROCESSING OF VIDEO CODING
US20200021839A1
Inter prediction method and apparatus based on history-based motion vector
US20200186820A1