Method for coding or decoding video signal and apparatus used for the same

Adaptive signaling of offset vectors in video encoding and decoding refines motion vectors, addressing data volume challenges in high-definition video services by enhancing inter prediction efficiency.

JP2025186547APending Publication Date: 2025-12-23GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025166634
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-11-08
Filing Date
2025-10-02
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

The increasing data volume of high-definition video services poses a challenge due to the limitations of existing video compression standards like HEVC, necessitating improved methods for refining motion vectors during video signal encoding and decoding.

Method used

A method for signaling offset vectors during video signal encoding and decoding, where the magnitude and direction of the offset vector are determined adaptively based on index information, allowing for refined motion vectors through merge candidates.

Benefits of technology

This approach enhances inter prediction efficiency by improving the accuracy of motion vectors, thereby optimizing video compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025186547000001_ABST
    Figure 2025186547000001_ABST
Patent Text Reader

Abstract

To provide a method for coding or decoding and a video decoding apparatus which refine a motion vector derived from merge candidates on the basis of an offset vector when a video signal is coded or decoded.SOLUTION: A method for decoding a video includes the steps of: generating a merge candidate list of a current block; specifying a merge candidate of the current block from merge candidates in the merge candidate list; deriving an offset vector from the current block; and deriving a motion vector of the current block by adding a motion vector of the merge candidate to the offset vector.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a video signal encoding / decoding method and an apparatus used in said method. [Background technology]

[0002] As display panels continue to grow larger, video services with higher image quality are becoming increasingly necessary. The biggest problem with high-definition video services is the significant increase in data volume. To address this issue, active research is underway to improve video compression rates. As a representative example, in 2009, the Motion Picture Experts Group (MPEG) and the Video Coding Experts Group (VCEG) of the International Telecommunication Union-Telecommunication (ITU-T) established the Joint Collaborative Team on Video Coding (JCT-VC). JCT-VC proposed the video compression standard HEVC (High Efficiency Video Coding), which was approved on January 25, 2013. Its compression performance is approximately twice that of H.264 / AVC. With the rapid growth of high-definition video services, the performance limitations of HEVC are gradually becoming apparent. Summary of the Invention [Problem to be solved by the invention]

[0003] It is an object of the present invention to provide a method for refining motion vectors derived from merge candidates based on offset vectors when encoding / decoding a video signal, and an apparatus for implementing the method.

[0004] SUMMARY OF THE INVENTION It is an object of the present invention to provide a method for signaling offset vectors when encoding / decoding a video signal, and an apparatus for carrying out said method.

[0005] The technical problems that the present invention aims to achieve are not limited to the technical problems mentioned above, and a person having ordinary skill in the art to which the present invention pertains can clearly understand other technical problems not mentioned in the following description. [Means for solving the problem]

[0006] The video signal decoding / encoding method of the present invention includes the steps of generating a merge candidate list for a current block, specifying a merge candidate for the current block from among the merge candidates included in the merge candidate list, deriving an offset vector for the current block, and deriving a motion vector for the current block by adding the motion vector of the merge candidate to the offset vector.

[0007] In the video signal decoding / encoding method according to the present invention, the magnitude of the offset vector can be determined based on first index information specifying any one of motion magnitude candidates.

[0008] In the video signal decoding / encoding method according to the present invention, at least one of the maximum or minimum values ​​of the motion magnitude candidates can be set to be different depending on the value of a flag indicating the range of the motion magnitude candidates.

[0009] In the video signal decoding / encoding method according to the present invention, the flag can be signaled at the picture level.

[0010] In the video signal decoding / encoding method according to the present invention, at least one of the maximum and minimum values ​​of the motion magnitude candidates may be set differently depending on the motion vector accuracy of the current block.

[0011] In the video signal decoding / encoding method according to the present invention, the magnitude of the offset vector can be obtained by applying a shift operation to a value represented by the motion magnitude candidate designated by the first index information.

[0012] In the video signal decoding / encoding method according to the present invention, the direction of the offset vector can be determined based on second index information specifying any one of vector direction candidates.

[0013] The above brief summary of the present invention features are merely exemplary embodiments of the detailed description of the invention set forth below and are not intended to limit the scope of the invention. [Effects of the Invention]

[0014] According to the present invention, inter prediction efficiency can be improved by refining the motion vectors of merging candidates based on the offset vectors.

[0015] According to the present invention, the efficiency of inter prediction can be improved by adaptively determining the magnitude and direction of the offset vector.

[0016] The effects obtainable by the present invention are not limited to the above effects, and a person having ordinary skill in the art to which the present invention pertains can clearly understand other effects not mentioned in the following description. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a block diagram illustrating a video encoder according to an embodiment of the present invention; [Figure 2] FIG. 2 is a block diagram illustrating a video decoder according to an embodiment of the present invention. [Figure 3] FIG. 2 illustrates a basic coding tree unit according to an embodiment of the present invention. [Figure 4]FIG. 10 is a diagram illustrating multiple division types of a coding block. [Figure 5] FIG. 1 illustrates division of a coding tree unit. [Figure 6] 1 is a flowchart illustrating an inter-prediction method according to an embodiment of the present invention. [Figure 7] FIG. 1 illustrates non-linear motion of an object. [Figure 8] 1 is a flowchart illustrating an affine motion-based inter-prediction method according to an embodiment of the present invention. [Figure 9] 10A and 10B are diagrams illustrating examples of affine seed vectors for each affine motion model. [Figure 10] FIG. 10 is a diagram showing an example of affine vectors of sub-blocks in a four-parameter motion model. [Figure 11] 10 is a flowchart illustrating a process of deriving motion information for a current block in merge mode. [Figure 12] FIG. 10 is a diagram illustrating an example of candidate blocks for deriving merge candidates. [Figure 13] FIG. 10 shows the location of a reference sample. [Figure 14] FIG. 10 is a diagram illustrating an example of candidate blocks for deriving merge candidates. [Figure 15a] FIG. 10 is a diagram illustrating an example of fluctuation in the position of a reference sample. [Figure 15b] FIG. 10 is a diagram illustrating an example of fluctuation in the position of a reference sample. [Figure 16] FIG. 10 is a diagram illustrating an example of fluctuation in the position of a reference sample. [Figure 17] FIG. 10 is a diagram showing offset vectors based on the values ​​of distance_idx, which indicates the magnitude of the offset vector, and direction_idx, which indicates the direction of the offset vector. [Figure 18] FIG. 10 is a diagram showing offset vectors based on the values ​​of distance_idx, which indicates the magnitude of the offset vector, and direction_idx, which indicates the direction of the offset vector. DETAILED DESCRIPTION OF THE INVENTION

[0018] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0019] Video encoding and decoding is performed on a block-by-block basis, for example, encoding / decoding operations such as transform, quantization, prediction, loop filtering, or reconstruction can be performed on coding blocks, transform blocks, or prediction blocks.

[0020] Hereinafter, the block to be coded / decoded is referred to as a “current block.” For example, according to the current coding / decoding process step, the current block can represent a coding block, a transformation block, or a prediction block.

[0021] It should be noted that the term "unit" used in this specification may be understood to refer to a basic unit for performing a specific encoding / decoding process, and "block" may be understood to refer to a sample array of a predetermined size. Unless otherwise specified, "block" and "unit" are used interchangeably. For example, in the embodiments described below, a coding block and a coding unit may be understood to have the same meaning.

[0022] FIG. 1 is a block diagram illustrating a video encoder according to an embodiment of the present invention.

[0023] Referring to FIG. 1, the video encoding device 100 may include an image division unit 110, prediction units 120 and 125, a transformation unit 130, a quantization unit 135, a rearrangement unit 160, an entropy encoding unit 165, an inverse quantization unit 140, an inverse transform unit 145, a filter unit 150, and a memory 155.

[0024] 1 are shown individually to represent different characteristic functions in a video encoding device, but do not represent that each component is composed of separate hardware or a single software assembly. That is, for ease of explanation, each component may be arranged such that at least two of the components are combined to form a single component, or a single component is divided into multiple components that perform the function. Such embodiments in which each component is integrated and those in which each component is separated are also within the scope of the present invention, provided that they do not deviate from the essence of the present invention.

[0025] It should be noted that some components are not essential components for performing the essential functions of the present invention, but are only optional components for improving performance. The present invention may be implemented by including only the members necessary to realize the essence of the present invention, other than the components used only to improve performance, and a structure including only essential components other than optional components used only to improve performance also falls within the scope of the present invention.

[0026] The image division unit 110 can divide an input image into at least one processing unit. In this case, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The image division unit 110 divides one image into a plurality of combinations of coding units, prediction units, and transform units. The image can be coded by selecting a combination of coding units, prediction units, and transform units based on a predetermined criterion (e.g., a cost function).

[0027] For example, an image can be divided into multiple coding units. To divide an image into coding units, a recursive tree structure such as a quad tree structure can be used to divide a video or largest coding unit into other coding units, with the largest coding unit as the root. The other coding units may have child nodes whose number is equal to the number of divided coding units. Coding units that cannot be divided due to some restrictions become leaf nodes. That is, assuming that one coding unit can only realize square division, one coding unit can be divided into a maximum of four other coding units.

[0028] Hereinafter, in the embodiments of the present invention, the encoding unit may refer to a unit that performs encoding, or may refer to a unit that performs decoding.

[0029] The prediction units in one coding unit may be divided into shapes such as at least one of squares or rectangles of the same size, and one prediction unit in one coding unit may also be divided into prediction units of different shapes and / or sizes than another prediction unit.

[0030] Performing intra prediction based on a coding unit If the prediction unit is not the smallest coding unit, intra prediction can be performed without the need to divide it into multiple NxN prediction units.

[0031] The prediction units 120 and 125 may include an inter prediction unit 120 that performs inter prediction and an intra prediction unit 125 that performs intra prediction. It may be determined whether to use inter prediction or intra prediction for a prediction unit, and specific information (e.g., intra prediction mode, motion vector, reference image, etc.) may be determined based on each prediction method. In this case, the processing unit that performs the prediction may be different from the processing unit that determines the prediction method and specific content. For example, the prediction method and prediction mode may be determined by the prediction unit, and the prediction may be performed by a transform unit. A residual value (residual block) between the generated prediction block and the original block may be input to the transform unit 130. Prediction mode information, motion vector information, etc. for prediction may be coded together with the residual value by the entropy coding unit 165 and transmitted to the decoder. When a specific coding mode is used, the original block may be directly coded and transmitted to the decoder without generating a prediction block by the prediction units 120 and 125.

[0032] The inter prediction unit 120 may predict a prediction unit based on information of at least one image immediately before or after the current image. In some cases, the prediction unit may also be predicted based on information of a coded portion of the current image. The inter prediction unit 120 may include a reference image interpolation unit, a motion prediction unit, and a motion compensation unit.

[0033] The reference image interpolation unit receives reference image information from memory 155 and can generate pixel information of integer pixels or less from the reference image. For luminance pixels, an 8-tap DCT-based interpolation filter with different filter coefficients can be used to generate pixel information of integer pixels or less in units of 1 / 4 pixels. For chrominance signals, a 4-tap DCT-based interpolation filter with different filter coefficients can be used to generate pixel information of integer pixels or less in units of 1 / 8 pixels.

[0034] The motion prediction unit can perform motion prediction based on the reference image interpolated by the reference image interpolation unit. A number of methods can be used to calculate a motion vector, such as a full search-based block matching algorithm (FBMA), a three-step search (TSS), and a new three-step search algorithm (NTS). According to the interpolated pixels, the motion vector may have a motion vector value in units of half pixels or quarter pixels. The motion prediction unit can predict the current prediction unit using different motion prediction methods. A number of motion prediction methods can be used, such as a skip method, a merge method, an advanced motion vector prediction (AMVP), and an intra block copy method.

[0035] The intra prediction unit 125 can generate a prediction unit based on neighboring reference pixel information of the current block, which is pixel information in the current image. If the neighboring block of the current prediction unit is a block on which inter prediction has been performed and the reference pixel is a pixel on which inter prediction has been performed, the reference pixel included in the block on which inter prediction has been performed can be used as reference pixel information of the neighboring block on which intra prediction has been performed. In other words, if the reference pixel is unavailable, at least one reference pixel from the available reference pixels can be used instead of the unavailable reference pixel information.

[0036] In intra prediction, prediction modes may include an angular prediction mode that uses reference pixel information based on the prediction direction and a non-angular mode that does not use direction information when performing prediction. The mode for predicting luma information may be different from the mode for predicting chroma information. To predict chroma information, intra prediction mode information for predicting luma information or predicted luma signal information may be used.

[0037] When performing intra prediction, if the size of the prediction unit is the same as the size of the transform unit, intra prediction can be performed on the prediction unit based on the pixel located to the left, the pixel located to the upper left, and the pixel located above.However, when performing intra prediction, if the size of the prediction unit is different from the size of the transform unit, intra prediction can be performed based on the reference pixel of the transform unit.Note that intra prediction using NxN division can only be applied to the smallest coding unit.

[0038] After applying an adaptive intra smoothing (AIS) filter to reference pixels based on the prediction mode, a predicted block can be generated using an intra prediction method. The type of adaptive intra smoothing filter applied to the reference pixels can vary. To perform the intra prediction method, the intra prediction mode of a current prediction unit can be predicted based on the intra prediction mode of prediction units located around the current prediction unit. When predicting the prediction mode of the current prediction unit using mode information predicted from surrounding prediction units, if the intra prediction mode of the current prediction unit is the same as the intra prediction mode of the surrounding prediction units, information indicating that the prediction mode of the current prediction unit is the same as that of the surrounding prediction units can be transmitted using predetermined flag information. If the prediction mode of the current prediction unit is different from the prediction mode of the surrounding prediction units, entropy coding can be performed to encode the prediction mode information of the current block.

[0039] In addition, a residual block including residual value information, which is a difference value between a prediction unit that performs prediction based on the prediction unit generated by the prediction units 120 and 125 and the original block of the prediction unit, can be generated. The generated residual block can be input to the conversion unit 130.

[0040] The transform unit 130 may perform transforms on the original block and the residual block containing residual value information between the prediction units generated by the prediction units 120 and 125 using a transform method such as a discrete cosine transform (DCT) or a discrete sine transform (DST). Here, the DCT transform kernel includes at least one of DCT2 or DCT8, and the DSTc includes DST7. Whether to apply DCT or DST to transform the residual block may be determined based on intra-prediction mode information of the prediction unit used to generate the residual block. Transformation of the residual block may also be skipped. A flag indicating whether to skip transformation of the residual block may be coded. Transformation may be skipped for residual blocks whose magnitude is equal to or less than a threshold, luma component, or chroma component (4:4:4 format or less).

[0041] The quantization unit 135 may quantize the values ​​transformed into the frequency domain by the transformation unit 130. The quantization coefficient may vary depending on the importance of the block or image. The values ​​calculated by the quantization unit 135 may be provided to the inverse quantization unit 140 and the rearrangement unit 160.

[0042] The rearrangement unit 160 may perform rearrangement of coefficient values ​​for the quantized residual values.

[0043] The rearrangement unit 160 may convert two-dimensional block shape coefficients into one-dimensional vector form using a coefficient scanning method. For example, the rearrangement unit 160 may scan DC coefficients or high-frequency region coefficients using a zig-zag scan method and convert them into one-dimensional vector form. Depending on the size of the transform unit and the intra prediction mode, vertical scanning, which scans two-dimensional block shape coefficients along the column direction, and horizontal scanning, which scans two-dimensional block shape coefficients along the row direction, may be used instead of zig-zag scanning. That is, whether to use zig-zag scanning, vertical scanning, or horizontal scanning may be determined based on the size of the transform unit and the intra prediction mode.

[0044] The entropy coding unit 165 can perform entropy coding based on the values ​​calculated by the rearrangement unit 160. For example, the entropy coding can use a plurality of coding methods such as Exponential Golomb coding, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC).

[0045] The entropy coding unit 165 can encode multiple information such as residual value coefficient information and block type information of the coding unit from the rearrangement unit 160 and the prediction units 120 and 125, prediction mode information, division unit information, prediction unit information and transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information.

[0046] The entropy coding unit 165 can perform entropy coding on the coefficient values ​​of the coding unit input from the rearrangement unit 160 .

[0047] The inverse quantization unit 140 and the inverse transform unit 145 perform inverse quantization on the multiple values ​​quantized by the quantization unit 135, and perform inverse transform on the values ​​transformed by the transform unit 130. A reconstructed block can be generated by merging the residual values ​​generated by the inverse quantization unit 140 and the inverse transform unit 145 with the prediction units predicted by the motion prediction unit, motion compensation unit, and intra prediction unit included in the prediction units 120 and 125.

[0048] The filter unit 150 may include at least one of a deblocking filter, an offset correction unit, and an adaptive loop filter (ALF).

[0049] A deblocking filter can remove block artifacts generated in a reconstructed image due to boundaries between blocks. To determine whether to perform deblocking, it can be determined whether to apply a deblocking filter to a current block based on pixels included in several columns or rows included in the block. When applying a deblocking filter to a block, a strong filter or a weak filter can be applied based on the required deblocking filtering strength. In addition, when performing vertical filtering or horizontal filtering in the process of using a deblocking filter, horizontal filtering and vertical filtering can be performed synchronously.

[0050] The offset correction unit can correct the offset between the deblocked video and the original video on a pixel-by-pixel basis. The offset correction can be performed on a specified image in the following manner: After dividing the pixels included in the video into a predetermined number of regions, the regions that require offset execution are determined, and an offset is applied to the corresponding region or an offset is applied taking into account edge information of each pixel.

[0051] Adaptive loop filtering (ALF) can be performed based on the comparison value between the filtered reconstructed image and the original video. After dividing the pixels in the video into predetermined groups, a filter to be used for each group is determined, allowing differential filtering to be performed for each group. Information regarding whether adaptive loop filtering is to be applied and luminance information can be transmitted according to the coding unit (CU). The shape and filter coefficients of the adaptive loop filter to be applied vary depending on the block. It is also possible to apply the same type (constant type) of adaptive loop filter regardless of the characteristics of the block to which it is applied.

[0052] The memory 155 can store the reconstructed blocks or images calculated by the filter unit 150, and can provide the stored reconstructed blocks or images to the prediction units 120 and 125 when performing inter prediction.

[0053] FIG. 2 is a block diagram illustrating a video decoder according to an embodiment of the present invention.

[0054] Referring to FIG. 2, the video decoder 200 may include an entropy decoding unit 210, a rearrangement unit 215, an inverse quantization unit 220, an inverse transform unit 225, a prediction unit 230, a prediction unit 235, a filter unit 240, and a memory 245.

[0055] When a video bitstream is input from a video encoder, the input bitstream can be decoded in steps that are the reverse of those taken by the video encoder.

[0056] The entropy decoding unit 210 may perform entropy decoding in a step that is the reverse of the step of entropy encoding performed by the entropy encoding unit of the video encoder. For example, to correspond to the method performed by the video encoder, multiple methods such as Exponential Golomb coding, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) may be applied.

[0057] The entropy decoding unit 210 can decode information related to intra-prediction and inter-prediction performed by the encoder.

[0058] The rearrangement unit 215 may perform rearrangement by rearranging a bitstream entropy decoded by the entropy decoding unit 210 in the encoding unit. The rearrangement may be performed by reconstructing a plurality of coefficients represented in a one-dimensional vector form into coefficients in a two-dimensional block shape. The rearrangement unit 215 may perform rearrangement by receiving information related to coefficient scanning performed by the encoding unit and performing reverse scanning according to the scanning order performed by the corresponding encoding unit.

[0059] The inverse quantization unit 220 may perform inverse quantization based on the quantization parameter provided by the encoder and the coefficient values ​​of the rearranged block.

[0060] The inverse transform unit 225 may perform an inverse discrete cosine transform or an inverse discrete sine transform on the quantization result performed by the video encoder. The inverse discrete cosine transform or the inverse discrete sine transform is the inverse transform of the transform performed by the transform unit, i.e., the inverse transform of the discrete cosine transform or the discrete sine transform. Here, the DCT transform kernel may include at least one of a DCT2 or a DCT8, and the DST transform kernel may include a DST7. Alternatively, if the video encoder skips a transform, the inverse transform unit 225 may not perform the inverse transform. The inverse transform may be performed using a transmission unit determined by the video encoder. The inverse transform unit 225 of the video decoder may selectively perform a transform method (e.g., DCT or DST) based on multiple pieces of information such as a prediction method, a size of a current block, and a prediction direction.

[0061] The prediction units 230, 235 can generate the prediction blocks based on information related to the generation of the prediction blocks provided by the entropy decoding unit 210 and previously decoded block or image information provided by the memory 245.

[0062] As described above, when intra prediction is performed in the same manner as the operation manner in a video encoder, if the size of the prediction unit is the same as the size of the transform unit, intra prediction is performed on the prediction unit based on the pixel located to the left, the pixel located to the upper left, and the pixel located above the prediction unit. When performing intra prediction, if the size of the prediction unit is different from the size of the transform unit, intra prediction can be performed based on reference pixels of the transform unit. Note that intra prediction using NxN division can be applied only to the smallest coding unit.

[0063] The prediction units 230 and 235 may include a prediction unit determination unit, an inter prediction unit, and an intra prediction unit. The prediction unit determination unit receives multiple pieces of information, such as prediction unit information input from the entropy decoding unit 210, prediction mode information for the intra prediction method, and motion prediction-related information for the inter prediction method, and classifies the prediction unit based on the current encoding unit and determines whether the prediction unit performs inter prediction or intra prediction. The inter prediction unit 230 may perform inter prediction on the current prediction unit based on information included in at least one image immediately before or immediately after the current image to which the current prediction unit belongs, using information necessary for performing inter prediction on the current prediction unit provided by the video encoder. Alternatively, inter prediction may be performed based on information on a reconstructed region of the current image to which the current prediction unit belongs.

[0064] To perform inter prediction, it is possible to determine, based on the coding unit, whether the motion prediction method of the prediction unit included in the corresponding coding unit is skip mode, merge mode, advanced motion vector prediction mode (AMVP mode), or intra block duplication mode.

[0065] The intra prediction unit 235 may generate a prediction block based on pixel information in the current image. If the prediction unit is a prediction unit for which intra prediction has been performed, the intra prediction may be performed based on intra prediction mode information of the prediction unit provided from the video encoder. The intra prediction unit 235 may include an adaptive intra smoothing (AIS) filter, a reference pixel interpolator, and a DC filter. The adaptive intra smoothing filter is a part that performs filtering on reference pixels of the current block and may determine whether to apply a filter based on the prediction mode of the current prediction unit. Adaptive intra smoothing filtering may be performed on reference pixels of the current block using the prediction mode of the prediction unit and adaptive intra smoothing filter information provided from the video encoder. If the prediction mode of the current block is a mode that does not perform adaptive intra smoothing filtering, the adaptive intra smoothing filter may not be applied.

[0066] Regarding the reference pixel interpolation unit, if the prediction mode of the prediction unit is a prediction unit that performs intra prediction based on pixel values ​​to be interpolated for reference pixels, reference pixels of integer or smaller pixel units can be generated by interpolating the reference pixels. If the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixels, interpolation of the reference pixels is not necessary. If the prediction mode of the current block is DC mode, the DC filter can generate a prediction block by filtering.

[0067] The reconstructed block or image may be provided to a filter unit 240. The filter unit 240 may include a deblocking filter, an offset correction unit, and an ALF.

[0068] The video decoder may receive, from the video encoder, information regarding whether to apply a deblocking filter to a corresponding block or image and, when applying the deblocking filter, information regarding whether to use a strong filter or a weak filter. The deblocking filter of the video decoder may receive the information regarding the deblocking filter provided by the video encoder, and the video decoder may perform deblocking filtering on the corresponding block.

[0069] The offset correction unit can perform offset correction on the reconstructed image based on the offset information and the type of offset correction applied to the video when encoding.

[0070] ALF can be applied to a coding unit based on information provided by the encoder regarding whether to apply ALF, ALF coefficient information, etc. Such ALF information may be provided by being included in a particular parameter set.

[0071] The memory 245 stores the reconstructed image or block, makes it available as a reference image or block, and is capable of providing the reconstructed image to an output.

[0072] FIG. 3 is a diagram illustrating a basic coding tree unit according to an embodiment of the present invention.

[0073] The largest coding block can be defined as a coding tree block. An image may be divided into multiple coding tree units (each coding tree unit has a size of Coding Tree Unit: CTU). The largest coding unit may be called the Largest Coding Unit (LCU). Figure 3 shows an example of dividing an image into multiple coding tree units.

[0074] The size of the coding tree unit may be defined at the picture level or the sequence level, so that the picture parameter set or the sequence parameter set can be used to signal information indicating the size of the coding tree unit.

[0075] For example, the coding tree unit size for all images in a sequence may be 128 x 128. Alternatively, the coding tree unit size may be determined to be either 128 x 128 or 256 x 256 at the image level. For example, the coding tree unit size for a first image may be 128 x 128, and the coding tree unit size for a second image may be 256 x 256.

[0076] Coding blocks can be generated by dividing the coding tree units. The coding blocks represent basic units for performing encoding / decoding processes. For example, prediction or transformation can be performed according to different coding blocks, or a predictive coding mode can be determined according to different coding blocks. Here, the predictive coding mode represents a method for generating a predicted image. For example, the predictive coding mode may include intra prediction (intra prediction), inter prediction (inter prediction), current picture referencing (CPR), intra block copy (IBC), or combined prediction. For a coding block, a predictive block related to the coding block can be generated using at least one predictive coding mode from intra prediction, inter prediction, current picture referencing, or combined prediction.

[0077] Information indicating the predictive coding mode of the current block may be signaled via the bitstream. For example, the information may be a one-bit flag indicating whether the predictive coding mode is intra mode or inter mode. Only when it is determined that the predictive coding mode of the current block is inter mode, can current picture reference or combined prediction be used.

[0078] The current image reference is used to obtain a prediction block of the current block from an encoded / decoded region of the current image using the current image as a reference image. Here, the current image refers to an image including the current block. Information indicating whether the current image reference is applied to the current block may be signaled via the bitstream. For example, the information may be a 1-bit flag. If the flag is true, the predictive coding mode of the current block may be determined as current image reference. If the flag is false, the prediction mode of the current block may be determined as inter prediction.

[0079] Alternatively, the predictive coding mode of the current block may be determined based on a reference image index. For example, if the reference image index points to the current image, the predictive coding mode of the current block may be determined as current image reference. If the reference image index points to an image other than the current image, the predictive coding mode of the current block may be determined as inter prediction. That is, current image reference is a prediction method that uses information on an encoded / decoded region in the current image, and inter prediction is a prediction method that uses information on another encoded / decoded image.

[0080] Combined prediction is a coding mode formed by combining two or more of intra prediction, inter prediction, and current image reference. For example, when combined prediction is applied, a first predicted block may be generated based on one of intra prediction, inter prediction, or current image reference, and a second predicted block may be generated based on the other. When generating the first predicted block and the second predicted block, a final predicted block may be generated by averaging and weighted addition of the first predicted block and the second predicted block. Information indicating whether combined prediction is applied may be signaled via a bitstream. The information may be a 1-bit flag.

[0081] FIG. 4 is a diagram showing a plurality of division types of a coding block.

[0082] A coding block can be divided into a plurality of coding blocks based on quadtree division, binary tree division, or ternary tree division. Also, a divided coding block can be further divided into a plurality of coding blocks based on quadtree division, binary tree division, or ternary tree division.

[0083] Quadtree partitioning is a partitioning technique that divides the current block into four blocks. As a result of quadtree partitioning, the current block can be divided into four square partitions (see "SPLIT_QT" in Figure 4(a)).

[0084] Binary tree partitioning is a partitioning technique that divides a current block into two blocks. The process of dividing the current block into two blocks along the vertical direction (i.e., using a vertical line that crosses the current block) is called vertical binary tree partitioning, and the process of dividing the current block into two blocks along the horizontal direction (i.e., using a horizontal line that crosses the current block) is called horizontal binary tree partitioning. After binary tree partitioning, the current block can be divided into two non-square partitions. "SPLIT_BT_VER" in Figure 4(b) represents the vertical binary tree partitioning result, and "SPLIT_BT_HOR" in Figure 4(c) represents the horizontal binary tree partitioning result.

[0085] Ternary tree partitioning is a partitioning technique that divides a current block into three blocks. The process of dividing the current block into three blocks along the vertical direction (i.e., using two vertical lines crossing the current block) is called vertical ternary tree partitioning, and the process of dividing the current block into three blocks along the horizontal direction (i.e., using two horizontal lines crossing the current block) is called horizontal ternary tree partitioning. After ternary tree partitioning, the current block can be divided into three non-square partitions. In this case, the width / height of the partition located at the center of the current block may be twice the width / height of the other partitions. 'SPLIT_TT_VER' in Figure 4(d) represents the result of vertical ternary tree partitioning, and 'SPLIT_TT_HOR' in Figure 4(e) represents the result of horizontal ternary tree partitioning.

[0086] The number of times a coding tree unit is divided can be defined as the partitioning depth. The maximum partitioning depth of a coding tree unit can be determined at the sequence or image level. Therefore, the maximum partitioning depth of a coding tree unit may vary depending on the sequence or image.

[0087] Alternatively, a maximum partitioning depth can be determined independently for each of multiple partitioning techniques. For example, the maximum partitioning depth that allows for quadtree partitioning may be different from the maximum partitioning depth that allows for binary tree and / or ternary tree partitioning.

[0088] The encoder may signal information indicating at least one of the partition shape or partition depth of the current block via the bitstream, and the decoder may determine the partition shape and partition depth of the coding tree unit based on the information parsed from the bitstream.

[0089] FIG. 5 is a diagram illustrating an example of division of a coding tree unit.

[0090] The process of dividing a coding block using a division technique such as quad-tree division, binary tree division, and / or ternary tree division can be called multi-tree partitioning.

[0091] The coding blocks generated by applying multi-tree division to a coding block can be called multiple sub-coding blocks. If the division depth of the coding block is k, the division depth of the multiple sub-coding blocks is k+1.

[0092] On the other hand, among multiple coding blocks with a division depth of k+1, a coding block with a division depth of k can be called a higher coding block.

[0093] The division type of the currently coded block may be determined based on at least one of the division shape of the upper coded block or the division type of the adjacent coded block. Here, the adjacent coded block is adjacent to the currently coded block, and may include at least one of the upper adjacent block, the left adjacent block, or the adjacent block adjacent to the upper left corner of the currently coded block. Here, the division type may include at least one of whether to perform quadtree division, whether to perform binary tree division, the binary tree division direction, whether to perform ternary tree division, and the ternary tree division direction.

[0094] To determine the partitioning shape of a coding block, information indicating whether the coding block is split can be signaled via the bitstream. The information is a 1-bit flag "split_cu_flag", which indicates that the coding block is split using a multi-tree partitioning technique if the flag is true.

[0095] If 'split_cu_flag' is true, information indicating whether the coding block has been quadtree split can be signaled via the bitstream. The information is a 1-bit flag 'split_qt_flag', and if the flag is true, the coding block may be split into four blocks.

[0096] For example, in the example shown in Fig. 5, the coding tree unit is quadtree divided to generate four coding blocks with a division depth of 1. It is also shown that quadtree division is applied again to the first and fourth coding blocks of the four coding blocks generated as a result of the quadtree division. Ultimately, four coding blocks with a division depth of 2 can be generated.

[0097] Note that by applying quadtree division to an encoding block with a division depth of 2 again, an encoding block with a division depth of 3 can be generated.

[0098] If quadtree partitioning has not been applied to a coding block, it may be determined whether to perform binary tree partitioning or ternary tree partitioning on the coding block by considering at least one of the size of the coding block, whether the coding block is located on an image boundary, the maximum partition depth, or the partition shape of adjacent blocks. If it is determined to perform binary tree partitioning or ternary tree partitioning on the coding block, information indicating the partitioning direction may be signaled via a bitstream. The information may be a one-bit flag "mtt_split_cu_vertical_flag." Based on the flag, it may be determined whether the partitioning direction is vertical or horizontal. Note that information indicating whether binary tree partitioning or ternary tree partitioning is applied to the coding block may be signaled via a bitstream. The information may be a one-bit flag "mtt_split_cu_binary_flag." Based on the flag, it may be determined whether to apply binary tree partitioning or ternary tree partitioning to the coding block.

[0099] For example, in the example shown in Figure 5, vertical binary tree partitioning is applied to a coding block with a partitioning depth of 1. Of the coding blocks generated as a result of the partitioning, vertical ternary tree partitioning is applied to the left coding block, and vertical binary tree partitioning is applied to the right coding block.

[0100] Inter-prediction is a predictive coding mode that predicts a current block using information of a previous image. For example, a block in the previous image at the same position as the current block (hereinafter referred to as a collocated block) can be used as a prediction block for the current block. Hereinafter, a prediction block generated based on a block at the same position as the current block is referred to as a collocated prediction block.

[0101] On the other hand, if an object in the previous image moves to another position in the current image, the current block can be effectively predicted based on the object's motion. For example, if the direction and size of the object's movement can be determined by comparing the previous image with the current image, a predicted block (or predicted image) of the current block can be generated taking into account the object's motion information. Hereinafter, the predicted block generated using the motion information may be referred to as a motion predicted block.

[0102] A residual block can be generated by subtracting a predicted block from a current block. In this case, if target motion is present, using a motion prediction block instead of a co-located prediction block can reduce the energy of the residual block and improve the compression performance of the residual block.

[0103] As described above, the process of generating a prediction block using motion information may be called motion compensated prediction. In most inter predictions, a prediction block may be generated based on motion compensated prediction.

[0104] The motion information may include at least one of a motion vector, a reference image index, a prediction direction, or a bidirectional weight index. The motion vector represents the movement direction and size of an object. The reference image index specifies a reference image of the current block from multiple reference images included in a reference image list. The prediction direction indicates one of unidirectional L0 prediction, unidirectional L1 prediction, or bidirectional prediction (L0 prediction and L1 prediction). At least one of L0 direction motion information or L1 direction motion information can be used based on the prediction direction of the current block. The bidirectional weight index specifies a weight to be applied to the L0 prediction block and a weight to be applied to the L1 prediction block.

[0105] FIG. 6 is a flowchart illustrating an inter prediction method according to an embodiment of the present invention.

[0106] Referring to FIG. 6, the inter prediction method includes a step of determining an inter prediction mode of a current block (S601), a step of obtaining motion information of the current block based on the determined inter prediction mode (S602), and a step of performing motion compensation prediction on the current block based on the obtained motion information (S603).

[0107] Here, the inter prediction mode represents a number of techniques for determining motion information of the current block, and may include an inter prediction mode using translation motion information and an inter prediction mode using affine motion information. For example, the inter prediction mode using translation motion information may include a merge mode and an advanced motion vector prediction mode. The inter prediction mode using affine motion information may include an affine merge mode and an affine motion vector prediction mode. According to the inter prediction mode, the motion information of the current block may be determined based on information analyzed from a neighboring block or a bitstream adjacent to the current block.

[0108] The inter prediction method using affine motion information will be described in detail below.

[0109] FIG. 7 is a diagram illustrating the nonlinear motion of an object.

[0110] The motion of an object in a video may be nonlinear. For example, as shown in the example of FIG. 7, nonlinear motion of the object may occur due to camera zoom-in, zoom-out, rotation, or affine transformation. When nonlinear motion of the object occurs, the motion of the object cannot be effectively represented by a translational motion vector. Therefore, in the part where nonlinear motion of the object occurs, affine motion is used instead of translational motion to improve coding efficiency.

[0111] FIG. 8 is a flowchart illustrating an affine motion-based inter prediction method according to an embodiment of the present invention.

[0112] Whether to apply an affine motion-based inter prediction technique to the current block may be determined based on information analyzed from the bitstream. Specifically, whether to apply an affine motion-based inter prediction technique to the current block may be determined based on at least one of a flag indicating whether to apply an affine merge mode to the current block or a flag indicating whether to apply an affine motion vector prediction mode to the current block.

[0113] When applying an affine motion-based inter prediction technique to a current block, an affine motion model of the current block can be determined (S801). The affine motion model may be determined by at least one of a six-parameter affine motion model or a four-parameter affine motion model. The six-parameter affine motion model represents affine motion using six parameters, and the four-parameter affine motion model represents affine motion using four parameters.

[0114] Equation 1 is a case where affine motion is expressed by six parameters. The affine motion represents the translational motion of a predetermined region determined by an affine seed vector.

number

[0115] When affine motion is represented using six parameters, complex motion can be represented, but the number of bits required for encoding each parameter increases, reducing coding efficiency. Therefore, affine motion can also be represented using four parameters. Equation 2 shows the case where affine motion is represented using four parameters.

number

[0116] Information for determining an affine motion model for a current block may be coded and signaled via a bitstream. For example, the information may be a 1-bit flag "affine_type_flag." A value of 0 of the flag indicates that a 4-parameter affine motion model is to be applied. A value of 1 of the flag indicates that a 6-parameter affine motion model is to be applied. The flag may be coded in units of a slice, an image block, or a block (e.g., a coding block or a coding tree unit). When a flag is transmitted using a signal at the slice level, the affine motion model determined at the slice level may be applied to all blocks belonging to the slice.

[0117] Alternatively, the affine motion model of the current block may be determined based on the affine inter prediction mode of the current block. For example, when the affine merge mode is applied, the affine motion model of the current block may be determined as a four-parameter motion model. On the other hand, when the affine motion vector prediction mode is applied, information for determining the affine motion model of the current block may be coded and signaled via a bitstream. For example, when the affine motion vector prediction mode is applied to the current block, the affine motion model of the current block may be determined based on a 1-bit flag "affine_type_flag."

[0118] Next, an affine seed vector for the current block can be derived (S802). If a four-parameter affine motion model is selected, motion vectors can be derived at two control points of the current block. If a six-parameter affine motion model is selected, motion vectors can be derived at three control points of the current block. The motion vectors at the control points can be called affine seed vectors. The control points may include at least one of the top left corner, the top right corner, or the bottom left corner of the current block.

[0119] FIG. 9 is a diagram showing examples of affine seed vectors for each affine motion model.

[0120] In a four-parameter affine motion model, an affine seed vector relating to two of the upper left corner, the upper right corner, or the lower left corner can be derived. For example, as shown in (a) of FIG. 9, when a four-parameter affine motion model is selected, an affine vector can be derived using an affine seed vector sv0 relating to the upper left corner of the current block (e.g., the upper left sample (x0, y0)) and an affine seed vector sv1 relating to the upper right corner of the current block (e.g., the upper right sample (x1, y1)). Also, instead of the affine seed vector relating to the upper left corner, an affine seed vector relating to the lower left corner can be used. Alternatively, instead of the affine seed vector relating to the upper right corner, an affine seed vector relating to the lower left corner can be used.

[0121] In a six-parameter affine motion model, affine seed vectors relating to the upper left corner, the upper right corner, and the lower left corner can be derived. For example, as shown in the example of FIG. 9(b), when a six-parameter affine motion model is selected, an affine vector can be derived using an affine seed vector sv0 relating to the upper left corner of the current block (e.g., the upper left sample (x0, y0)), an affine seed vector sv1 relating to the upper right corner of the current block (e.g., the upper right sample (x1, y1)), and an affine seed vector sv2 relating to the upper left corner of the current block (e.g., the upper left sample (x2, y2)).

[0122] In the embodiments described below, in a four-parameter affine motion model, the affine seed vectors of the top-left control point and the top-right control point are referred to as the first affine seed vector and the second affine seed vector, respectively. In the embodiments described below that use the first affine seed vector and the second affine seed vector, at least one of the first affine seed vector and the second affine seed vector can be replaced with the affine seed vector of the bottom-left control point (third affine seed vector) or the affine seed vector of the bottom-right control point (fourth affine seed vector).

[0123] In the six-parameter affine motion model, the affine seed vectors of the top-left control point, the top-right control point, and the bottom-left control point are referred to as the first affine seed vector, the second affine seed vector, and the third affine seed vector, respectively. In an embodiment described below in which the first affine seed vector, the second affine seed vector, and the third affine seed vector are used, at least one of the first affine seed vector, the second affine seed vector, and the third affine seed vector can be replaced with the affine seed vector of the bottom-right control point (fourth affine seed vector).

[0124] The affine seed vector can be used to derive affine vectors according to different sub-blocks (S803). Here, the affine vectors represent translational motion vectors derived based on the affine seed vectors. The affine vectors of sub-blocks can be called affine sub-block motion vectors or sub-block motion vectors.

[0125] FIG. 10 is a diagram showing an example of affine vectors of sub-blocks in a four-parameter motion model.

[0126] The affine vector of the sub-block can be derived based on the position of the control point, the position of the sub-block, and the affine seed vector. For example, Equation 3 shows an example of derivation of the affine sub-block vector.

number

[0127] In Equation 3, (x, y) represents the position of a sub-block. Here, the position of the sub-block represents the position of a reference sample included in the sub-block. The reference sample may be a sample located in the upper left corner of the sub-block, or at least one sample located at the center in the x-axis or y-axis coordinate. (x0, y0) represents the position of the first control point, and (sv0x, sv0y) represent the first affine seed vector. Note that (x1, y1) represents the position of the second control point, and (sv1x, sv1y) represent the second affine seed vector.

[0128] If the first and second control points correspond to the upper left and upper right corners of the current block, respectively, then x1-x0 can be set to a value equal to the width of the current block.

[0129] Next, motion compensation prediction can be performed on each sub-block using the affine vector of each sub-block (S804). After performing motion compensation prediction, a prediction block for each sub-block can be generated. The prediction block of the sub-block can be set as the prediction block of the current block.

[0130] Next, the inter prediction method using translational motion information will be described in detail.

[0131] The motion information of the current block may be derived from the motion information of another block of the current block. Here, the other block may be a block that is preferentially encoded / decoded using inter prediction rather than the current block. A merge mode may be defined as a case where the motion information of the current block is the same as the motion information of another block. Furthermore, a motion vector prediction mode may be defined as a case where the motion vector of another block is set as a predicted value of the motion vector of the current block.

[0132] FIG. 11 is a flowchart illustrating the process of deriving motion information for a current block in merge mode.

[0133] Merge candidates for the current block can be derived (S1101). The merge candidates for the current block are derived from a block that precedes the current block and is coded / decoded using inter prediction.

[0134] FIG. 12 is a diagram showing an example of candidate blocks for deriving merge candidates.

[0135] The candidate blocks may include at least one of adjacent blocks including samples adjacent to the current block or non-adjacent blocks including samples not adjacent to the current block. Hereinafter, samples for determining candidate blocks are referred to as reference samples. Note that reference samples adjacent to the current block are referred to as adjacent reference samples, and reference samples not adjacent to the current block are referred to as non-adjacent reference samples.

[0136] The neighboring reference sample may be included in the column adjacent to the leftmost column of the current block or the row adjacent to the topmost row of the current block. For example, if the coordinates of the top-left sample of the current block are (0,0), one of the block including the reference sample at the (-1,H-1) position, the block including the reference sample at the (W-1,-1) position, the block including the reference sample at the (W,-1) position, the block including the reference sample at the (-1,H) position, or the block including the reference sample at the (-1,-1) position may be used as a candidate block. As shown in the drawing, neighboring blocks with indexes 0 to 4 can be used as candidate blocks.

[0137] The non-adjacent reference sample represents a sample in which at least one of the x-axis distance or y-axis distance from the reference sample adjacent to the current block has a predetermined value. For example, one of the following can be used as a candidate block: a block including a reference sample whose x-axis distance from the left reference sample is a predetermined value; a block including a non-adjacent sample whose y-axis distance from the above reference sample is a predetermined value; or a block including a non-adjacent sample whose x-axis distance and y-axis distance from the above-left reference sample are predetermined values. The predetermined value may be an integer such as 4, 8, 12, 16, etc. As shown in the drawing, at least one of the blocks with indexes 5 to 26 can be used as a candidate block.

[0138] A non-adjacent reference sample can be a sample that is not located on the same vertical, horizontal, or diagonal line as an adjacent reference sample.

[0139] FIG. 13 is a diagram showing the positions of the reference samples.

[0140] 13, the x-coordinate of the upper non-adjacent reference sample may be set to be different from the x-coordinate of the upper adjacent reference sample. For example, when the position of the upper adjacent reference sample is (W-1, -1), the position of the upper non-adjacent reference sample that is N away from the upper adjacent reference sample along the y-axis may be set to ((W / 2)-1, -1-N), and the position of the upper non-adjacent reference sample that is 2N away from the upper adjacent reference sample along the y-axis may be set to (0, -1-2N). In other words, the position of the non-adjacent reference sample may be determined based on the position of the adjacent reference sample and the distance from the adjacent reference sample.

[0141] Hereinafter, among the candidate blocks, a candidate block that includes adjacent reference samples will be referred to as an adjacent block, and a block that includes non-adjacent reference samples will be referred to as a non-adjacent block.

[0142] If the distance between the current block and a candidate block is equal to or greater than a threshold, the candidate block may be set as unavailable as a merging candidate. The threshold may be determined based on the size of the coding tree unit. For example, the threshold may be the height of the coding tree unit (ctu_height) or a value obtained by adding or subtracting an offset value to the height of the coding tree unit (e.g., ctu_height ± N). The offset value N is a value predefined in the encoder and decoder, and may be 4, 8, 16, 32, or ctu_height.

[0143] If the difference value between the y-axis coordinate of the current block and the y-axis coordinate of the sample included in the candidate block is greater than a threshold, the candidate block can be determined to be unavailable as a merging candidate.

[0144] Alternatively, a candidate block that does not belong to the same coding tree unit as the current block can be set as unavailable as a merging candidate. For example, if a reference sample exceeds the upper boundary of the coding tree unit to which the current block belongs, the candidate block containing the reference sample can be set as unavailable as a merging candidate.

[0145] When the upper boundary of the current block is adjacent to the upper boundary of a coding tree unit, if multiple candidate blocks are set as unavailable as merge candidates, the encoding / decoding efficiency of the current block will be reduced. To solve this problem, the number of candidate blocks located above the current block can be made greater than the number of candidate blocks located to the left of the current block by setting the candidate blocks.

[0146] FIG. 14 is a diagram showing an example of candidate blocks for deriving merge candidates.

[0147] 14, the candidate blocks can be an upper block belonging to a row of N blocks above the current block and a left block belonging to a row of M blocks to the left of the current block. In this case, by making M larger than N, the number of left candidate blocks can be made larger than the number of above candidate blocks.

[0148] For example, the difference between the y-axis coordinate of a reference sample in the current block and the y-axis coordinate of a block above that can be used as a candidate block can be set to N times the height of the current block or less, and the difference between the x-axis coordinate of a reference sample in the current block and the x-axis coordinate of a block to the left that can be used as a candidate block can be set to M times the width of the current block or less.

[0149] For example, in the example shown in FIG. 14, the blocks belonging to the two block columns above the current block and the blocks belonging to the five block columns to the left of the current block are set as candidate blocks.

[0150] As another example, if a candidate block does not belong to the same coding tree unit as the current block, a merge candidate can be derived instead of the candidate block using a block that belongs to the same coding tree unit as the current block or a block that contains a reference sample adjacent to the boundary of the coding tree unit.

[0151] FIG. 15 is a diagram showing an example of fluctuation in the position of the reference sample.

[0152] If a reference sample is included in a coding tree unit different from the current block and the reference sample is not adjacent to the boundary of the coding tree unit, a reference sample adjacent to the boundary of the coding tree unit can be used instead of the reference sample to determine a candidate block reference sample.

[0153] For example, in the examples shown in Figures 15(a) and 15(b), if the upper boundary of the current block and the upper boundary of the coding tree unit touch each other, the reference samples above the current block belong to a coding tree unit different from the current block, and the samples adjacent to the upper boundary of the coding tree unit can replace the reference samples that do not adjacent to the upper boundary of the coding tree unit among the reference samples that belong to a coding tree unit different from the current block.

[0154] For example, as shown in the example of Figure 15(a), the reference sample at position 6 is replaced with the sample located at position 6' of the upper boundary of the coding tree unit. As shown in the example of Figure 15(b), the reference sample at position 15 is replaced with the sample located at position 15' of the upper boundary of the coding tree unit. In this case, the y coordinate of the replacement sample may be changed to an adjacent position in the coding tree unit, and the x coordinate of the replacement sample may be set to be the same as the reference sample. For example, the sample at position 6' may have the same x coordinate as the sample at position 6, and the sample at position 15' may have the same x coordinate as the sample at position 15.

[0155] Alternatively, the x-coordinate of the replacement sample can be determined by adding or subtracting an offset value to or from the x-coordinate of the reference sample. For example, if an adjacent reference sample and a non-adjacent reference sample located above the current block have the same x-coordinate, the x-coordinate of the replacement sample can be determined by adding or subtracting an offset value to or from the x-coordinate of the reference sample. The purpose of this is to prevent the replacement sample for a non-adjacent reference sample from being located at the same position as other non-adjacent or adjacent reference samples.

[0156] FIG. 16 is a diagram showing an example of fluctuation in the position of the reference sample.

[0157] When a sample located at the boundary of a coding tree unit replaces a reference sample that is included in a coding tree unit different from the current block and is not adjacent to the boundary of the coding tree unit, the value obtained by adding or subtracting an offset value to the x coordinate of the reference sample can be used as the x coordinate of the replacement sample.

[0158] 16 , the reference sample at position 6 and the reference sample at position 15 may be replaced with the sample at position 6′ and the sample at position 15′, respectively, which have the same y-coordinate as the row adjacent to the upper boundary of the coding tree unit. In this case, the x-coordinate of the sample at position 6′ may be set to have a difference value of W / 2 from the x-coordinate of the reference sample at position 6, and the x-coordinate of the sample at position 15′ may be set to have a difference value of W−1 from the x-coordinate of the reference sample at position 15.

[0159] Different from the examples shown in Figures 15 and 16, the y coordinate of the row located above the top row of the current block or the y coordinate of the upper boundary of the coding tree unit may also be set as the y coordinate of the replacement sample.

[0160] Although not shown, a sample to replace the reference sample may be determined based on the left boundary of the coding tree unit. For example, if the reference sample is not included in the same coding tree unit as the current block and is not adjacent to the left boundary of the coding tree unit, the reference sample may be replaced with a sample adjacent to the left boundary of the coding tree unit. In this case, the replacement sample may have the same y coordinate as the reference sample, or may have a y coordinate obtained by adding or subtracting an offset value to the y coordinate of the reference sample.

[0161] Subsequently, the block containing the replacement sample is taken as a candidate block, and based on the candidate block, a merging candidate for the current block can be derived.

[0162] Merging candidates can also be derived from temporally neighboring blocks in a different image than the current block, for example, from collocated blocks in a collocated image.

[0163] The motion information of the merging candidate may be set to be the same as the motion information of the candidate block, for example, at least one of the motion vector, reference image index, prediction direction, or bidirectional weight index of the candidate block may be the motion information of the merging candidate.

[0164] A merge candidate list including merge candidates may be generated (S1102). The merge candidates may be classified into adjacent merge candidates derived from adjacent blocks adjacent to the current block and non-adjacent merge candidates derived from non-adjacent blocks.

[0165] Indices for multiple merge candidates in the merge candidate list may be assigned according to a predetermined order. For example, indices assigned to adjacent merge candidates may have lower values ​​than indices assigned to non-adjacent merge candidates. Alternatively, indices may be assigned to each merge candidate based on the index of each block shown in FIG. 12 or FIG. 14.

[0166] If the merge candidate list includes multiple merge candidates, at least one of the multiple merge candidates may be selected (S1103). In this case, information indicating whether motion information of the current block is derived from a neighboring merge candidate may be signaled via the bitstream. The information may be a 1-bit flag. For example, a syntax element isAdjancentMergeFlag indicating whether motion information of the current block is derived from a neighboring merge candidate may be signaled via the bitstream. If the syntax element isAdjancentMergeFlag has a value of 1, the motion information of the current block may be derived based on a neighboring merge candidate. On the other hand, if the syntax element isAdjancentMergeFlag has a value of 0, the motion information of the current block may be derived based on a non-neighboring merge candidate.

[0167] Table 1 shows a syntax table including the syntax element isAdjancentMergeFlag. [Table 1(1)]

[0168] [Table 1(2)]

[0169] Information for specifying one of multiple merge candidates can be signaled via the bitstream, for example, information indicating the index of one of the merge candidates included in the merge candidate list can be signaled via the bitstream.

[0170] If isAdjacentMergeflag is 1, a signal can be used to transmit a syntax element merge_idx for determining one of the adjacent merge candidates. The maximum value of the syntax element merge_idx can be a value whose difference from the number of adjacent merge candidates is 1.

[0171] If isAdjacentMergeflag is 0, the signal can be used to transmit the syntax element NA_merge_idx to determine one of the non-adjacent merge candidates. The syntax element NA_merge_idx indicates a value obtained by subtracting the index of the non-adjacent merge candidate from the number of adjacent merge candidates. The decoder can select a non-adjacent merge candidate by adding the number of adjacent merge candidates to the index based on NA_merge_idx.

[0172] If the number of merge candidates included in the merge candidate list is less than a threshold, merge candidates included in the inter motion information list may be added to the merge candidate list. Here, the threshold may be the maximum number of merge candidates included in the merge candidate list or a value obtained by subtracting an offset value from the maximum number of merge candidates. The offset value may be an integer such as 1 or 2. The inter motion information list may include merge candidates derived based on blocks coded / decoded before the current block.

[0173] The inter motion information list includes merge candidates derived from blocks in the current image that are coded / decoded based on inter prediction. For example, the motion information of the merge candidates included in the inter motion information list may be set equal to the motion information of the blocks that are coded / decoded based on inter prediction. Here, the motion information may include at least one of a motion vector, a reference image index, a prediction direction, or a bidirectional weight index. For ease of interpretation, merge candidates included in the inter motion information list are referred to as inter merge candidates.

[0174] When selecting a merge candidate for the current block, the motion vector of the merge candidate is set as an initial motion vector. Furthermore, motion compensation prediction of the current block can be performed by adding or subtracting an offset vector to or from the initial motion vector. Deriving a new motion vector by adding or subtracting an offset vector to or from the motion vector of the merge candidate may be defined as a merge motion differential coding method.

[0175] Information indicating whether the merge offset coding method is used can be signaled via the bitstream. The information may be a 1-bit flag, merge_offset_vector_flag. For example, a value of 1 in merge_offset_vector_flag indicates that the merge motion differential coding method is applied to the current block. When the merge motion differential coding method is applied to the current block, the motion vector of the current block can be derived by adding or subtracting an offset vector from the motion vector of the merge candidate. A value of 0 in merge_offset_vector_flag indicates that the merge motion differential coding method is not applied to the current block. When the merge offset coding method is not applied, the motion vector of the merge candidate can be used as the motion vector of the current block.

[0176] It can be signaled only when the value of the skip flag indicating whether to apply the skip mode is true or the value of the merge flag indicating whether to apply the merge mode is true. For example, when the value of the skip_flag indicating whether to apply the skip mode to the current block is 1 or the value of the merge_flag indicating whether to apply the merge mode to the current block is 1, the merge_offset_vector_flag can be coded and signaled.

[0177] If it is determined that the merge offset encoding method is to be applied to the current block, at least one of information specifying one of the merge candidates included in the merge candidate list, information indicating the magnitude of the offset vector, and information indicating the direction of the offset vector can be separately signaled.

[0178] Information can be signaled via the bitstream to determine the maximum number of merge candidates that can be included in a merge candidate list, for example, the maximum number of merge candidates that can be included in a merge candidate list can be 6 or a smaller integer.

[0179] When it is determined that the merge offset coding method is to be applied to the current block, only a predetermined maximum number of merge candidates can be used as the initial motion vector of the current block. That is, the number of merge candidates available for the current block can be adaptively determined based on whether the merge offset coding method is to be applied. For example, when the value of merge_offset_vector_flag is set to 0, the maximum number of merge candidates available for the current block may be set to M. When the value of merge_offset_vector_flag is set to 1, the maximum number of merge candidates available for the current block may be set to N. Here, M represents the maximum number of merge candidates included in the merge candidate list, and N represents an integer equal to or less than M.

[0180] For example, if M is 6 and N is 2, the two merge candidates with the smallest indexes among those included in the merge candidate list may be set as available for the current block. Thus, the motion vector of the merge candidate with index value 0 or the motion vector of the merge candidate with index value 1 may be set as the initial motion vector of the current block. If M and N are the same (e.g., M and N are 2), all merge candidates included in the merge candidate list may be set as available for the current block.

[0181] Alternatively, whether a neighboring block can be used as a merge candidate may be determined based on whether a merge motion differential coding method is applied to the current block. For example, if the value of merge_offset_vector_flag is 1, at least one of the neighboring blocks adjacent to the upper right corner, the lower left corner, and the lower left corner of the current block may be set as unavailable as a merge candidate. Therefore, when the merge motion differential coding method is applied to the current block, the motion vector of at least one of the neighboring blocks adjacent to the upper right corner, the lower left corner, and the lower left corner of the current block cannot be set as an initial motion vector. Alternatively, if the value of merge_offset_vector_flag is 1, the temporal neighboring blocks of the current block may be set as unavailable as merge candidates.

[0182] When applying the merge motion difference coding method to the current block, at least one of the paired merge candidates and the zero merge candidate may be set not to be used. Therefore, if the value of merge_offset_vector_flag is 1, at least one of the paired merge candidates or the zero merge candidate may not be added to the merge candidate list even if the number of merge candidates included in the merge candidate list is less than the maximum number.

[0183] The motion vector of a merge candidate may be set as the initial motion vector of the current block. In this case, if there are multiple merge candidates available for the current block, information specifying one of the multiple merge candidates may be signaled via the bitstream. For example, if the maximum number of merge candidates included in the merge candidate list is greater than one, information merge_idx indicating one of the multiple merge candidates may be signaled via the bitstream. That is, in a merge offset coding method, a merge candidate may be specified by information merge_idx for specifying one of the multiple merge candidates. The initial motion vector of the current block may be set as the motion vector of the merge candidate indicated by merge_idx.

[0184] On the other hand, if the number of merge candidates available for the current block is one, signaling of information for specifying a merge candidate may be omitted. For example, if the maximum number of merge candidates that can be included in the merge candidate list is one or less, signaling of information merge_idx for specifying a merge candidate may be omitted. That is, in the merge offset encoding method, when one merge candidate is included in the merge candidate list, encoding of information merge_idx for specifying a merge candidate may be omitted, and an initial motion vector may be determined based on the merge candidate included in the merge candidate list. The motion vector of the merge candidate may be set as the initial motion vector of the current block.

[0185] As another example, after determining merge candidates for the current block, it may be determined whether to apply the merge motion differential coding method to the current block. For example, if the maximum number of merge candidates that can be included in the merge candidate list is greater than one, information merge_idx for specifying one of the merge candidates may be signaled. After selecting a merge candidate based on merge_idx, decoding may be performed on merge_offset_vector_flag, which indicates whether to apply the merge motion differential coding method to the current block. Table 2 illustrates a syntax table according to the above embodiment. [Table 2(1)]

[0186] [Table 2(2)]

[0187] As another example, after determining merge candidates for the current block, it may be determined whether to apply the merge motion differential coding method to the current block only if the determined merge candidate index is smaller than the maximum number of merge candidates available when applying the merge motion differential coding method. For example, only if the value of index information merge_idx is less than N, it may be coded and signaled for merge_offset_vector_flag, which indicates whether to apply the merge motion differential coding method to the current block. If the value of index information merge_idx is greater than or equal to N, coding for merge_offset_vector_flag may be omitted. If coding for merge_offset_vector_flag is omitted, it may be determined that the merge motion differential coding method is not applied to the current block.

[0188] Alternatively, after determining a merge candidate for the current block, whether to apply the merge motion differential coding method to the current block may be determined taking into account whether the determined merge candidate has bidirectional motion information or unidirectional motion information. For example, merge_offset_vector_flag, which indicates whether to apply the merge motion differential coding method to the current block, may be coded and signaled only when the value of index information merge_idx is less than N and the merge candidate selected based on the index information has bidirectional motion information. Optionally, merge_offset_vector_flag, which indicates whether to apply the merge motion differential coding method to the current block, may be coded and signaled only when the value of index information merge_idx is less than N and the merge candidate selected based on the index information has unidirectional motion information.

[0189] Alternatively, whether to apply the merge motion differential coding method may be determined based on at least one of the size of the current block, the shape of the current block, and whether the current block touches a boundary of a coding tree unit. If at least one of the size of the current block, the shape of the current block, and whether the current block touches a boundary of a coding tree unit does not satisfy a predetermined condition, coding of merge_offset_vector_flag, which indicates whether to apply the merge motion differential coding method to the current block, may be omitted.

[0190] When selecting a merge candidate, the motion vector of the merge candidate may be set as the initial motion vector of the current block. Then, information indicating the magnitude and direction of the offset vector may be decoded to determine the offset vector. The offset vector may have a horizontal component or a vertical component.

[0191] The information indicating the magnitude of the offset vector may be index information indicating one of the motion magnitude candidates. For example, index information distance_idx indicating one of the motion magnitude candidates may be signaled via a bitstream. Table 3 shows the binarization of the index information distance_idx and the value of the variable DistFromMergeMV for determining the magnitude of the offset vector based on distance_idx. [Table 3]

[0192] The magnitude of the offset vector can be derived by dividing the variable DistFromMergeMV by a predetermined value. Equation 4 shows an example of determining the magnitude of the offset vector.

number

[0193] According to Equation 4, the value obtained by dividing the variable DistFromMegeMV by 4 or the value obtained by shifting the variable DistFromMergeMV by 2 to the left can be set as the magnitude of the offset vector.

[0194] More or fewer motion magnitude candidates than the example shown in Table 3 may be used, or the range of motion vector offset magnitude candidates may be set differently from the example shown in Table 3. For example, the magnitude of the horizontal or vertical component of the offset vector may be set to two sample distances or less. Table 4 shows the binarization of the index information distance_idx and the value of the variable DistFromMergeMV for determining the magnitude of the offset vector based on distance_idx. [Table 4]

[0195] Alternatively, the range of motion vector offset magnitude candidates may be set differently based on the motion vector precision. For example, if the motion vector precision of the current block is fractional-pel, the value of the variable DistFromMergeMV corresponding to the value of the index information distance_idx may be 1, 2, 4, 8, 16, etc. Here, fractional pel includes at least one of 1 / 16 pixel, 1 / 8 pixel (octo-pel), 1 / 4 pixel (quarter-pel), and 1 / 2 pixel (half-pel). On the other hand, if the motion vector precision of the current block is integer pixel, the value of the variable DistFromMergeMV corresponding to the value of the index information distance_idx may be 4, 8, 16, 32, 64, etc. In other words, the table for determining the variable DistFromMergeMV may be set differently based on the motion vector precision of the current block.

[0196] For example, if the motion vector precision of the current block or merging candidate is 1 / 4 pixel, the variable DistFromMergeMV represented by distance_idx can be derived using Table 3. On the other hand, if the motion vector precision of the current block or merging candidate is integer pixel, the value of the variable DistFromMergeMV represented by distance_idx in Table 3 can be multiplied by N (for example, 4 times) and the derived value can be used as the value of the variable DistFromMergeMV.

[0197] Information for determining motion vector accuracy can be signaled via the bitstream. For example, information can be transmitted by signal at the sequence level, image level, slice level, or block level. Therefore, the range of motion magnitude candidates can be set differently based on the information related to motion vector accuracy signaled via the bitstream. Alternatively, the motion vector accuracy can be determined based on the merging candidate of the current block. For example, the motion vector accuracy of the current block can be set to be the same as the motion vector accuracy of the merging candidate.

[0198] Alternatively, information for determining a search range for an offset vector can be signaled via a bitstream. At least one of the number of motion magnitude candidates, the minimum value of the motion magnitude candidates, and the maximum value of the motion magnitude candidates can be determined based on the search range. For example, a flag "merge_offset_vector_flag" for determining the search range for the offset vector can be signaled via a bitstream. The information can be signaled via a sequence header, a picture header, or a slice header.

[0199] For example, if the value of merge_offset_extend_range_flag is 0, the magnitude of the offset vector can be 2 or less. Therefore, the maximum value of DistFromMergeMV can be 8. On the other hand, if the value of merge_offset_extend_range_flag is 1, the magnitude of the offset vector can be 32 sample distances. Therefore, the maximum value of DistFromMergeMV can be 128.

[0200] The magnitude of the offset vector can be determined using a flag indicating whether the magnitude of the offset vector is greater than a threshold. For example, a flag distance_flag indicating whether the magnitude of the offset vector is greater than a threshold can be signaled via the bitstream. The threshold may be 1, 2, 4, 8, or 16. For example, if distance_flag is 1, it indicates that the magnitude of the offset vector is greater than 4. On the other hand, if distance_flag is 0, it indicates that the magnitude of the offset vector is equal to or less than 4.

[0201] If the magnitude of the offset vector is greater than the threshold, the index information distance_idx can be used to derive a difference value between the magnitude of the offset vector and the threshold. Alternatively, if the magnitude of the offset vector is equal to or less than the threshold, the index information distance_idx can be used to determine the magnitude of the offset vector. Table 5 shows a syntax table for the process of encoding distance_flag and distance_idx. [Table 5(1)]

[0202] [Table 5(2)]

[0203] Equation 5 shows an example of deriving a variable DistFromMergeMV for determining the magnitude of the offset vector using distance_flag and distance_idx.

number

[0204] In Equation 5, the value of distance_flag can be 1 or 0. The value of distance_idx can be 1, 2, 4, 8, 16, 32, 64, 128, etc. N represents a coefficient determined by a threshold. For example, if the threshold is 4, N can be 16.

[0205] The information indicating the direction of the offset vector may be index information indicating one of vector direction candidates. For example, index information "direction_idx" indicating one of the vector direction candidates may be signaled via a bitstream. Table 6 shows the binarization of the index information "direction_idx" and the direction of the offset vector according to "direction_idx." [Table 6]

[0206] In Table 6, sign[0] indicates the horizontal direction, and sign[1] indicates the vertical direction. +1 indicates that the value of the x or y component of the offset vector is a positive number (+), and -1 indicates that the value of the x or y component of the offset vector is a negative number (-). Equation 6 shows an example of determining an offset vector based on the magnitude and direction of the offset vector.

number

[0207] In Equation 6, offsetMV[0] indicates the vertical component of the offset vector, and offsetMV[1] indicates the horizontal component of the offset vector.

[0208] FIG. 17 is a diagram showing offset vectors based on the values ​​of distance_idx, which indicates the magnitude of the offset vector, and direction_idx, which indicates the direction of the offset vector.

[0209] In the example shown in FIG. 7, the magnitude and direction of the offset vector can be determined based on the values ​​of distance_idx and direction_idx. The maximum magnitude of the offset vector can be set to a threshold value or less. Here, the threshold value may have a value predefined by the encoder and decoder. For example, the threshold value may be a distance of 32 samples. Alternatively, the threshold value can be determined based on the magnitude of the initial motion vector. For example, the horizontal threshold value can be set based on the magnitude of the horizontal component of the initial motion vector, and the vertical threshold value can be set based on the vertical component of the initial motion vector.

[0210] If the merge candidate has bidirectional motion information, the L0 motion vector of the merge candidate can be set as the L0 initial motion vector of the current block, and the L1 motion vector of the merge candidate can be set as the L1 initial motion vector of the current block. In this case, the L0 offset vector and the L1 offset vector can be determined by considering the output order difference value (hereinafter referred to as the L0 difference value) between the L0 reference image of the merge candidate and the current image and the output order difference value (hereinafter referred to as the L1 difference value) between the L1 reference image of the merge candidate and the current image.

[0211] First, if the symbols of the L0 differential value and the L1 differential value are the same, the L0 offset vector and the L1 offset vector can be set to be the same, whereas if the symbols of the L0 differential value and the L1 differential value are different, the L1 offset vector can be set in the opposite direction to the L0 offset vector.

[0212] The magnitude of the L0 offset vector and the magnitude of the L1 offset vector can be set to be the same. Alternatively, the magnitude of the L1 offset vector can be determined by shrinking the L0 offset vector based on the L0 difference value and the L1 difference value.

[0213] For example, Equation 7 shows the L0 offset vector and the L1 offset vector when the symbols of the L0 differential value and the L1 differential value are the same.

number

[0214] In Equation 7, offsetMVL0[0] indicates the horizontal component of the L0 offset vector, offsetMVL0[1] indicates the vertical component of the L0 offset vector, offsetMVL1[0] indicates the horizontal component of the L1 offset vector, and offsetMVL1[1] indicates the vertical component of the L1 offset vector.

[0215] Equation 8 shows the L0 offset vector and the L1 offset vector when the symbols of the L0 differential value and the L1 differential value are different.

number

[0216] It is also possible to define four or more vector direction candidates. Tables 7 and 8 show examples of defining eight vector direction candidates. [Table 7] [Table 8]

[0217] In Tables 7 and 8, if the absolute values ​​of sign[0] and sign[1] are greater than 0, it indicates that the offset vector is diagonal. When using Table 6, the magnitude of the x-axis and y-axis components of the diagonal offset vector is abs(offsetMV), and when using Table 7, the magnitude of the x-axis and y-axis components of the diagonal offset vector is abs(offsetMV / 2).

[0218] FIG. 18 is a diagram showing offset vectors based on the values ​​of distance_idx, which indicates the magnitude of the offset vector, and direction_idx, which indicates the direction of the offset vector.

[0219] FIG. 18(a) is an example in which Table 6 is applied, and FIG. 18(b) is an example in which Table 7 is applied.

[0220] Information for determining at least one of the number or magnitude of vector direction candidates can be signaled via the bitstream. For example, a flag, merge_offset_direction_range_flag, for determining vector direction candidates can be signaled via the bitstream. The flag can be transmitted by signaling at the sequence level, the image level, or the slice level. For example, if the flag value is 0, the four vector direction candidates shown in Table 6 can be used. On the other hand, if the flag value is 1, the eight vector direction candidates shown in Table 7 or Table 8 can be used.

[0221] Alternatively, at least one of the number or magnitude of the vector direction candidates can be determined based on the magnitude of the offset vector. For example, when the value of the variable DistFromMergeMV for determining the magnitude of the offset vector is equal to or less than a threshold, the eight vector direction candidates shown in Table 7 or Table 8 can be used. On the other hand, when the value of the variable DistFromMergeMV is greater than the threshold, the four vector direction candidates shown in Table 6 can be used.

[0222] Alternatively, at least one of the number or magnitude of vector direction candidates can be determined based on the value MVx of the x component and the value MVy of the y component of the initial motion vector. For example, if the difference or the absolute value of the difference between MVx and MVy is equal to or less than a threshold, the eight vector direction candidates shown in Table 7 or Table 8 can be used. On the other hand, if the difference or the absolute value of the difference between MVx and MVy is greater than the threshold, the four vector direction candidates shown in Table 6 can be used.

[0223] The motion vector of the current block can be derived by adding the initial motion vector to the offset vector. Equation 9 shows an example of determining the motion vector of the current block.

number

[0224] In Equation 9, mvL0 indicates the L0 motion vector of the current block, mvL1 indicates the L1 motion vector of the current block, mergeMVL0 indicates the L0 initial motion vector of the current block (i.e., the L0 motion vector of the merge candidate), and mergeMVL1 indicates the L1 initial motion vector of the current block. [0] indicates the horizontal component of the motion vector, and [1] indicates the vertical component of the motion vector.

[0225] Intra prediction is a method of predicting a current block using coded / decoded reconstructed samples around the current block. In this case, the intra prediction of the current block can use reconstructed samples before applying an in-loop filter.

[0226] Intra prediction techniques include matrix-based intra prediction and general intra prediction that takes into account the directionality of surrounding reconstructed samples. Information indicating the intra prediction technique of the current block may be signaled via the bitstream. The information may be a one-bit flag. Alternatively, the intra prediction technique of the current block may be determined based on at least one of the position, size, and shape of the current block or the intra prediction techniques of neighboring blocks. For example, if the current block straddles an image boundary, the current block may be set to not apply matrix-based intra prediction.

[0227] Matrix-based intra prediction is a method of obtaining a prediction block of a current block based on matrix multiplication of a matrix stored in an encoder and a decoder with reconstructed samples surrounding the current block. Information indicating one of a plurality of stored matrices can be signaled via a bitstream. The decoder can determine a matrix to be used for intra prediction of the current block based on the information and the size of the current block.

[0228] General intra prediction is a method of obtaining a predicted block related to a current block based on a non-angular intra prediction mode or an angular intra prediction mode.

[0229] A residual image can be derived by subtracting the original image from the predicted image. In this case, when the residual image is converted to the frequency domain, removing high-frequency components from the frequency components does not significantly degrade the subjective image quality of the video. Therefore, reducing the value of the high-frequency components or setting them to zero has the effect of improving compression efficiency without causing obvious visual distortion. To reflect this characteristic, the residual image can be decomposed into two-dimensional frequency components by transforming the current block. The transformation can be performed using a transform technique such as a discrete cosine transform (DCT) or a discrete sine transform (DST).

[0230] After transforming a current block using a DCT or DST, the transformed current block can be transformed again. In this case, the transformation based on the DCT or DST is defined as the first transformation, and the process of transforming the block to which the first transformation has been applied again is called the second transformation.

[0231] The first transform can be performed using any one of a number of candidate transform kernels, for example, any one of DCT2, DCT8, or DCT7.

[0232] Different transform kernels can be used for the horizontal and vertical directions, and information representing a combination of horizontal and vertical transform kernels can be signaled via the bitstream.

[0233] The first and second transforms are performed by different units. For example, the first transform may be performed on an 8x8 block, and the second transform may be performed on a 4x4 sub-block of the transformed 8x8 block. In this case, the transform coefficients of the remaining area where the second transform is not performed may be set to 0.

[0234] Alternatively, a first transform can be performed on a 4x4 block and a second transform can be performed on a region of size 8x8 that contains the transformed 4x4 block.

[0235] Information can be signaled via the bitstream indicating whether or not to perform a second transformation.

[0236] In the decoder, the inverse transform of the second transform (second inverse transform) can be performed, and the inverse transform of the first transform (first inverse transform) can be performed on the result. As a result of performing the second inverse transform and the first inverse transform, a residual signal of the current block can be obtained.

[0237] Quantization is used to reduce the energy of a block, and the quantization process involves dividing the transform coefficients by a certain constant, which may be derived from a quantization parameter, which may be defined as a value between 1 and 63.

[0238] After performing the transform and quantization at the encoder, the decoder can obtain the residual block by inverse quantization and inverse transform, and at the decoder, the predicted block and the residual block can be added to obtain the reconstructed block of the current block.

[0239] Once a reconstructed block of the current block is obtained, in-loop filtering can be performed to reduce information loss that occurs in the quantization and encoding processes. The in-loop filter may include at least one of a deblocking filter, a sample adaptive offset filter (SAO), or an adaptive loop filter (ALF).

[0240] It is within the scope of the present invention to use the embodiments described with emphasis on the decoding or encoding process in an encoding or decoding process. It is also within the scope of the present invention to modify the embodiments described in a given order in a different order than that described.

[0241] Although the embodiments have been described based on a series of steps or flowcharts, this does not limit the chronological order of the invention. Furthermore, steps may be executed simultaneously or in a different order, as needed. In the above embodiments, the components (e.g., units, modules, etc.) constituting the block diagrams may each be realized as hardware devices or software. Furthermore, multiple components may be combined and executed as a single hardware device or software. The embodiments may be implemented in the form of program instructions. The program instructions may be executed by various computer components and recorded on a computer-readable storage medium. The computer-readable storage medium may include program instructions, data files, data structures, etc., alone or in combination. Examples of computer-readable storage media include magnetic media such as hard disks, flexible disks, and magnetic tapes; optical media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specially configured to store and execute program instructions, such as ROM, RAM, flash memory, etc. The hardware devices may be configured to operate as one or more software modules to perform the processes according to the present invention, or vice versa. [Industrial Applicability]

[0242] The present invention is applicable to electronic devices that encode / decode video.

Claims

1. 1. A video decoding method comprising: generating a merge candidate list for the current block; designating a merge candidate for the current block from among the merge candidates included in the merge candidate list; deriving an offset vector for the current block; deriving a motion vector of the current block by adding the motion vector of the merging candidate to the offset vector; the offset vector is determined based on first index information and second index information, the first index information is used to designate one of motion magnitude candidates, and the second index information is used to determine a direction of the offset vector; A range of motion vector offset magnitude candidates is set to be different depending on the motion vector accuracy of the current block; The video decoding method, wherein the magnitude of the offset vector is obtained by shifting the value represented by the motion magnitude candidate specified by the first index information left by two bits.

2. at least one of the maximum value and the minimum value of the motion magnitude candidates is set to be different based on a value of a flag indicating a range of the motion magnitude candidates; The flag is signaled at the image level. The video decoding method of claim 1 .

3. If the motion vector precision of the current block is a fractional pixel, the value of the variable corresponding to the value of the index information is 1, 2, 4, 8, or 16; or If the motion vector precision of the current block is integer pixel, the value of the variable corresponding to the value of the index information is 4, 8, 16, 32 or 64; The video decoding method of claim 1 .

4. The subpixel is a quarter pixel.

4. The video decoding method of claim 3.

5. the information for determining the motion vector precision is signaled at a picture level in the bitstream. The video decoding method of claim 1 .

6. 1. A video encoding method comprising: generating a merge candidate list for the current block; designating a merge candidate for the current block from among the merge candidates included in the merge candidate list; deriving an offset vector for the current block; deriving a motion vector of the current block by adding the motion vector of the merging candidate to the offset vector; the offset vector is determined based on first index information and second index information, the first index information is used to designate one of motion magnitude candidates, and the second index information is used to determine a direction of the offset vector; A range of motion vector offset magnitude candidates is set to be different depending on the motion vector accuracy of the current block; A video encoding method, wherein the magnitude of the offset vector is obtained by shifting the value represented by the motion magnitude candidate specified by the first index information by two bits to the left.

7. at least one of the maximum value and the minimum value of the motion magnitude candidates is set to be different based on a value of a flag indicating a range of the motion magnitude candidates; The flag is signaled at the image level.

7. The video encoding method of claim 6.

8. If the motion vector precision of the current block is a fractional pixel, the value of the variable corresponding to the value of the index information is 1, 2, 4, 8, or 16; or If the motion vector precision of the current block is integer pixel, the value of the variable corresponding to the value of the index information is 4, 8, 16, 32 or 64; 7. The video encoding method of claim 6.

9. The subpixel is a quarter pixel.

9. The video encoding method of claim 8.

10. the information for determining the motion vector precision is signaled at a picture level in the bitstream.

7. The video encoding method of claim 6.

11. 6. A computer-readable storage medium having stored thereon a computer program and a bitstream, the computer program, when executed by a processor, causing the processor to perform the video decoding method of any one of claims 1 to 5 to generate the bitstream.

12. 11. A computer-readable storage medium having stored thereon a computer program and a bitstream, the computer program, when executed by a processor, causing the processor to perform the video encoding method of any one of claims 6 to 10 to generate the bitstream.

Citation Information

Patent Citations

  • Method for encoding and decoding motion information and device for encoding and decoding motion information

    WO2019054736A1

  • Systems and methods for performing motion vector prediction for video coding using motion vector predictor origins

    WO2019151093A1

  • Encoding method and device thereof, and decoding method and device thereof

    WO2019168244A1

  • Coding device, decoding device, coding method, and decoding method

    WO2020017367A1