Video signal encoding / decoding method and apparatus therefor

CN120547327BActive Publication Date: 2026-09-15GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510655807.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-11-08
Filing Date
2019-11-08
Publication Date
2026-09-15
Estimated Expiration
2039-11-08

AI Technical Summary

Technical Problem

高清视频服务的最大问题在于数据量大幅增加,为了解决这种问题,正在积极进行用于提高视频压缩率的研究

Benefits of technology

[0018] According to the present invention, the efficiency of inter-frame prediction can be improved by refining the motion vectors of the merged candidates based on the offset vector.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120547327B_ABST
    Figure CN120547327B_ABST
Patent Text Reader

Abstract

The video decoding method according to the present application comprises the steps of: generating a merge candidate list of a current block; determining a merge candidate of the current block from the merge candidates included in the merge candidate list; deriving an offset vector of the current block; and deriving a motion vector of the current block by adding the offset vector to a motion vector of the merge candidate.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This case is a divisional application of Chinese national phase patent application 201980071283.X, which was filed on November 8, 2019, under international patent application PCT / KR2019 / 015194. Technical Field

[0002] This invention relates to a video signal encoding / decoding method and an apparatus for the method. Background Technology

[0003] With the trend of increasingly larger display panels, there is a growing need for higher-quality video services. The biggest problem with high-definition video services is the significant increase in data volume. To address this issue, research is actively underway to improve video compression rates. As a representative example, in 2009, the Moving Picture Experts Group (MPEG) and the Video Coding Experts Group (VCEG) of the International Telecommunication Union-Telecommunication (ITU-T) established the Joint Collaborative Team on Video Coding (JCT-VC). JCT-VC proposed the video compression standard HEVC (High Efficiency Video Coding), which was approved on January 25, 2013, and its compression performance is approximately twice that of H.264 / AVC. However, with the rapid development of high-definition video services, the limitations of HEVC have gradually become apparent. Summary of the Invention

[0004] Technical problems to be solved

[0005] The purpose of this invention is to provide a method for refining motion vectors derived from merging candidates based on offset vectors when encoding / decoding video signals, and an apparatus for performing the method.

[0006] The purpose of this invention is to provide a method for transmitting an offset vector using a signal during the encoding / decoding of a video signal, and an apparatus for performing the method.

[0007] The technical problems to be solved by the present invention are not limited to those mentioned above, and those skilled in the art to which the present invention pertains will clearly understand other technical problems not mentioned through the following description.

[0008] Technical solution

[0009] The video signal decoding / encoding method according to the present invention includes the following steps: generating a merging candidate list for the current block; determining a merging candidate for the current block from the merging candidates included in the merging candidate list; deriving an offset vector for the current block; and deriving a motion vector for the current block by adding the offset vector to the motion vector of the merging candidate.

[0010] In the video signal decoding / encoding method according to the present invention, the size of the offset vector can be determined based on the first index information of any of the specified motion size candidates.

[0011] In the video signal decoding / encoding method according to the present invention, at least one of the maximum or minimum values ​​of the motion size candidate can be set differently depending on the value of the flag indicating the range of the motion size candidate.

[0012] In the video signal decoding / encoding method according to the present invention, the flag can be transmitted as an image-level signal.

[0013] In the video signal decoding / encoding method according to the present invention, at least one of the maximum or minimum values ​​of the motion size candidate can be set differently depending on the motion vector accuracy of the current block.

[0014] In the video signal decoding / encoding method according to the present invention, the magnitude of the offset vector can be obtained by shifting the value represented by the motion size candidate specified by the first index information.

[0015] In the video signal decoding / encoding method according to the present invention, the direction of the offset vector can be determined based on the second index information of any of the candidate vector directions.

[0016] The features briefly outlined above are merely exemplary embodiments of the invention as described in the detailed description to follow, and do not limit the scope of the invention.

[0017] Invention Effects

[0018] According to the present invention, the efficiency of inter-frame prediction can be improved by refining the motion vectors of the merged candidates based on the offset vector.

[0019] According to the present invention, the efficiency of inter-frame prediction can be improved by adaptively determining the magnitude and direction of the offset vector.

[0020] The effects that can be obtained in this invention are not limited to those described above, and other effects not mentioned will be clearly understood by those skilled in the art through the following description. Attached Figure Description

[0021] Figure 1 This is a block diagram of a video encoder according to an embodiment of the present invention.

[0022] Figure 2 This is a block diagram of a video decoder according to an embodiment of the present invention.

[0023] Figure 3 This is a diagram illustrating the basic coding tree unit of an embodiment of the present invention.

[0024] Figures 4(a)-4(e) This is a diagram illustrating the various partitioning types of coded blocks.

[0025] Figure 5 This is a diagram illustrating the partitioning of coding tree units.

[0026] Figure 6 This is a flowchart of the inter-frame prediction method according to an embodiment of the present invention.

[0027] Figure 7 It is a diagram showing the nonlinear motion of an object.

[0028] Figure 8 This is a flowchart illustrating an inter-frame prediction method based on affine motion according to an embodiment of the present invention.

[0029] Figure 9 This is a diagram showing an example of the affine seed vector for each affine motion model.

[0030] Figure 10 This is a diagram showing an example of the affine vectors of a sub-block under a 4-parameter motion model.

[0031] Figure 11 This is a flowchart of the process of exporting motion information of the current block in merge mode.

[0032] Figure 12 This is a diagram showing an example of a candidate block used to derive merge candidates.

[0033] Figure 13 This is a diagram showing the location of the reference sample.

[0034] Figure 14 This is a diagram showing an example of a candidate block used to derive merge candidates.

[0035] Figure 15 This is a diagram illustrating an example of changing the position of a reference sample.

[0036] Figure 16 This is a diagram illustrating an example of changing the position of a reference sample.

[0037] Figure 17This is a graph showing the offset vector based on the values ​​of distance_idx, which indicates the magnitude of the offset vector, and direction_idx, which indicates the direction of the offset vector.

[0038] Figure 18 This is a graph showing the offset vector based on the values ​​of distance_idx, which indicates the magnitude of the offset vector, and direction_idx, which indicates the direction of the offset vector. Detailed Implementation

[0039] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0040] Video encoding and decoding are performed on a block-by-block basis. For example, encoding / decoding processes such as transform, quantization, prediction, in-loop filtering, or reconstruction can be performed on encoded blocks, transform blocks, or prediction blocks.

[0041] Hereinafter, the block to be encoded / decoded will be referred to as the "current block". For example, depending on the current encoding / decoding process step, the current block can represent an encoded block, a transform block, or a prediction block.

[0042] Additionally, as used herein, the term "unit" refers to a basic unit used to perform a specific encoding / decoding process, and "block" can be understood as representing a sample array of a predetermined size. Unless otherwise stated, "block" and "unit" are used interchangeably. For example, in the embodiments described later, encoding block and encoding unit can be understood to have the same meaning.

[0043] Figure 1 This is a block diagram of a video encoder according to an embodiment of the present invention.

[0044] Reference Figure 1 The video encoding device 100 may include an image segmentation unit 110, a prediction unit 120 and 125, a transformation unit 130, a quantization unit 135, a rearrangement unit 160, an entropy coding unit 165, an inverse quantization unit 140, an inverse transformation unit 145, a filter unit 150, and a memory 155.

[0045] Figure 1 The components shown are illustrated individually to illustrate the distinct functionalities of the video encoding device and do not imply that each component is composed of separate hardware or a single software component. That is, for ease of explanation, the components are arranged such that at least two components are combined into one, or one component is divided into multiple components, thereby performing functions. Such embodiments of integrated components and embodiments of separated components are also within the scope of this invention, provided they do not depart from its spirit.

[0046] Furthermore, some structural elements are not essential structural elements for performing the essential functions of this invention, but rather optional structural elements used only to improve performance. This invention can be implemented by including only the components necessary for realizing the essence of the invention, excluding the structural elements used only to improve performance, and structures including only the essential structural elements, excluding the optional structural elements used only to improve performance, are also within the scope of this invention.

[0047] The image partitioning unit 110 can divide the input image into at least one processing unit. In this case, the processing unit can be a prediction unit (PU), a transformation unit (TU), or a coding unit (CU). The image partitioning unit 110 divides an image into a combination of multiple coding units, prediction units, and transformation units, and can select a combination of coding units, prediction units, and transformation units to encode the image based on a predetermined criterion (e.g., a cost function).

[0048] For example, an image can be divided into multiple coding units. To segment an image into coding units, a recursive tree structure such as a quadtree structure can be used. A video or the largest coding unit can be used as the root, and the coding unit can be divided into additional coding units with a number of child nodes equivalent to the number of coding units in the division. Coding units that are no longer divided according to certain constraints become leaf nodes. That is, when it is assumed that a coding unit can only be divided into squares, a coding unit can be divided into a maximum of four other coding units.

[0049] In the embodiments of the present invention, the encoding unit may mean a unit that performs encoding, or it may mean a unit that performs decoding.

[0050] A prediction unit within a coding unit can be divided into at least one shape of the same size, such as a square or a rectangle, or a prediction unit within a coding unit can be divided into shapes and / or sizes different from those of another prediction unit.

[0051] When the prediction unit for intra-frame prediction based on the coding unit is not the smallest coding unit, intra-frame prediction can be performed without splitting into multiple prediction units N×N.

[0052] Prediction units 120 and 125 may include an inter-frame prediction unit 120 that performs inter-frame prediction and an intra-frame prediction unit 125 that performs intra-frame prediction. It can be determined whether inter-frame prediction or intra-frame prediction is used for the prediction unit, and specific information (e.g., intra-frame prediction mode, motion vector, reference image, etc.) is determined based on each prediction method. In this case, the processing unit performing the prediction may be different from the processing unit that determines the prediction method and specific content. For example, the prediction unit may determine the prediction method and prediction mode, and the transformation unit may perform the prediction. The residual value (residual block) between the generated prediction block and the original block can be input to the transformation unit 130. Furthermore, the prediction mode information, motion vector information, etc., used for prediction, along with the residual value, can be encoded in the entropy coding unit 165 and transmitted to the decoder. When using a specific coding mode, the original block can also be directly encoded and transmitted to the decoder without generating a prediction block through the prediction units 120 and 125.

[0053] The inter-frame prediction unit 120 can predict prediction units based on information from at least one of the previous or next images of the current image. In some cases, it can also predict prediction units based on information from a portion of the encoded region within the current image. The inter-frame prediction unit 120 may include a reference image interpolation unit, a motion prediction unit, and a motion compensation unit.

[0054] The reference image interpolation unit receives reference image information from memory 155 and can generate pixel information of integer pixels or less from the reference image. For luminance pixels, in order to generate pixel information of integer pixels or less in 1 / 4 pixel units, an 8th-order DCT-based interpolation filter with different filter coefficients can be used. For chrominance signals, in order to generate pixel information of integer pixels or less in 1 / 8 pixel units, a 4th-order DCT-based interpolation filter with different filter coefficients can be used.

[0055] The motion prediction unit can perform motion prediction based on a reference image interpolated by the reference image interpolation unit. Various methods can be used to calculate motion vectors, such as the Full Search-based Block Matching Algorithm (FBMA), the Three-Step Search (TSS), and the New Three-Step Search Algorithm (NTS). Motion vectors can have values ​​in units of 1 / 2 pixel or 1 / 4 pixel based on the interpolated pixels. Different motion prediction methods can be used in the motion prediction unit to predict the current prediction unit. These methods include skipping, merging, Advanced Motion Vector Prediction (AMVP), and Intra Block Copying.

[0056] The intra-prediction unit 125 can generate prediction units based on reference pixel information surrounding the current block, which serves as pixel information within the current image. When the neighboring block of the current prediction unit is a block that has already undergone inter-frame prediction, and the reference pixel is a pixel that has undergone inter-frame prediction, the reference pixel included in the block that has undergone inter-frame prediction can be used as the reference pixel information for the surrounding block that has undergone intra-frame prediction. That is, when a reference pixel is unavailable, at least one of the available reference pixels can be used to replace the unavailable reference pixel information.

[0057] In intra-frame prediction, the prediction mode can have an angular prediction mode that uses reference pixel information in the prediction direction and a non-angular mode that does not use direction information when performing prediction. The mode used to predict luminance information and the mode used to predict chrominance information can be different. To predict chrominance information, either the intra-frame prediction mode information used for predicting luminance information or the predicted luminance signal information can be applied.

[0058] When performing intra-frame prediction, if the size of the prediction unit is the same as the size of the transform unit, intra-frame prediction can be performed based on pixels to the left, upper left, and upper right of the prediction unit. However, when performing intra-frame prediction, if the size of the prediction unit is different from the size of the transform unit, intra-frame prediction can be performed using reference pixels based on the transform unit. Furthermore, intra-frame prediction using an N×N partition only for the smallest coding unit can be applied.

[0059] Intra-prediction methods can generate prediction blocks after applying an Adaptive Intra Smoothing (AIS) filter to a reference pixel based on the prediction mode. The type of AIS filter used for the reference pixel may vary. To perform intra-prediction, the intra-prediction mode of the current prediction unit can be predicted from the intra-prediction modes of prediction units existing in the vicinity of the current prediction unit. When using mode information predicted from surrounding prediction units to predict the prediction mode of the current prediction unit, if the intra-prediction modes of the current prediction unit and those of the surrounding prediction units are the same, predetermined flag information can be used to convey information indicating that the prediction modes of the current prediction unit and those of the surrounding prediction units are the same. If the prediction modes of the current prediction unit and those of the surrounding prediction units are different, entropy coding can be performed to encode the prediction mode information of the current block.

[0060] Furthermore, a residual block can be generated that includes residual value information, which is the difference between the original block of the prediction unit and the prediction unit that performs prediction based on the prediction unit generated in the prediction units 120 and 125. The generated residual block can be input to the transformation unit 130.

[0061] In the transform unit 130, transform methods such as Discrete Cosine Transform (DCT) or Discrete Sine Transform (DST) can be used to transform the original block and the residual block, which includes residual information between the prediction units generated by the prediction units 120 and 125. The DCT transform kernel includes at least one of DCT2 or DCT8, and the DST transform kernel includes DST7. Whether to apply DCT or DST to transform the residual block can be determined based on the intra-frame prediction mode information of the prediction units used to generate the residual block. Transformation of the residual block can also be skipped. A flag indicating whether to skip the transformation of the residual block can be encoded. Transformation skipping is allowed for residual blocks with a size below a threshold, or for luma or chroma components (4:4:4 format or below).

[0062] The quantization unit 135 can quantize the values ​​that have been transformed into the frequency domain in the transformation unit 130. The quantization coefficients can be changed according to the importance of the block or video. The values ​​calculated in the quantization unit 135 can be provided to the inverse quantization unit 140 and the rearrangement unit 160.

[0063] The rearrangement unit 160 can rearrange the coefficient values ​​of the quantized residual values.

[0064] The rearrangement unit 160 can transform 2D block shape coefficients into 1D vector form using a coefficient scanning method. For example, the rearrangement unit 160 can use a zig-zag scan method to scan the DC coefficients and even the coefficients in the high-frequency domain, and transform them into 1D vector form. Depending on the size of the transform unit and the intra-frame prediction mode, instead of zig-zag scanning, vertical scanning along the column direction and horizontal scanning along the row direction can also be used to scan the 2D block shape coefficients. That is, the choice between zig-zag scanning, vertical scanning, and horizontal scanning can be determined based on the size of the transform unit and the intra-frame prediction mode.

[0065] The entropy coding unit 165 can perform entropy coding based on the value calculated by the rearrangement unit 160. For example, entropy coding can use various coding methods such as Exponential Golomb code, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC).

[0066] The entropy coding unit 165 can encode various information such as residual coefficient information, block type information, prediction mode information, partitioning unit information, prediction unit information and transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information of the coding units originating from the rearrangement unit 160 and the prediction units 120 and 125.

[0067] The coefficient values ​​of the coding units input from the rearrangement unit 160 can be entropy encoded in the entropy coding unit 165.

[0068] The inverse quantization unit 140 and the inverse transform unit 145 perform inverse quantization on the multiple values ​​quantized in the quantization unit 135, and perform inverse transform on the values ​​transformed in the transform unit 130. The residual values ​​generated in the inverse quantization unit 140 and the inverse transform unit 145 can be merged with the prediction units predicted by the motion prediction unit, motion compensation unit and intra-frame prediction unit included in the prediction units 120 and 125 to generate a reconstructed block.

[0069] The filter unit 150 may include at least one of a deblocking filter, an offset correction unit, and an adaptive loop filter (ALF).

[0070] Deblocking filters remove block distortion generated in the reconstructed image due to the boundaries between blocks. To determine whether to perform deblocking, the number of pixels in the columns or rows included in the block can be used to decide whether to apply a deblocking filter to the current block. When applying a deblocking filter to a block, a strong or weak filter can be applied depending on the desired deblocking intensity. Furthermore, during the use of deblocking filters, horizontal and vertical filtering can be processed simultaneously when performing vertical and horizontal filtering.

[0071] The offset correction unit can correct the offset between the video being deblocked and the original video on a pixel-by-pixel basis. To perform offset correction on a specific image, the following methods can be used: after dividing the pixels included in the video into a predetermined number of regions, determine the region to be offset and apply the offset to the corresponding region, or apply the offset by taking into account the edge information of each pixel.

[0072] Adaptive Loop Filtering (ALF) can be performed based on a comparison between the filtered reconstructed image and the original video. After dividing the pixels in the video into predetermined groups, filtering can be performed differently for each group by determining a filter to be used for the corresponding group. Information related to whether adaptive loop filtering is applied, along with the luminance signal, can be transmitted per coding unit (CU). The shape and filter coefficients of the adaptive loop filter to be applied can vary depending on the block. Furthermore, it is possible to apply the same type (fixed type) of adaptive loop filter regardless of the characteristics of the block to which it is applied.

[0073] The memory 155 can store the reconstructed blocks or images calculated by the filter unit 150, and can provide the stored reconstructed blocks or images to the prediction units 120 and 125 when performing inter-frame prediction.

[0074] Figure 2 This is a block diagram of a video decoder according to an embodiment of the present invention.

[0075] Reference Figure 2 The video decoder 200 may include an entropy decoding unit 210, a rearrangement unit 215, an inverse quantization unit 220, an inverse transform unit 225, a prediction unit 230, a prediction unit 235, a filter unit 240, and a memory 245.

[0076] When inputting a video bitstream from a video encoder, the input bitstream can be decoded by following the reverse steps of the video encoder.

[0077] The entropy decoding unit 210 can perform entropy decoding in the reverse order of the entropy encoding steps performed in the entropy encoding unit of the video encoder. For example, corresponding to the methods performed in the video encoder, various methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) can be applied.

[0078] The entropy decoding unit 210 can decode information related to intra-frame prediction and inter-frame prediction performed by the encoder.

[0079] The rearrangement unit 215 can perform rearrangement based on a method used in the encoding unit to rearrange the bitstream that has been entropily decoded by the entropy decoding unit 210. Multiple coefficients represented in one-dimensional vector form can be reconstructed into two-dimensional block-shaped coefficients for rearrangement. The rearrangement unit 215 receives information related to the coefficient scan performed in the encoding unit and can perform rearrangement by performing a reverse scan based on the scan order performed in the corresponding encoding unit.

[0080] The inverse quantization unit 220 can perform inverse quantization based on the quantization parameters provided by the encoder and the coefficient values ​​of the rearranged blocks.

[0081] The inverse transform unit 225 can perform inverse discrete cosine transform and inverse discrete sine transform on the quantization result performed by the video encoder. These inverse discrete cosine transforms and inverse discrete sine transforms are inverse transforms of the transforms performed in the transform unit, i.e., inverse transforms of the discrete cosine transform and discrete sine transform. The DCT transform kernel can include at least one of DCT2 or DCT8, and the DST transform kernel can include DST7. Alternatively, if the transform is skipped in the video encoder, the inverse transform unit 225 may not perform the inverse transform. The inverse transform can be performed based on the transmission unit determined in the video encoder. In the inverse transform unit 225 of the video decoder, a transform method (e.g., DCT or DST) can be selectively performed based on multiple pieces of information such as the prediction method, the size of the current block, and the prediction direction.

[0082] Prediction units 230 and 235 can generate prediction blocks based on information related to prediction block generation provided by entropy decoding unit 210 and previously decoded block or image information provided by memory 245.

[0083] As described above, when intra-prediction is performed in the same manner as in the video encoder, if the size of the prediction unit is the same as the size of the transform unit, intra-prediction is performed on the prediction unit based on the pixels to its left, the pixels to its upper left, and the pixels above it. If the size of the prediction unit is different from the size of the transform unit, intra-prediction can be performed using reference pixels based on the transform unit. Furthermore, intra-prediction using only N×N partitioning for the smallest coding unit can also be applied.

[0084] Prediction units 230 and 235 may include a prediction unit determination unit, an inter-frame prediction unit, and an intra-frame prediction unit. The prediction unit determination unit receives various information input from the entropy decoding unit 210, such as prediction unit information, prediction mode information of the intra-frame prediction method, and motion prediction-related information of the inter-frame prediction method. It classifies prediction units according to the current coding unit and determines whether the prediction unit is performing inter-frame prediction or intra-frame prediction. The inter-frame prediction unit 230 can use the information required for inter-frame prediction of the current prediction unit provided by the video encoder and perform inter-frame prediction on the current prediction unit based on information included in at least one of the previous or next images of the current image to which the current prediction unit belongs. Alternatively, inter-frame prediction can also be performed based on information from a portion of the reconstructed region within the current image to which the current prediction unit belongs.

[0085] In order to perform inter-frame prediction, it is possible to determine, based on the coding unit, which of the following modes of motion prediction method is used for the prediction units included in the corresponding coding unit: Skip Mode, Merge Mode, Advanced Motion Vector Prediction Mode (AMVP Mode), or Intra-block Copy Mode.

[0086] The intra-prediction unit 235 can generate prediction blocks based on pixel information within the current image. When the prediction unit is one that has already performed intra-prediction, intra-prediction can be performed based on the intra-prediction mode information of the prediction unit provided by the video encoder. The intra-prediction unit 235 may include an adaptive intra-smoothing (AIS) filter, a reference pixel interpolation unit, and a DC filter. The adaptive intra-smoothing filter is the part that performs filtering on the reference pixels of the current block, and whether to apply the filter can be determined according to the prediction mode of the current prediction unit. Adaptive intra-smoothing filtering can be performed on the reference pixels of the current block using the prediction mode of the prediction unit provided by the video encoder and the adaptive intra-smoothing filter information. If the prediction mode of the current block is a mode that does not perform adaptive intra-smoothing filtering, then the adaptive intra-smoothing filter may not be applied.

[0087] For the reference pixel interpolation unit, if the prediction mode of the prediction unit is a prediction unit that performs intra-frame prediction based on pixel values ​​interpolated from reference pixels, then reference pixels with integer values ​​or smaller pixel units can be generated by interpolating the reference pixels. If the prediction mode of the current prediction unit is a prediction mode that generates prediction blocks without interpolating reference pixels, then interpolation of reference pixels is not required. If the prediction mode of the current block is DC mode, then the DC filter can generate prediction blocks by filtering.

[0088] The reconstructed blocks or images can be provided to the filter unit 240. The filter unit 240 may include a deblocking filter, an offset correction unit, and an ALF.

[0089] Information related to whether to apply a deblocking filter to a corresponding block or image can be received from the video encoder, as well as information related to whether a strong or weak filter is applied when applying the deblocking filter. Information related to the deblocking filter provided by the video encoder can be received from the video decoder's deblocking filter, and deblocking filtering can be performed on the corresponding block at the video decoder.

[0090] The offset correction unit can perform offset correction on the reconstructed video based on the type and amount of offset correction used during video encoding.

[0091] The ALF can be applied to the coding unit based on information provided by the encoder, such as whether the ALF is applied and ALF coefficient information. This ALF information can be provided by including it in a specific parameter set.

[0092] The memory 245 stores the reconstructed image or block, such that the image or block can be used as a reference image or reference block, and the reconstructed image can be provided to the output unit.

[0093] Figure 3 This is a diagram illustrating the basic coding tree unit of an embodiment of the present invention.

[0094] The largest coding block can be defined as the coding tree block. An image can be divided into multiple coding tree units (CTUs). The coding tree unit is the largest coding unit and can also be called the largest coding unit (LCU). Figure 3 An example of dividing an image into multiple coding tree units is shown.

[0095] The size of a coding tree unit can be defined at the image level or the sequence level. Therefore, information representing the size of a coding tree unit can be transmitted via signals using either an image parameter set or a sequence parameter set.

[0096] For example, the size of the coding tree unit for the entire image within the sequence can be set to 128×128. Alternatively, either 128×128 or 256×256 at the image level can be determined as the size of the coding tree unit. For example, the size of the coding tree unit in the first image can be set to 128×128, and the size of the coding tree unit in the second image can be set to 256×256.

[0097] Coded blocks can be generated by dividing the coding tree into units. A coded block represents the basic unit used for encoding / decoding processing. For example, prediction or transformation can be performed on different coded blocks, or prediction coding modes can be determined on different coded blocks. The prediction coding mode represents the method for generating the predicted image. For example, prediction coding modes can include intra-prediction, inter-prediction, current picture referencing (CPR, or intra-block copy (IBC)), or combined prediction. For a coded block, at least one of the prediction coding modes—intra-prediction, inter-prediction, current picture referencing, or combined prediction—can be used to generate the prediction block associated with that coded block.

[0098] Information representing the predictive coding mode of the current block can be transmitted via a bitstream signal. For example, this information could be a 1-bit flag indicating whether the predictive coding mode is intra-frame or inter-frame. Current image reference or combined prediction can be used only if the predictive coding mode of the current block is determined to be inter-frame.

[0099] The current image reference is used to set the current image as the reference image and obtain the prediction block of the current block from the encoded / decoded regions within the current image. Here, the current image means the image that includes the current block. Information indicating whether the current image reference is applied to the current block can be sent via a bitstream signal. For example, this information could be a 1-bit flag. When the flag is true, the prediction coding mode of the current block can be determined as the current image reference; when the flag is false, the prediction mode of the current block can be determined as inter-frame prediction.

[0100] Alternatively, the predictive coding mode for the current block can be determined based on a reference image index. For example, when the reference image index points to the current image, the predictive coding mode for the current block can be determined as current image reference. When the reference image index points to another image instead of the current image, the predictive coding mode for the current block can be determined as inter-frame prediction. That is, current image reference is a prediction method that uses information from already encoded / decoded regions within the current image, and inter-frame prediction is a prediction method that uses information from other encoded / decoded images.

[0101] Combinatorial prediction represents a coding mode composed of two or more of intra-frame prediction, inter-frame prediction, and current image reference. For example, when applying combinatorial prediction, a first prediction block can be generated based on one of intra-frame prediction, inter-frame prediction, or the current image reference, and a second prediction block can be generated based on another. If a first and a second prediction block are generated, a final prediction block can be generated by averaging or weighted summing the first and second prediction blocks. Information indicating whether combinatorial prediction is applied can be transmitted via a bitstream signal. This information can be a 1-bit flag.

[0102] Figures 4(a)-4(e) This is a diagram illustrating the various partitioning types of coded blocks.

[0103] A coded block can be divided into multiple coded blocks based on quadtree partitioning, binary tree partitioning, or ternary tree partitioning. Furthermore, the divided coded blocks can be further divided into multiple coded blocks based on quadtree partitioning, binary tree partitioning, or ternary tree partitioning.

[0104] Quadtree partitioning is a partitioning technique that divides the current block into four blocks. As a result of quadtree partitioning, the current block can be divided into four square partitions (refer to “SPLIT_QT” in part 4(a) of Figure 4).

[0105] Binary tree partitioning refers to a partitioning technique that divides the current block into two blocks. The process of partitioning the current block along a vertical direction (i.e., using a vertical line crossing the current block) is called vertical binary tree partitioning, and the process of partitioning the current block along a horizontal direction (i.e., using a horizontal line crossing the current block) is called horizontal binary tree partitioning. After binary tree partitioning, the current block can be divided into two non-square partitions. "SPLIT_BT_VER" in part 4(b) represents the result of vertical binary tree partitioning, and "SPLIT_BT_HOR" in part 4(c) represents the result of horizontal binary tree partitioning.

[0106] Ternary tree partitioning refers to a partitioning technique that divides the current block into three blocks. The process of dividing the current block into three blocks along a vertical direction (i.e., using two vertical lines crossing the current block) is called vertical ternary tree partitioning, and the process of dividing the current block into three blocks along a horizontal direction (i.e., using two horizontal lines crossing the current block) is called horizontal ternary tree partitioning. After ternary tree partitioning, the current block can be divided into three non-square partitions. In this case, the width / height of the partition located at the center of the current block can be twice the width / height of the other partitions. "SPLIT_TT_VER" in part 4(d) represents the result of vertical ternary tree partitioning, and "SPLIT_TT_HOR" in part 4(e) represents the result of horizontal ternary tree partitioning.

[0107] The number of times a coding tree unit is divided can be defined as the partitioning depth. The maximum partitioning depth of a coding tree unit can be determined at the sequence or image level. Therefore, the maximum partitioning depth of a coding tree unit can vary depending on different sequences or images.

[0108] Alternatively, the maximum partitioning depth can be determined individually for each of the multiple partitioning techniques. For example, the maximum partitioning depth allowed for quadtree partitioning can be different from the maximum partitioning depth allowed for binary tree partitioning and / or ternary tree partitioning.

[0109] The encoder can transmit information representing at least one of the partition shape or partition depth of the current block via a bitstream. The decoder can determine the partition shape and partition depth of the coding tree unit based on the information parsed from the bitstream.

[0110] Figure 5 This is a diagram illustrating the partitioning of coding tree units.

[0111] The process of dividing coding blocks using partitioning techniques such as quadtree partitioning, binary tree partitioning, and / or ternary tree partitioning is called multitree partitioning.

[0112] The coded blocks generated by applying a multi-way tree partitioning to the coded block can be called multiple downstream coded blocks. When the partitioning depth of the coded block is k, the partitioning depth of the multiple downstream coded blocks is set to k+1.

[0113] On the other hand, for multiple coding blocks with a partitioning depth of k+1, the coding block with a partitioning depth of k can be called the upstream coding block.

[0114] The partition type of the current coding block can be determined based on at least one of the partition shape of the upstream coding block or the partition type of the adjacent coding blocks. The adjacent coding blocks are adjacent to the current coding block and can include at least one of the current coding block's upper adjacent block, left adjacent block, or adjacent block to its upper left corner. The partition type can include at least one of whether to partition into a quadtree, whether to partition into a binary tree, the binary tree partition direction, whether to partition into a ternary tree, or the ternary tree partition direction.

[0115] To determine the shape of the coded block partition, information indicating whether the coded block has been partitioned can be sent via a bitstream signal. This information is a 1-bit flag "split_cu_flag," and when the flag is true, it indicates that the coded block has been partitioned using a multi-way tree partitioning technique.

[0116] When "split_cu_flag" is true, information indicating whether the coded block has been partitioned by a quadtree can be sent via a bitstream signal. This information is a 1-bit flag "split_qt_flag". When this flag is true, the coded block can be divided into 4 blocks.

[0117] For example, in Figure 5 The example shown illustrates how the coding tree unit is partitioned by a quadtree to generate four coding blocks with a partition depth of 1. Furthermore, the example illustrates applying quadtree partitioning again to the first and fourth coding blocks generated as a result of the quadtree partitioning. Ultimately, four coding blocks with a partition depth of 2 can be generated.

[0118] Furthermore, a code block with a partition depth of 3 can be generated by applying a quadtree partition to the code block with a partition depth of 2 again.

[0119] When a quadtree partition is not applied to the coded block, it can be determined whether to perform a binary tree partition or a ternary tree partition by considering at least one of the following: the size of the coded block, whether the coded block is located at an image boundary, the maximum partition depth, or the partition shape of adjacent blocks. When it is determined whether to perform a binary tree partition or a ternary tree partition, information indicating the partition direction can be transmitted via a bitstream signal. This information can be a 1-bit flag "mtt_split_cu_vertical_flag". The partition direction (vertical or horizontal) can be determined based on this flag. Alternatively, information indicating whether a binary tree partition or a ternary tree partition is applied to the coded block can be transmitted via a bitstream signal. This information can be a 1-bit flag "mtt_split_cu_binary_flag". The binary tree partition or ternary tree partition can be determined based on this flag.

[0120] For example, in Figure 5The example shown illustrates the application of a vertical binary tree partitioning to a coded block with a partitioning depth of 1, the application of a vertical ternary tree partitioning to the left coded block in the resulting coded block, and the application of a vertical binary tree partitioning to the right coded block.

[0121] Inter-frame prediction refers to using information from the previous image to predict the predictive coding mode of the current block. For example, a block in the previous image that is at the same position as the current block (hereinafter referred to as a collocated block) can be set as the prediction block for the current block. Hereinafter, the prediction block generated based on the block at the same position as the current block will be called a collocated prediction block.

[0122] On the other hand, if an object that existed in the previous image has moved to a different position in the current image, the object's motion can be used to effectively predict the current block. For example, if the direction and size of the object's movement can be known by comparing the previous and current images, the object's motion information can be considered to generate a predicted block (or predicted image) for the current block. Hereinafter, the predicted block generated using motion information can be referred to as a motion prediction block.

[0123] Residual blocks can be generated by subtracting prediction blocks from the current block. In this case, when there is motion of the object, motion prediction blocks can be used instead of co-position prediction blocks, thereby reducing the energy of the residual blocks and improving their compression performance.

[0124] As mentioned above, the process of generating prediction blocks using motion information can be called motion-compensated prediction. In most inter-frame predictions, prediction blocks can be generated based on motion-compensated prediction.

[0125] Motion information may include at least one of motion vectors, reference image indices, prediction directions, or bidirectional weighted indexes. Motion vectors represent the direction and magnitude of an object's movement. Reference image indices specify the reference image for the current block among a list of reference images. Prediction directions refer to any one of unidirectional L0 prediction, unidirectional L1 prediction, or bidirectional prediction (L0 and L1 prediction). Motion information in either the L0 or L1 direction can be used based on the prediction direction of the current block. Bidirectional weighted indexes specify the weights applied to the L0 prediction block and the weights applied to the L1 prediction block.

[0126] Figure 6 This is a flowchart of the inter-frame prediction method according to an embodiment of the present invention.

[0127] refer to Figure 6The inter-frame prediction method includes the following steps: determining the inter-frame prediction mode of the current block (S601); obtaining motion information of the current block according to the determined inter-frame prediction mode (S602); and performing motion compensation prediction on the current block based on the obtained motion information (S603).

[0128] Inter-frame prediction modes represent various techniques used to determine the motion information of the current block, and can include inter-frame prediction modes using translational motion information and inter-frame prediction modes using affine motion information. For example, inter-frame prediction modes using translational motion information can include merging mode and advanced motion vector prediction mode, while inter-frame prediction modes using affine motion information can include affine merging mode and affine motion vector prediction mode. Based on the inter-frame prediction mode, the motion information of the current block can be determined based on neighboring blocks adjacent to the current block or information parsed from the bitstream.

[0129] The following section details the inter-frame prediction method using affine motion information.

[0130] Figure 7 It is a diagram illustrating the nonlinear motion of the object.

[0131] The motion of objects within a video may be non-linear. For example, ... Figure 7 The example shown may involve non-linear motion of the object, such as camera zoom-in, zoom-out, rotation, or affine transformation. When non-linear motion occurs, it is impossible to effectively represent the object's motion using translational motion vectors. Therefore, in parts where non-linear motion occurs, affine motion can be used instead of translational motion, thereby improving coding efficiency.

[0132] Figure 8 This is a flowchart illustrating an inter-frame prediction method based on affine motion according to an embodiment of the present invention.

[0133] Whether to apply an affine motion-based inter-frame prediction technique to the current block can be determined based on information parsed from the bitstream. Specifically, whether to apply an affine motion-based inter-frame prediction technique to the current block can be determined based on at least one of a flag indicating whether an affine merging mode is applied to the current block or a flag indicating whether an affine motion vector prediction mode is applied to the current block.

[0134] When an inter-frame prediction technique based on affine motion is applied to the current block, the affine motion model of the current block can be determined (S801). The affine motion model can be determined by at least one of a 6-parameter affine motion model or a 4-parameter affine motion model. The 6-parameter affine motion model uses 6 parameters to represent the affine motion, while the 4-parameter affine motion model uses 4 parameters to represent the affine motion.

[0135] Equation 1 represents the case where affine motion is expressed using 6 parameters. Affine motion represents translational motion within a predetermined region determined by the affine seed vector.

[0136] Equation 1

[0137] While using six parameters to represent affine motion allows for the representation of complex motions, the increased number of bits required to encode each parameter reduces encoding efficiency. Therefore, affine motion can also be represented using four parameters. Equation 2 illustrates the case of representing affine motion using four parameters.

[0138] Equation 2

[0139] Information used to determine the affine motion model for the current block can be encoded and transmitted via a bitstream signal. For example, this information could be a 1-bit flag, "affine_type_flag". A value of 0 indicates the application of a 4-parameter affine motion model, and a value of 1 indicates the application of a 6-parameter affine motion model. The flag can be encoded in units of stripes, tiles, or blocks (e.g., coded blocks or coded tree units). When the flag is transmitted at the strip level, the affine motion model determined at that strip level can be applied to all blocks belonging to that strip.

[0140] Alternatively, the affine motion model of the current block can be determined based on the affine inter-frame prediction mode of the current block. For example, when applying the affine merging mode, the affine motion model of the current block can be determined as a 4-parameter motion model. On the other hand, when applying the affine motion vector prediction mode, the information used to determine the affine motion model of the current block can be encoded and transmitted as a signal via a bitstream. For example, when applying the affine motion vector prediction mode to the current block, the affine motion model of the current block can be determined based on a 1-bit flag "affine_type_flag".

[0141] Next, the affine seed vector of the current block can be exported (S802). When a 4-parameter affine motion model is selected, motion vectors at two control points of the current block can be exported. On the other hand, when a 6-parameter affine motion model is selected, motion vectors at three control points of the current block can be exported. The motion vectors at the control points can be called affine seed vectors. Control points can include at least one of the upper left, upper right, or lower left corners of the current block.

[0142] Figure 9 This is a diagram showing an example of the affine seed vector for each affine motion model.

[0143] In a 4-parameter affine motion model, two related affine seed vectors can be derived from the top left, top right, or bottom left corners. For example, ... Figure 9 In the example shown in section (a), when the 4-parameter affine motion model is selected, the affine vectors can be derived by using the affine seed vector sv0 associated with the top-left corner of the current block (e.g., the top-left sample (x0, y0)) and the affine seed vector sv1 associated with the top-right corner of the current block (e.g., the top-right sample (x1, y1)). Alternatively, the affine seed vector associated with the bottom-left corner can be used instead of the affine seed vector associated with the top-left corner, or vice versa.

[0144] In a 6-parameter affine motion model, affine seed vectors related to the top-left, top-right, and bottom-left corners can be derived. For example, ... Figure 9 In the example shown in section (b), when the 6-parameter affine motion model is selected, the affine vectors can be derived by using the affine seed vector sv0 associated with the top left corner of the current block (e.g., the top left sample (x0, y0)), the affine seed vector sv1 associated with the top right corner of the current block (e.g., the top right sample (x1, y1)), and the affine seed vector sv2 associated with the top left corner of the current block (e.g., the top left sample (x2, y2)).

[0145] In the embodiments described later, under the 4-parameter affine motion model, the affine seed vectors of the upper left control point and the upper right control point are referred to as the first affine seed vector and the second affine seed vector, respectively. In the embodiments using the first and second affine seed vectors described later, at least one of the first and second affine seed vectors can be replaced by the affine seed vector of the lower left control point (the third affine seed vector) or the affine seed vector of the lower right control point (the fourth affine seed vector).

[0146] Furthermore, in the 6-parameter affine motion model, the affine seed vectors of the upper left control point, upper right control point, and lower left control point are respectively referred to as the first affine seed vector, the second affine seed vector, and the third affine seed vector. In the embodiments using the first, second, and third affine seed vectors described later, at least one of the first, second, and third affine seed vectors can be replaced by the affine seed vector of the lower right control point (the fourth affine seed vector).

[0147] Affine vectors can be derived for different sub-blocks using an affine seed vector (S803). Here, the affine vector represents the translational motion vector derived based on the affine seed vector. The affine vector of a sub-block can be called the affine sub-block motion vector or the sub-block motion vector.

[0148] Figure 10 This is a diagram showing an example of the affine vectors of a sub-block under a 4-parameter motion model.

[0149] The affine vector of a sub-block can be derived based on the position of the control point, the position of the sub-block, and the affine seed vector. For example, Equation 3 shows an example of deriving the affine sub-block vector.

[0150] Equation 3

[0151] In Equation 3, (x, y) represents the position of the sub-block. The position of the sub-block refers to the position of the reference sample included within it. The reference sample can be the sample located at the top left corner of the sub-block, or at least one sample located at the center of the x-axis or y-axis coordinate system. (x0, y0) represents the position of the first control point, and (sv0x, sv0y) represents the first affine seed vector. Additionally, (x1, y1) represents the position of the second control point, and (sv1x, sv1y) represents the second affine seed vector.

[0152] When the first control point and the second control point correspond to the top left corner and the top right corner of the current block, respectively, x1-x0 can be set to the same value as the width of the current block.

[0153] Then, motion compensation prediction can be performed on each sub-block using the affine vectors of each sub-block (S804). After performing motion compensation prediction, prediction blocks associated with each sub-block can be generated. The prediction blocks of the sub-blocks can be set as the prediction blocks of the current block.

[0154] Next, we will explain in detail the inter-frame prediction method that uses translational motion information.

[0155] Motion information for the current block can be derived from the motion information of other blocks. These other blocks can be those that are prioritized for inter-frame prediction encoding / decoding compared to the current block. Setting the motion information of the current block to be the same as that of other blocks is defined as a merging mode. Furthermore, setting the motion vectors of other blocks to the predicted values ​​of the motion vectors of the current block is defined as a motion vector prediction mode.

[0156] Figure 11 This is a flowchart of the process of exporting motion information of the current block in merge mode.

[0157] Merging candidates for the current block can be exported (S1101). Merging candidates for the current block can be exported from blocks that were encoded / decoded using inter-frame prediction before the current block.

[0158] Figure 12 This is a diagram showing an example of a candidate block used to derive merge candidates.

[0159] Candidate blocks can include at least one of the following: neighboring blocks containing samples adjacent to the current block, or non-neighboring blocks containing samples not adjacent to the current block. Hereinafter, the samples used to determine candidate blocks will be designated as reference samples. Furthermore, reference samples adjacent to the current block will be referred to as neighboring reference samples, and reference samples not adjacent to the current block will be referred to as non-neighboring reference samples.

[0160] Adjacent reference samples can be included in the adjacent column of the leftmost column of the current block or the adjacent row of the topmost row of the current block. For example, if the coordinates of the top-left sample of the current block are (0, 0), then at least one of the following blocks—a block including a reference sample at position (-1, H-1), a block including a reference sample at position (W-1, -1), a block including a reference sample at position (W, -1), a block including a reference sample at position (-1, H), or a block including a reference sample at position (-1, -1)—can be used as candidate blocks. Referring to the accompanying drawings, adjacent blocks with indices 0 to 4 can be used as candidate blocks.

[0161] A non-adjacent reference sample refers to a sample whose x-axis distance or y-axis distance to the reference sample adjacent to the current block has a predefined value. For example, a block containing a reference sample whose x-axis distance to the left reference sample is a predefined value, a block containing a non-adjacent sample whose y-axis distance to the upper reference sample is a predefined value, or a block containing non-adjacent samples whose x-axis and y-axis distances to the upper-left reference sample are both predefined values ​​can be used as a candidate block. The predefined value can be an integer such as 4, 8, 12, 16, etc. Referring to the accompanying drawings, at least one of the blocks with indices from 5 to 26 can be used as a candidate block.

[0162] Samples that are not on the same vertical, horizontal, or diagonal line as adjacent reference samples can be set as non-adjacent reference samples.

[0163] Figure 13 This is a diagram showing the location of the reference sample.

[0164] like Figure 13 The example shown allows setting the x-coordinate of a non-adjacent upper reference sample to be different from that of the adjacent upper reference sample. For instance, when the position of the adjacent upper reference sample is (W-1, -1), the position of a non-adjacent upper reference sample that is N away from the adjacent upper reference sample along the y-axis can be set to ((W / 2)-1, -1-N), and the position of a non-adjacent upper reference sample that is 2N away from the adjacent upper reference sample along the y-axis can be set to (0, -1-2N). That is, the position of a non-adjacent reference sample can be determined based on the position of the adjacent reference sample and the distance between them.

[0165] In the following text, a candidate block containing an adjacent reference sample is called a neighboring block, and a block containing a non-adjacent reference sample is called a non-adjacent block.

[0166] When the distance between the current block and a candidate block is greater than or equal to a threshold, the candidate block can be set as unusable as a merging candidate. The threshold can be determined based on the size of the coding tree unit. For example, the threshold can be set to the height of the coding tree unit (ctu_height), or the height of the coding tree unit plus or minus an offset value (e.g., ctu_height ± N). The offset value N is a predefined value in the encoder and decoder, and can be set to 4, 8, 16, 32, or ctu_height.

[0167] If the difference between the y-axis coordinate of the current block and the y-axis coordinate of the samples included in the candidate block is greater than a threshold, the candidate block can be determined as unsuitable for merging.

[0168] Alternatively, candidate blocks that do not belong to the same coding tree unit as the current block can be set as unsuitable for merging. For example, when the reference sample exceeds the upper boundary of the coding tree unit to which the current block belongs, candidate blocks that include the reference sample can be set as unsuitable for merging.

[0169] If the upper boundary of the current block is adjacent to the upper boundary of a coding tree unit, multiple candidate blocks will be determined as unsuitable for merging, which will reduce the encoding / decoding efficiency of the current block. To resolve this issue, candidate blocks can be configured such that the number of candidate blocks above the current block is greater than the number of candidate blocks to the left of the current block.

[0170] Figure 14This is a diagram showing an example of a candidate block used to derive merge candidates.

[0171] like Figure 14 The example shown allows setting the top block of the N blocks above the current block and the left block of the M blocks to the left of the current block as candidate blocks. In this case, by setting M to be greater than N, the number of left candidate blocks can be set to be greater than the number of top candidate blocks.

[0172] For example, the difference between the y-axis coordinate of the reference sample within the current block and the y-axis coordinate of the block above which can be used as a candidate block can be set to no more than N times the height of the current block. Additionally, the difference between the x-axis coordinate of the reference sample within the current block and the x-axis coordinate of the block to the left of which can be used as a candidate block can be set to no more than M times the width of the current block.

[0173] For example, such as Figure 14 The example shown illustrates setting the blocks belonging to the two blocks above the current block and the five blocks belonging to the left of the current block as candidate blocks.

[0174] As another example, when a candidate block does not belong to the same coding tree unit as the current block, a merge candidate can be derived by using a block that belongs to the same coding tree unit as the current block, or a block that contains a reference sample adjacent to the boundary of the coding tree unit, instead of the candidate block.

[0175] Figure 15 This is a diagram illustrating an example of changing the position of a reference sample.

[0176] When a reference sample is included in a coding tree unit that is different from the current block, and the reference sample is not adjacent to the boundary of the coding tree unit, a candidate block reference sample can be determined by using a reference sample adjacent to the boundary of the coding tree unit instead of the reference sample.

[0177] For example, in Figure 15 (a) and Figure 15 In the example shown in (b), when the upper boundary of the current block touches the upper boundary of the coding tree unit, the reference sample above the current block belongs to a coding tree unit different from the current block. A reference sample belonging to a coding tree unit different from the current block that is not adjacent to the upper boundary of the coding tree unit can be replaced with a sample adjacent to the upper boundary of the coding tree unit.

[0178] For example, such as Figure 15 As shown in example (a), the reference sample at position 6 is replaced with the sample at position 6', which is located at the upper boundary of the coding tree unit, as follows. Figure 15As shown in example (b), the reference sample at position 15 is replaced with the sample at position 15', which is located at the upper boundary of the coding tree unit. In this case, the y-coordinate of the replacement sample can be changed to that of an adjacent position in the coding tree unit, and the x-coordinate of the replacement sample can be set to be the same as that of the reference sample. For example, the sample at position 6' can have the same x-coordinate as the sample at position 6, and the sample at position 15' can have the same x-coordinate as the sample at position 15.

[0179] Alternatively, the x-coordinate of the replacement sample can be set by adding or subtracting the offset value from the x-coordinate of the reference sample. For example, when the x-coordinates of adjacent and non-adjacent reference samples above the current block are the same, the x-coordinate of the replacement sample can be set by adding or subtracting the offset value from the x-coordinate of the reference sample. This is to prevent the replacement sample used to replace a non-adjacent reference sample from being in the same position as other non-adjacent or adjacent reference samples.

[0180] Figure 16 This is a diagram illustrating an example of changing the position of a reference sample.

[0181] When replacing a reference sample that is included in a different coding tree unit than the current block and is not adjacent to the boundary of the coding tree unit with a sample located at the boundary of the coding tree unit, the x-coordinate of the replacement sample can be set by adding or subtracting the offset value to the x-coordinate of the reference sample.

[0182] For example, in Figure 16 In the example shown, the reference sample at position 6 and the reference sample at position 15 can be replaced with the sample at position 6' and the sample at position 15', respectively, having the same y-coordinate as the row adjacent to the upper boundary of the coding tree unit. In this case, the x-coordinate of the sample at position 6' can be set to a value where the difference between its x-coordinate and the x-coordinate of the reference sample at position 6 is W / 2, and the x-coordinate of the sample at position 15' can be set to a value where the difference between its x-coordinate and the x-coordinate of the reference sample at position 15 is W-1.

[0183] Unlike Figure 15 and Figure 16 In the example shown, the y-coordinate of the row above the top row of the current block or the y-coordinate of the upper boundary of the coding tree unit can also be set to the y-coordinate of the replacement sample.

[0184] Although not illustrated, the sample replacing the reference sample can also be determined based on the left boundary of the coding tree unit. For example, when the reference sample is not included in the same coding tree unit as the current block and is not adjacent to the left boundary of the coding tree unit, the reference sample can be replaced with a sample adjacent to the left boundary of the coding tree unit. In this case, the replacement sample can have the same y-coordinate as the reference sample, or it can have a y-coordinate obtained by adding or subtracting an offset value to the y-coordinate of the reference sample.

[0185] Then, the block containing the replacement sample can be set as a candidate block, and the merge candidates for the current block can be derived based on the candidate blocks.

[0186] Merge candidates can also be derived from temporally adjacent blocks included in images different from the current block. For example, merge candidates can be derived from co-located blocks included in co-located images.

[0187] The motion information of the merged candidate can be set to be the same as that of the candidate block. For example, at least one of the motion vector, reference image index, prediction direction, or bidirectional weighted index of the candidate block can be set as the motion information of the merged candidate.

[0188] A list of merge candidates, including merge candidates, can be generated (S1102). The merge candidates can be classified into adjacent merge candidates derived from adjacent blocks adjacent to the current block, and non-adjacent merge candidates derived from non-adjacent blocks.

[0189] The indices of multiple merge candidates within the merge candidate list can be assigned in a predetermined order. For example, the index assigned to an adjacent merge candidate can have a smaller value than the index assigned to a non-adjacent merge candidate. Alternatively, based on... Figure 12 or Figure 14 The index shown for each block can be assigned to each merge candidate.

[0190] When the merge candidate list includes multiple merge candidates, at least one of the multiple merge candidates can be selected (S1103). At this time, information indicating whether the motion information of the current block is derived from adjacent merge candidates can be sent via a bit stream signal. The information can be a 1-bit flag. For example, the syntax element isAdjancentMergeFlag indicating whether the motion information of the current block is derived from adjacent merge candidates can be sent via a bit stream signal. When the value of the syntax element isAdjancentMergeFlag is 1, the motion information of the current block can be derived based on adjacent merge candidates. On the other hand, when the value of the syntax element isAdjancentMergeFlag is 0, the motion information of the current block can be derived based on non-adjacent merge candidates.

[0191] Table 1 shows the syntax table including the syntax element isAdjancentMergeFlag.

[0192] Table 1

[0193]

[0194]

[0195] Information specifying any one of multiple merge candidates can be transmitted via a bitstream signal. For example, information indicating the index of any merge candidate included in the merge candidate list can be transmitted via a bitstream signal.

[0196] When isAdjacentMergeflag is 1, the syntax element merge_idx can be signaled to determine which of the adjacent merge candidates is being merged. The maximum value of the syntax element merge_idx can be set to a value that is 1 greater than the difference between the number of adjacent merge candidates.

[0197] When isAdjacentMergeflag is 0, the syntax element NA_merge_idx can be signaled to determine any of the non-adjacent merge candidates. The syntax element NA_merge_idx indicates the value obtained by subtracting the index of the non-adjacent merge candidate from the number of adjacent merge candidates. The decoder can select a non-adjacent merge candidate by adding the number of adjacent merge candidates to the index determined by NA_merge_idx.

[0198] When the number of merge candidates in the merge candidate list is less than a threshold, merge candidates included in the inter-frame motion information list can be added to the merge candidate list. The threshold can be a value calculated from the maximum number of merge candidates the merge candidate list can include, or the maximum number of merge candidates minus an offset. The offset can be an integer such as 1 or 2. The inter-frame motion information list can include merge candidates derived based on blocks encoded / decoded prior to the current block.

[0199] The inter-frame motion information list includes merging candidates derived from blocks encoded / decoded based on inter-frame prediction within the current image. For example, the motion information of the merging candidates included in the inter-frame motion information list can be set to be the same as the motion information of the blocks encoded / decoded based on inter-frame prediction. The motion information can include at least one of motion vectors, reference image indices, prediction directions, or bidirectional weighted indexes. For ease of explanation, the merging candidates included in the inter-frame motion information list are referred to as inter-frame merging candidates.

[0200] When a merge candidate is selected for the current block, the motion vector of the selected merge candidate is set as the initial motion vector, and motion compensation prediction for the current block can be performed using the motion vector derived by adding or subtracting the offset vector from the initial motion vector. Deriving a new motion vector by adding or subtracting the offset vector from the motion vector of the merge candidate can be defined as a merge motion difference coding method.

[0201] Information indicating whether the merge offset encoding method is used can be transmitted via a bitstream signal. This information can be a 1-bit flag, `merge_offset_vector_flag`. For example, a value of 1 for `merge_offset_vector_flag` indicates that the merge motion interpolation encoding method is applied to the current block. When the merge motion interpolation encoding method is applied to the current block, the motion vector of the current block can be derived by adding or subtracting the offset vector from the motion vector of the merge candidate. A value of 0 for `merge_offset_vector_flag` indicates that the merge motion interpolation encoding method is not applied to the current block. When the merge offset encoding method is not applied, the motion vector of the merge candidate can be set as the motion vector of the current block.

[0202] The flag can only be signaled when the value of the skip flag indicating whether to apply the skip mode or the value of the merge flag indicating whether to apply the merge mode. For example, when the value of skip_flag indicating whether to apply the skip mode to the current block is 1, or when the value of merge_flag indicating whether to apply the merge mode to the current block is 1, merge_offset_vector_flag can be encoded and signaled.

[0203] When it is determined that the merge offset encoding method will be applied to the current block, at least one of the following can be signaled: information specifying any of the merge candidates included in the merge candidate list, information indicating the magnitude of the offset vector, and information indicating the direction of the offset vector.

[0204] Information used to determine the maximum number of merge candidates that can be included in the merge candidate list can be transmitted via a bit stream using signals. For example, the maximum number of merge candidates that can be included in the merge candidate list can be set to an integer of 6 or less.

[0205] When it is determined that the merge offset encoding method will be applied to the current block, only a preset maximum number of merge candidates can be set as the initial motion vector of the current block. That is, the number of merge candidates that the current block can use can be adaptively determined depending on whether the merge offset encoding method is applied. For example, when the value of `merge_offset_vector_flag` is set to 0, the maximum number of merge candidates that the current block can use can be set to M, while when the value of `merge_offset_vector_flag` is set to 1, the maximum number of merge candidates that the current block can use can be set to N. Here, M represents the maximum number of merge candidates that the merge candidate list can include, and N represents an integer equal to or less than M.

[0206] For example, when M is 6 and N is 2, the two merge candidates with the smallest indices in the merge candidate list can be set as available for the current block. Therefore, the motion vector of the merge candidate with index 0 or the motion vector of the merge candidate with index 1 can be set as the initial motion vector for the current block. When M and N are the same (e.g., when M and N are 2), all merge candidates in the merge candidate list can be set as available for the current block.

[0207] Alternatively, whether adjacent blocks can be used as merge candidates can be determined based on whether the merge motion interpolation coding method is applied to the current block. For example, when the value of merge_offset_vector_flag is 1, at least one of the adjacent blocks adjacent to the top-right corner, the adjacent blocks adjacent to the bottom-left corner, and the adjacent blocks adjacent to the bottom-left corner of the current block can be set as unusable as merge candidates. Therefore, when the merge motion interpolation coding method is applied to the current block, the motion vectors of at least one of the adjacent blocks adjacent to the top-right corner, the adjacent blocks adjacent to the bottom-left corner, and the adjacent blocks adjacent to the bottom-left corner of the current block cannot be set as the initial motion vector. Alternatively, when the value of merge_offset_vector_flag is 1, the temporal adjacent blocks of the current block can be set as unusable as merge candidates.

[0208] When the merge motion difference encoding method is applied to the current block, it can be configured not to use at least one of the pairwise merge candidates and the zero merge candidate. Therefore, when the value of merge_offset_vector_flag is 1, even if the number of merge candidates included in the merge candidate list is less than the maximum number, at least one of the pairwise merge candidates or the zero merge candidate may not be added to the merge candidate list.

[0209] The motion vector of a merge candidate can be set as the initial motion vector of the current block. In this case, when there are multiple merge candidates available for the current block, information specifying any one of the multiple merge candidates can be sent via a bitstream signal. For example, when the maximum number of merge candidates that the merge candidate list can include is greater than one, information indicating any one of the multiple merge candidates, `merge_idx`, can be sent via a bitstream signal. That is, under the merge offset encoding method, a merge candidate can be specified using the information `merge_idx` to specify any one of the multiple merge candidates. The initial motion vector of the current block can be set as the motion vector of the merge candidate indicated by `merge_idx`.

[0210] On the other hand, when the number of merge candidates available for the current block is one, the signaling of information specifying the merge candidate can be omitted. For example, when the maximum number of merge candidates that the merge candidate list can include is no greater than one, the signaling of the merge_idx information specifying the merge candidate can be omitted. That is, under the merge offset encoding method, when a merge candidate is included in the merge candidate list, the encoding of the merge_idx information specifying the merge candidate can be omitted, and the initial motion vector can be determined based on the merge candidates included in the merge candidate list. The motion vector of the merge candidate can be set as the initial motion vector of the current block.

[0211] As another example, after determining the merge candidates for the current block, it can be determined whether to apply the merge motion interpolation coding method to the current block. For example, when the maximum number of merge candidates that the merge candidate list can include is greater than 1, the information merge_idx, specifying any of the merge candidates, can be signaled. After selecting a merge candidate based on merge_idx, the merge_offset_vector_flag indicating whether to apply the merge motion interpolation coding method to the current block can be decoded. Table 2 is a syntax table illustrating the embodiment described above.

[0212] Table 2

[0213]

[0214]

[0215] As another example, after determining the merge candidates for the current block, it can be determined whether to apply the merge motion interpolation (MIO) method to the current block only if the index of the determined merge candidate is less than the maximum number of merge candidates that can be used when applying the merge motion interpolation method. For example, the merge_offset_vector_flag, which indicates whether to apply the merge motion interpolation method to the current block, can only be encoded and signaled if the value of the index information merge_idx is less than N. When the value of the index information merge_idx is equal to or greater than N, the encoding of merge_offset_vector_flag can be omitted. If the encoding of merge_offset_vector_flag is omitted, it can be determined that the merge motion interpolation method has not been applied to the current block.

[0216] Alternatively, after determining the merge candidates for the current block, it is possible to consider whether the determined merge candidates have bidirectional or unidirectional motion information to determine whether to apply the merge motion difference encoding method to the current block. For example, the merge_offset_vector_flag indicating whether to apply the merge motion difference encoding method to the current block is encoded and signaled only if the value of the index information merge_idx is less than N and the merge candidate selected by the index information has bidirectional motion information. Alternatively, the merge_offset_vector_flag indicating whether to apply the merge motion difference encoding method to the current block is encoded and signaled only if the value of the index information merge_idx is less than N and the merge candidate selected by the index information has unidirectional motion information.

[0217] Alternatively, the decision to apply the merge motion interpolation coding method can be based on at least one of the following: the size of the current block, the shape of the current block, and whether the current block is in contact with the boundary of a coding tree unit. When at least one of the following conditions is not met, the encoding of the merge_offset_vector_flag indicating whether to apply the merge motion interpolation coding method to the current block can be omitted.

[0218] When a merge candidate is selected, its motion vector can be set as the initial motion vector of the current block. Then, information indicating the magnitude and direction of the offset vector can be decoded to determine the offset vector. The offset vector can have a horizontal or vertical component.

[0219] The information indicating the magnitude of the offset vector can be index information indicating any of the motion size candidates. For example, the index information distance_idx indicating any of the motion size candidates can be signaled via a bitstream. Table 3 shows the binarization of the index information distance_idx and the value of the variable DistFromMergeMV used to determine the magnitude of the offset vector based on distance_idx.

[0220] Table 3

[0221]

[0222] The size of the offset vector can be derived by dividing the variable DistFromMergeMV by a preset value. Equation 4 shows an example of determining the size of the offset vector.

[0223] Equation 4

[0224] According to Equation 4, the value obtained by dividing the variable DistFromMergeMV by 4 or by shifting the variable DistFromMergeMV to the left by 2 can be set as the size of the offset vector.

[0225] More or fewer motion size candidates can be used compared to the examples shown in Table 3, or the range of motion vector offset size candidates can be set differently from the examples shown in Table 3. For example, the magnitude of the horizontal or vertical component of the offset vector can be set to no more than two sample distances. Table 4 shows the binarization of the index information distance_idx and the value of the variable DistFromMergeMV used to determine the magnitude of the offset vector based on distance_idx.

[0226] Table 4

[0227]

[0228] Alternatively, the range of candidate motion vector offset sizes can be set differently based on the motion vector precision. For example, when the motion vector precision of the current block is fractional-pixel, the value of the variable DistFromMergeMV corresponding to the value of the index information distance_idx can be set to 1, 2, 4, 8, 16, etc. Here, fractional pixels include at least one of 1 / 16 pixel, one-eighth pixel, one-quarter pixel, or half pixel. On the other hand, when the motion vector precision of the current block is integer pixels, the value of the variable DistFromMergeMV corresponding to the value of the index information distance_idx can be set to 4, 8, 16, 32, 64, etc. That is, the table used to determine the variable DistFromMergeMV can be set differently depending on the motion vector precision of the current block.

[0229] For example, when the motion vector precision of the current block or merge candidate is a quarter pixel, the variable DistFromMergeMV, represented by distance_idx, can be derived using Table 3. On the other hand, when the motion vector precision of the current block or merge candidate is an integer pixel, the value of the variable DistFromMergeMV can be derived by taking N times (e.g., 4 times) the value of the variable DistFromMergeMV indicated by distance_idx in Table 3.

[0230] Information used to determine the motion vector precision can be transmitted via bitstream signals. For example, information can be transmitted at the sequence level, image level, slice level, or block level. Therefore, the range of motion size candidates can be set differently based on the motion vector precision-related information transmitted via bitstream signals. Alternatively, the motion vector precision can be determined based on the merging candidates of the current block. For example, the motion vector precision of the current block can be set to the same as the motion vector precision of the merging candidates.

[0231] Alternatively, information for determining the search range of the offset vector can be transmitted via a bitstream using signals. At least one of the following can be determined based on the search range: the number of motion size candidates, the minimum among the motion size candidates, and the maximum among the motion size candidates. For example, a flag `merge_offset_vector_flag` for determining the search range of the offset vector can be transmitted via a bitstream using signals. This information can be transmitted via a sequence header, image header, or stripe header using signals.

[0232] For example, when the value of merge_offset_extend_range_flag is 0, the size of the offset vector can be set to no more than 2. Therefore, the maximum value of DistFromMergeMV can be set to 8. On the other hand, when the value of merge_offset_extend_range_flag is 1, the size of the offset vector can be set to no more than 32 sample distances. Therefore, the maximum value of DistFromMergeMV can be set to 128.

[0233] The size of the offset vector can be determined using a flag indicating whether its size is greater than a threshold. For example, the flag `distance_flag` indicating whether the offset vector's size is greater than a threshold can be signaled via a bitstream. The threshold can be 1, 2, 4, 8, or 16. For example, a `distance_flag` of 1 indicates that the offset vector's size is greater than 4. On the other hand, a `distance_flag` of 0 indicates that the offset vector's size is 4 or less.

[0234] When the size of the offset vector is greater than the threshold, the difference between the offset vector size and the threshold can be derived using the index information distance_idx. Alternatively, when the size of the offset vector is less than or equal to the threshold, the size of the offset vector can be determined using the index information distance_idx. Table 5 is a syntax table illustrating the encoding process of distance_flag and distance_idx.

[0235] Table 5

[0236]

[0237]

[0238] Equation 5 shows an example of using distance_flag and distance_idx to derive the variable DistFromMergeMV for determining the size of the offset vector.

[0239] Equation 5

[0240] In Equation 5, the value of distance_flag can be set to 1 or 0. The value of distance_idx can be set to 1, 2, 4, 8, 16, 32, 64, 128, etc. N represents a coefficient determined by the threshold. For example, when the threshold is 4, N can be set to 16.

[0241] The information indicating the direction of the offset vector can be index information indicating any of the vector direction candidates. For example, the index information direction_idx indicating any of the vector direction candidates can be transmitted as a signal via a bitstream. Table 6 shows the binarization of the index information direction_idx and the direction of the offset vector according to direction_idx.

[0242] Table 6

[0243]

[0244] In Table 6, sign[0] indicates the horizontal direction, and sign[1] indicates the vertical direction. +1 indicates that the x-component or y-component of the offset vector is positive (+), and -1 indicates that the x-component or y-component of the offset vector is negative (-). Equation 6 shows an example of determining the offset vector based on its magnitude and direction.

[0245] Equation 6

[0246] In Equation 6, offsetMV[0] indicates the vertical component of the offset vector, and offsetMV[1] indicates the horizontal component of the offset vector.

[0247] Figure 17 This is a graph showing the offset vector based on the values ​​of distance_idx, which indicates the magnitude of the offset vector, and direction_idx, which indicates the direction of the offset vector.

[0248] As in Figure 17 In the example shown, the magnitude and direction of the offset vector can be determined based on the values ​​of distance_idx and direction_idx. The maximum size of the offset vector can be set to not exceed a threshold. Here, the threshold can have a value predefined by the encoder and decoder. For example, the threshold could be 32 sample distances. Alternatively, the threshold can be determined based on the magnitude of the initial motion vector. For example, the horizontal threshold can be set based on the magnitude of the horizontal component of the initial motion vector, and the vertical threshold can be set based on the magnitude of the vertical component of the initial motion vector.

[0249] When the merging candidate has bidirectional motion information, the L0 motion vector of the merging candidate can be set as the L0 initial motion vector of the current block, and the L1 motion vector of the merging candidate can be set as the L1 initial motion vector of the current block. In this case, the L0 offset vector and L1 offset vector can be determined by taking into account the output order difference (hereinafter referred to as L0 difference) between the L0 reference image of the merging candidate and the current image and the output order difference (hereinafter referred to as L1 difference) between the L1 reference image of the merging candidate and the current image.

[0250] First, when the L0 and L1 differences have the same sign, the L0 and L1 offset vectors can be set to be the same. On the other hand, when the L0 and L1 differences have different signs, the L1 offset vector can be set in the opposite direction to the L0 offset vector.

[0251] The sizes of the L0 offset vector and the L1 offset vector can be set to be the same. Alternatively, the size of the L1 offset vector can be determined by scaling the L0 offset vector based on the L0 difference and the L1 difference.

[0252] For example, Equation 7 shows the L0 offset vector and L1 offset vector when the signs of the L0 difference and L1 difference are the same.

[0253] Equation 7

[0254] In Equation 7, offsetMVL0[0] indicates the horizontal component of the L0 offset vector, and offsetMVL0[1] indicates the vertical component of the L0 offset vector. offsetMVL1[0] indicates the horizontal component of the L1 offset vector, and offsetMVL1[1] indicates the vertical component of the L1 offset vector.

[0255] Equation 8 shows the L0 offset vector and L1 offset vector when the signs of the L0 difference and L1 difference are different.

[0256] Equation 8

[0257] More than four vector direction candidates can also be defined. Tables 7 and 8 show examples of defining eight vector direction candidates.

[0258] Table 7

[0259]

[0260] Table 8

[0261]

[0262] In Tables 7 and 8, an absolute value greater than 0 for sign[0] and sign[1] indicates that the offset vector is in the diagonal direction. When using Table 6, the magnitudes of the x-axis and y-axis components of the diagonal offset vector are set to abs(offsetMV), while when using Table 7, the magnitudes of the x-axis and y-axis components of the diagonal offset vector are set to abs(offsetMV / 2).

[0263] Figure 18This is a graph showing the offset vector based on the values ​​of distance_idx, which indicates the magnitude of the offset vector, and direction_idx, which indicates the direction of the offset vector.

[0264] Figure 18 (a) is an example of applying Table 6, and Figure 18 (b) is an example of applying Table 7.

[0265] Information for determining at least one of the number or size of vector direction candidates can be transmitted via a bitstream signal. For example, the flag `merge_offset_direction_range_flag` for determining vector direction candidates can be transmitted via a bitstream signal. The flag can be transmitted at the sequence level, image level, or strip level. For example, when the flag value is 0, the four vector direction candidates illustrated in Table 6 can be used. On the other hand, when the flag value is 1, the eight vector direction candidates illustrated in Table 7 or Table 8 can be used.

[0266] Alternatively, at least one of the number or size of vector direction candidates can be determined based on the magnitude of the offset vector. For example, when the value of the variable DistFromMergeMV, used to determine the magnitude of the offset vector, is equal to or less than a threshold, the eight vector direction candidates illustrated in Table 7 or Table 8 can be used. On the other hand, when the value of the variable DistFromMergeMV is greater than the threshold, the four vector direction candidates illustrated in Table 6 can be used.

[0267] Alternatively, at least one of the number or size of vector direction candidates can be determined based on the values ​​of the x-component MVx and the y-component MVy of the initial motion vector. For example, when the difference or absolute value of the difference between MVx and MVy is less than or equal to a threshold, the eight vector direction candidates illustrated in Table 7 or Table 8 can be used. On the other hand, when the difference or absolute value of the difference between MVx and MVy is greater than a threshold, the four vector direction candidates illustrated in Table 6 can be used.

[0268] The motion vector of the current block can be derived by adding the offset vector to the initial motion vector. Equation 9 shows an example of determining the motion vector of the current block.

[0269] Equation 9

[0270] In Equation 9, mvL0 indicates the L0 motion vector of the current block, and mvL1 indicates the L1 motion vector of the current block. mergeMVL0 indicates the initial L0 motion vector of the current block (i.e., merging candidate L0 motion vectors), and mergeMVL1 indicates the initial L1 motion vector of the current block. [0] indicates the horizontal component of the motion vector, and [1] indicates the vertical component of the motion vector.

[0271] Intra-frame prediction uses reconstructed samples that have already been encoded / decoded from the surrounding blocks to predict the current block. In this case, intra-frame prediction of the current block can use reconstructed samples before the application of the in-loop filter.

[0272] Intra-prediction techniques include matrix-based intra-prediction and general intra-prediction that takes into account the directionality with surrounding reconstructed samples. Information indicating the intra-prediction technique for the current block can be signaled via a bitstream. This information may be a 1-bit flag. Alternatively, the intra-prediction technique for the current block can be determined based on at least one of the intra-prediction techniques of the current block's position, size, shape, or neighboring blocks. For example, when the current block crosses an image boundary, the current block is set not to apply matrix-based intra-prediction.

[0273] Matrix-based intra-frame prediction is a method that obtains the predicted block for the current block by performing matrix multiplication between the matrices stored in the encoder and decoder and the reconstructed samples surrounding the current block. Information specifying any one of the stored matrices can be sent via a bitstream signal. The decoder can then determine the matrix for intra-frame prediction of the current block based on this information and the size of the current block.

[0274] Intra-frame prediction is a method that uses either non-angular intra-frame prediction mode or angular intra-frame prediction mode to obtain the prediction block associated with the current block.

[0275] The residual image can be derived by subtracting the predicted image from the original image. In this case, when the residual image is transformed into the frequency domain, even if high-frequency components are removed, the subjective image quality of the video is not significantly degraded. Therefore, reducing the value of high-frequency components or setting the value of high-frequency components to 0 can improve compression efficiency without causing significant visual distortion. Reflecting these characteristics, the current block can be transformed to decompose the residual image into 2D frequency components. This transformation can be performed using transformation techniques such as Discrete Cosine Transform (DCT) or Discrete Sine Transform (DST).

[0276] After transforming the current block using DCT or DST, the transformed current block can be transformed again. In this case, the DCT- or DST-based transformation can be defined as the first transformation, and the process of transforming the block again using the first transformation can be called the second transformation.

[0277] The first transform can be performed using any of a number of transform kernel candidates. For example, the first transform can be performed using any of DCT2, DCT8, or DCT7.

[0278] Different transform cores can be used for the horizontal and vertical directions. Information representing combinations of horizontal and vertical transform cores can also be transmitted as signals via bitstreams.

[0279] The execution units for the first and second transformations will be different. For example, the first transformation can be performed on an 8×8 block, and the second transformation can be performed on the 4×4 sub-blocks within the transformed 8×8 block. In this case, the transformation coefficients of the remaining regions where the second transformation is not performed can also be set to 0.

[0280] Alternatively, a first transformation can be performed on a 4×4 block, and a second transformation can be performed on an 8×8 region of the 4×4 block that includes the transformation.

[0281] Information indicating whether to perform the second transformation can be sent via a bitstream signal.

[0282] The decoder can perform the inverse of the second transform (second inverse transform), and the result of the second transform can be subjected to the inverse of the first transform (first inverse transform). The residual signal of the current block can be obtained as the result of the execution of the second inverse transform and the first inverse transform.

[0283] Quantization is used to reduce the energy of the block, and the quantization process involves dividing the transformation coefficients by a specific constant. This constant can be derived from quantization parameters, which can be defined as values ​​between 1 and 63.

[0284] If transform and quantization are performed in the encoder, the decoder can obtain the residual block through inverse quantization and inverse transform. The decoder then adds the predicted block and the residual block together to obtain the reconstructed block of the current block.

[0285] If a reconstructed block of the current block is obtained, in-loop filtering can be used to reduce information loss during quantization and encoding. In-loop filters can include at least one of a deblocking filter, a sample adaptive offset filter (SAO), or an adaptive loop filter (ALF).

[0286] Embodiments described with a focus on the decoding or encoding process are also included within the scope of this invention. Variations of multiple embodiments described in a predetermined order, in a different order than those described, are also included within the scope of this invention.

[0287] The embodiments have been described based on a series of steps or flowcharts, but this does not limit the chronological order of the invention, and they can be performed simultaneously or in a different order as needed. Furthermore, in the above embodiments, the structural elements constituting the block diagrams (e.g., units, modules, etc.) can also be implemented as hardware devices or software, and multiple structural elements can be combined to implement a single hardware device or software. The embodiments can be implemented in the form of program instructions, which can be executed by various computer components and recorded in a computer-readable recording medium. The computer-readable recording medium can individually or in combination include program instructions, data files, data structures, etc. Examples of computer-readable recording media can include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floppy optical disks; and hardware devices specifically configured to store and execute program instructions, such as ROMs, RAMs, and flash memory. The hardware device can be configured to operate as one or more software modules to perform the processing according to the invention, and vice versa.

[0288] [Industrial Applicability]

[0289] This invention can be applied to electronic devices that encode / decode video.

Claims

1. A video decoding method, comprising the following steps: Generate a list of candidate blocks to merge; The merge candidate for the current block is determined from the merge candidates included in the merge candidate list; The offset vector of the current block is derived, and the size of the offset vector is determined based on the first index information, wherein the first index information is used to indicate any of the motion size candidates; The step of determining the size of the offset vector based on the first index information includes: The magnitude of the offset vector is determined by left-shifting the candidate motion size indicated by the first index information by two binary bits; and The motion vector of the current block is determined by adding the offset vector to the motion vector of the merging candidate. Specifically, at least one of the maximum or minimum values ​​of the motion size candidate is determined differently depending on the motion vector accuracy of the current block.

2. The video decoding method according to claim 1, wherein, At least one of the maximum or minimum values ​​of the motion size candidate is set differently depending on the value of the flag indicating the range of the motion size candidate.

3. The video decoding method according to claim 2, wherein, The flag is transmitted as an image-level signal.

4. The video decoding method according to claim 2, wherein, The direction of the offset vector is determined based on the second index information of any of the candidate vector directions.

5. The video decoding method according to claim 1, wherein, Different ranges of candidate motion vector offset sizes are set based on the different accuracy of the motion vector.

6. The video decoding method according to claim 1, wherein, When the motion vector precision of the current block is a fraction of pixels, the variable value corresponding to the first index information is set to 1, 2, 4, 8, or 16.

7. The video decoding method according to claim 1, wherein, When the motion vector precision of the current block is an integer pixel, the variable value corresponding to the first index information is set to 4, 8, 16, 32 or 64.

8. A video encoding method, comprising the following steps: Generate a list of candidate blocks to merge; Select a merge candidate for the current block from the merge candidates included in the merge candidate list; The offset vector of the current block is derived, and the size of the offset vector is determined based on the first index information, wherein the first index information is used to indicate any of the motion size candidates; The step of determining the size of the offset vector based on the first index information includes: The magnitude of the offset vector is determined by left-shifting the candidate motion size indicated by the first index information by two binary bits; and The motion vector of the current block is determined by adding the offset vector to the motion vector of the merging candidate. Specifically, at least one of the maximum or minimum values ​​of the motion size candidate is determined differently depending on the motion vector accuracy of the current block.

9. The video encoding method according to claim 8, wherein, A flag indicating the range of motion size candidates is encoded, wherein at least one of the maximum or minimum values ​​of the motion size candidates is set differently depending on the value of the flag.

10. The video encoding method according to claim 9, wherein, The logo is encoded at the image level.

11. The video encoding method according to claim 8, wherein, The motion size candidate has a value derived by shifting the magnitude of the offset vector.

12. The video encoding method according to claim 8, wherein, The second index information is encoded, which is used to specify a vector direction candidate that indicates the direction of the offset vector among a plurality of vector direction candidates.

13. A video decoding device, comprising: Inter-frame prediction unit, the inter-frame prediction unit being used for: Generate a list of candidate blocks to merge; The merge candidate for the current block is determined from the merge candidates included in the merge candidate list; The offset vector of the current block is derived, and the size of the offset vector is determined based on the first index information, wherein the first index information is used to indicate any of the motion size candidates; The inter-frame prediction unit determines the magnitude of the offset vector based on the first index information, including: The inter-frame prediction unit determines the magnitude of the offset vector based on a left-shifting operation of two binary bits on the motion magnitude candidate value indicated by the first index information; and The inter-frame prediction unit derives the motion vector of the current block by adding the offset vector to the motion vector of the merging candidate. The inter-frame prediction unit sets at least one of the maximum or minimum values ​​of the motion size candidates differently depending on the motion vector accuracy of the current block.

14. A video encoding device, comprising: Inter-frame prediction unit, the inter-frame prediction unit being used for: Generate a list of candidate blocks to merge; The merge candidate for the current block is determined from the merge candidates included in the merge candidate list; The offset vector of the current block is derived, and the size of the offset vector is determined based on the first index information, wherein the first index information is used to indicate any of the motion size candidates; The inter-frame prediction unit determines the magnitude of the offset vector based on the first index information, including: The inter-frame prediction unit determines the magnitude of the offset vector based on a left-shifting operation of two binary bits on the motion magnitude candidate value indicated by the first index information; and The inter-frame prediction unit derives the motion vector of the current block by adding the offset vector to the motion vector of the merging candidate. The inter-frame prediction unit sets at least one of the maximum or minimum values ​​of the motion size candidates differently depending on the motion vector accuracy of the current block.

15. A computer storage medium, characterized in that, The computer storage medium stores a video decoding / encoding implementation program. When the video decoding / encoding implementation program is applied to a decoding system, it implements the video decoding method as described in any one of claims 1-7; or, when the video decoding / encoding implementation program is applied to an encoding system, it implements the video encoding method as described in any one of claims 8-12.

16. A method for transmitting a bit stream, characterized in that, The video encoding method according to any one of claims 8-12 is used to generate the bitstream; and the bitstream is transmitted.

Citation Information

Patent Citations

  • Method and apparatus for coefficient scan based on partition mode of prediction unit

    CN107493474A

  • Image decoding apparatus and image encoding apparatus

    JP2013223049A