Image encoding / decoding method and device, and recording medium storing bitstream

By improving the prediction accuracy of sub-block-based temporal motion vector candidates through the derivation and correction of motion vectors, the method addresses the challenge of efficiently encoding and decoding high-resolution and high-quality images, thereby reducing transmission and storage costs.

WO2025110783A1PCT designated stage expired Publication Date: 2025-05-30HYUNDAI MOTOR CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/018589
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-22
Filing Date
2024-11-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images, such as UHD images, leads to a significant increase in image data, resulting in higher transmission and storage costs. Existing image encoding/decoding technologies struggle to efficiently handle these high-resolution and high-quality images.

Method used

The proposed method improves the prediction accuracy of sub-block-based temporal motion vector candidates by deriving an initial motion vector from an adjacent block and correcting it within a predetermined search range. This method determines a co-position reference block based on the final motion vector and derives sub-block temporal motion vector candidates from this reference block.

Benefits of technology

The improved prediction accuracy of motion vector candidates enhances the overall encoding/decoding efficiency, reducing the transmission and storage costs associated with high-resolution and high-quality images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024018589_30052025_PF_FP_ABST
    Figure KR2024018589_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides an image decoding method comprising the steps of: deriving an initial motion vector used for deriving sub-block temporal motion vector candidates of sub-blocks of a current block from a block adjacent to the current block; deriving a final motion vector by correcting the initial motion vector within a predetermined search range; determining a co-located reference block referred to by the current block within a co-located reference picture of the current block on the basis of the final motion vector; and deriving the sub-block temporal motion vector candidates of the sub-blocks of the current block from the co-located reference block.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding / decoding method, device, and recording medium storing bitstream

[0001] The present invention relates to a video encoding / decoding method, a device, and a recording medium storing a bitstream. Specifically, the present invention relates to a video encoding / decoding method, a device, and a recording medium storing a bitstream using improved sub-block-based temporal motion vector prediction.

[0002] Recently, the demand for high-resolution, high-quality images, such as UHD (Ultra High Definition) images, is increasing across various application fields. As image data becomes higher in resolution and quality, the relative amount of data increases compared to conventional image data. Therefore, transmitting image data using existing media such as wired or wireless broadband lines or storing it using existing storage media leads to increased transmission and storage costs. To address these issues arising from the increasing resolution and quality of image data, high-efficiency image encoding / decoding technologies for higher-resolution and higher-quality images are required.

[0003] By performing inter-screen prediction on a sub-block basis, blocks can be predicted more accurately. To improve the prediction accuracy of sub-block-based block encoding / decoding, the prediction accuracy of sub-block-based motion vector candidates needs to be improved. Various methods are being discussed to improve the prediction accuracy of sub-block-based temporal motion vector candidates among sub-block-based motion vector candidates.

[0004] The purpose of the present invention is to provide a video encoding / decoding method and device with improved encoding / decoding efficiency.

[0005] In addition, the present invention aims to provide a recording medium storing a bitstream generated by the image decoding method or device provided in the present invention.

[0006] A video decoding method according to one embodiment of the present invention may include a step of deriving an initial motion vector used for deriving sub-block temporal motion vector candidates of sub-blocks of a current block from an adjacent block of the current block, a step of deriving a final motion vector by correcting the initial motion vector within a predetermined search range, a step of determining a same-position reference block referenced by the current block within a same-position reference picture of the current block based on the final motion vector, and a step of deriving sub-block temporal motion vector candidates of the sub-blocks of the current block from the same-position reference block.

[0007] According to one embodiment, the adjacent block of the current block used for deriving the initial motion vector may be characterized as being one of a plurality of candidate adjacent blocks of the current block.

[0008] According to one embodiment, the plurality of candidate adjacent blocks may be characterized by including at least one of one or more adjacent blocks adjacent to the lower left of the current block, one or more adjacent blocks adjacent to the upper right of the current block, and one or more adjacent blocks adjacent to the upper left of the current block.

[0009] According to one embodiment, the adjacent block of the current block used for deriving the initial motion vector may be characterized in that it is determined as the candidate adjacent block that is available first according to a predetermined search order among a plurality of candidate adjacent blocks of the current block.

[0010] According to one embodiment, it may be characterized in that if the collocated reference picture of the current block and the reference picture used for prediction of the candidate adjacent block are the same, the candidate adjacent block is determined to be available.

[0011] According to one embodiment, even if the same-position reference picture of the current block and the reference picture used for prediction of the candidate adjacent block are different, it may be characterized in that the candidate adjacent block is determined to be available.

[0012] According to one embodiment, the adjacent block of the current block used for deriving the initial motion vector is characterized in that it is determined as a candidate adjacent block with the least distortion among a plurality of candidate adjacent blocks of the current block according to template matching, and the template matching may be characterized in that it is performed by determining distortion of the reference template based on a comparison between a current template adjacent to the current block and a reference template adjacent to the candidate adjacent block.

[0013] According to one embodiment, the template matching may be performed on a candidate adjacent block in which the same position reference picture of the current block and the reference picture used for prediction of the candidate adjacent block are the same, and the adjacent block of the current block is determined as the candidate adjacent block with the smallest distortion among the candidate adjacent blocks for which the template matching was performed.

[0014] According to one embodiment, the adjacent block of the current block used for deriving the initial motion vector may be characterized in that it is determined as a candidate adjacent block with the least distortion based on block matching among a plurality of candidate adjacent blocks of the current block.

[0015] According to one embodiment, the adjacent block of the current block may be determined by block matching reference block information indicating a candidate adjacent block used for deriving a same-position reference block among a plurality of candidate adjacent blocks of the current block.

[0016] According to one embodiment, the block matching reference block information may be characterized in that it is encoded with codewords of different lengths according to a predetermined order for the plurality of candidate adjacent blocks.

[0017] According to one embodiment, the predetermined order for the plurality of candidate adjacent blocks may be determined according to a distortion determined by template matching of the plurality of candidate adjacent blocks.

[0018] According to one embodiment, in the step of deriving the final motion vector, a motion vector pointing to a reference template with the smallest distortion according to template matching within the predetermined search range is determined as the final motion vector, and the template matching may be performed by determining distortion of the reference template based on a comparison between a current template adjacent to the current block and the reference template.

[0019] According to one embodiment, in the step of deriving the final motion vector, a motion vector pointing to a block matching reference block having the smallest distortion due to block matching within the predetermined search range is determined as the final motion vector, the block matching is performed by determining distortion of the block matching reference block based on a comparison between an adjacent block of the current block and the block matching reference block, and the block matching reference block may be a block located within a predetermined search range from a reference block referenced by an adjacent block of the current block.

[0020] According to one embodiment, in the step of deriving the final motion vector, the final motion vector may be derived by correcting the initial motion vector within a predetermined search range based on differential motion vector information indicating the difference between the initial motion vector and the final motion vector.

[0021] A video encoding method according to one embodiment of the present invention may include a step of deriving an initial motion vector used for deriving sub-block temporal motion vector candidates of sub-blocks of a current block from an adjacent block of the current block, a step of deriving a final motion vector by correcting the initial motion vector within a predetermined search range, a step of determining a same-position reference block referenced by the current block within a same-position reference picture of the current block based on the final motion vector, and a step of deriving sub-block temporal motion vector candidates of the sub-blocks of the current block from the same-position reference block.

[0022] A non-transitory computer-readable recording medium according to one embodiment of the present invention stores a bitstream generated by the image encoding method.

[0023] A transmission method according to one embodiment of the present invention transmits a bitstream generated by the image encoding method.

[0024] The features briefly summarized above of the present invention are merely exemplary aspects of the detailed description of the present invention described below and do not limit the scope of the present invention.

[0025] The present invention proposes various embodiments of a method for determining a sub-block based temporal motion vector prediction candidate.

[0026] Additionally, the present invention proposes various embodiments of a method for correcting a sub-block-based temporal motion vector prediction candidate.

[0027] According to the above various embodiments, as the prediction accuracy of the sub-block-based temporal motion vector prediction candidate is improved, the overall encoding efficiency can be improved.

[0028] Figure 1 is a block diagram showing the configuration according to one embodiment of an encoding device to which the present invention is applied.

[0029] Figure 2 is a block diagram showing the configuration according to one embodiment of a decryption device to which the present invention is applied.

[0030] FIG. 3 is a diagram schematically showing a video coding system to which the present invention can be applied.

[0031] Figure 4 illustrates a method for determining a same-position reference block corresponding to a current block using a motion vector of a left adjacent block of the current block.

[0032] Figure 5 illustrates a method for determining a same-position reference block corresponding to a current block using motion vectors of multiple adjacent blocks of the current block.

[0033] Figure 6 illustrates a method for determining a same-position reference block corresponding to a current block based on the most suitable motion vector among the motion vectors of multiple adjacent blocks of the current block based on template matching.

[0034] Figure 7 illustrates a method for determining a same-position reference block corresponding to a current block based on the most suitable motion vector among the motion vectors of multiple adjacent blocks of the current block according to block matching.

[0035] Figure 8 illustrates an embodiment of a method for correcting an initial motion vector using template matching.

[0036] Figure 9 illustrates an embodiment of a method for correcting an initial motion vector using block matching.

[0037] Fig. 10 illustrates another embodiment of a method for correcting an initial motion vector using block matching.

[0038] Fig. 11 illustrates an embodiment of a method for determining sub-block-based temporal motion vector prediction candidates.

[0039] Figure 12 illustrates an example of a content streaming system to which an embodiment according to the present invention can be applied.

[0040] A video decoding method according to one embodiment of the present invention may include a step of deriving an initial motion vector used for deriving sub-block temporal motion vector candidates of sub-blocks of a current block from an adjacent block of the current block, a step of deriving a final motion vector by correcting the initial motion vector within a predetermined search range, a step of determining a same-position reference block referenced by the current block within a same-position reference picture of the current block based on the final motion vector, and a step of deriving sub-block temporal motion vector candidates of the sub-blocks of the current block from the same-position reference block.

[0041] The present invention is susceptible to various modifications and embodiments. Therefore, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and substitutes falling within the spirit and scope of the present invention. In the drawings, similar reference numerals designate the same or similar functions throughout. The shape and size of elements in the drawings may be provided by way of example only for clarity. The detailed description of the exemplary embodiments described below refers to the accompanying drawings, which illustrate specific embodiments. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that the various embodiments, while different, are not necessarily mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the present invention. Furthermore, it should be understood that the location or arrangement of individual components within each disclosed embodiment may be modified without departing from the spirit and scope of the embodiment. Accordingly, the detailed description set forth below is not intended to be taken in a limiting sense, and the scope of the exemplary embodiments, if properly described, is defined only by the appended claims, along with the full scope equivalents to which such claims are entitled.

[0042] In the present invention, terms such as first, second, etc. may be used to describe various components, but the components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, the first component may be referred to as the second component, and similarly, the second component may also be referred to as the first component. The term "and / or" includes a combination of multiple related described items or any of multiple related described items.

[0043] The components shown in the embodiments of the present invention are independently depicted to represent different characteristic functions, and do not mean that each component is composed of separate hardware or a single software component. That is, each component is listed and included as a separate component for convenience of explanation, and at least two components among each component may be combined to form a single component, or a single component may be divided into multiple components to perform a function, and such integrated and separate embodiments of each component are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.

[0044] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In addition, some components of the present invention are not essential components that perform essential functions in the present invention and may be optional components merely for performance enhancement. The present invention may be implemented by including only components essential to realizing the essence of the present invention, excluding components used only for performance enhancement, and a structure including only essential components, excluding optional components used only for performance enhancement, is also within the scope of the present invention.

[0045] In embodiments, the term "at least one" may mean one of a number greater than or equal to 1, such as 1, 2, 3, and 4. In embodiments, the term "a plurality of" may mean one of a number greater than or equal to 2, such as 2, 3, and 4.

[0046] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In describing the embodiments of this specification, if it is determined that a detailed description of a related known configuration or function may obscure the gist of this specification, the detailed description will be omitted. The same reference numerals will be used for identical components in the drawings, and duplicate descriptions of identical components will be omitted.

[0047]

[0048] Glossary of Terms

[0049] Hereinafter, “video” may mean a single picture constituting a video, or may refer to the video itself. For example, “encoding and / or decoding of a video” may mean “encoding and / or decoding of a video,” or may mean “encoding and / or decoding of one of the videos constituting the video.”

[0050] Hereinafter, the terms "video" and "movie" may be used interchangeably and have the same meaning. Furthermore, the target image may be an encoding target image, which is the target of encoding, and / or a decoding target image, which is the target of decoding. Furthermore, the target image may be an input image input to an encoding device, or an input image input to a decoding device. Here, the target image may have the same meaning as the current image.

[0051] Hereinafter, “image”, “picture”, “frame” and “screen” may be used with the same meaning and may be used interchangeably.

[0052] Hereinafter, the term "target block" may refer to an encoding target block, which is the target of encoding, and / or a decoding target block, which is the target of decoding. Furthermore, the target block may refer to a current block, which is the target of current encoding and / or decoding. For example, the terms "target block" and "current block" may be used interchangeably and have the same meaning.

[0053] Hereinafter, "block" and "unit" may be used with the same meaning and may be used interchangeably. In addition, "unit" may mean including a luminance component block and a corresponding chroma component block to distinguish it from a block. For example, a coding tree unit (CTU) may be composed of one luma component (Y) coding tree block (CTB) and two chroma component (Cb, Cr) coding tree blocks associated with it.

[0054] Hereinafter, “sample,” “pixel,” and “pixel” may be used interchangeably and have the same meaning. Here, a sample may represent a basic unit that constitutes a block.

[0055] Hereinafter, “inter” and “between screens” may be used interchangeably and have the same meaning.

[0056] Hereinafter, “intra” and “within screen” may be used interchangeably and have the same meaning.

[0057]

[0058] Figure 1 is a block diagram showing the configuration according to one embodiment of an encoding device to which the present invention is applied.

[0059] The encoding device (100) may be an encoder, a video encoding device, or an image encoding device. A video may include one or more images. The encoding device (100) may sequentially encode one or more images.

[0060] Referring to FIG. 1, the encoding device (100) may include an image segmentation unit (110), an intra prediction unit (120), a motion prediction unit (121), a motion compensation unit (122), a switch (115), a subtractor (113), a transformation unit (130), a quantization unit (140), an entropy encoding unit (150), an inverse quantization unit (160), an inverse transformation unit (170), an adder (117), a filter unit (180), and a reference picture buffer (190).

[0061] Additionally, the encoding device (100) can generate a bitstream including encoded information through encoding an input image and output the generated bitstream. The generated bitstream can be stored in a computer-readable recording medium or can be streamed via a wired / wireless transmission medium.

[0062] The video segmentation unit (110) can segment the input video into various forms to increase the efficiency of video encoding / decoding. That is, the input video is composed of multiple pictures, and one picture can be hierarchically segmented and processed for compression efficiency, parallel processing, etc. For example, one picture can be segmented into one or more tiles or slices, which can then be segmented into multiple Coding Tree Units (CTUs). Alternatively, one picture can first be segmented into multiple sub-pictures defined as groups of rectangular slices, and each sub-picture can then be segmented into the tiles / slices. Here, the sub-pictures can be utilized to support the function of partially independently encoding / decoding and transmitting the picture. Since multiple sub-pictures can each be individually restored, there is an advantage of easy editing in applications that configure multi-channel input into a single picture. In addition, tiles can be segmented horizontally to generate bricks. Here, a brick can be utilized as the basic unit of parallel processing within a picture. In addition, one CTU can be recursively split into a quadtree (QT), and the terminal node of the split can be defined as a coding unit (CU). The CU can be split into a prediction unit (PU) and a transformation unit (TU), and prediction and splitting can be performed. Meanwhile, the CU can be utilized as a prediction unit and / or a transformation unit itself. Here, for flexible splitting, each CTU can be recursively split into a multi-type tree (MTT) as well as a quadtree (QT). Splitting of a CTU into a multi-type tree can start from the terminal node of a QT, and the MTT can be composed of a binary tree (BT) and a triple tree (TT).For example, the MTT structure can be divided into vertical binary split mode (SPLIT_BT_VER), horizontal binary split mode (SPLIT_BT_HOR), vertical ternary split mode (SPLIT_TT_VER), and horizontal ternary split mode (SPLIT_TT_HOR). In addition, the minimum block size (MinQTSize) of the quad tree of the luminance block during splitting can be set to 16x16, the maximum block size (MaxBtSize) of the binary tree can be set to 128x128, and the maximum block size (MaxTtSize) of the triple tree can be set to 64x64. In addition, the minimum block size (MinBtSize) of the binary tree and the minimum block size (MinTtSize) of the triple tree can be set to 4x4, and the maximum depth (MaxMttDepth) of the multi-type tree can be set to 4. Additionally, to improve the encoding efficiency of the I slice, a dual tree can be applied that uses different CTU partition structures for the luminance and chrominance components. On the other hand, in the P and B slices, the luminance and chrominance CTBs (Coding Tree Blocks) within the CTU can be partitioned into a single tree that shares the coding tree structure.

[0063] The encoding device (100) may perform encoding on the input image in intra mode and / or inter mode. Alternatively, the encoding device (100) may perform encoding on the input image in a third mode (e.g., IBC mode, Palette mode, etc.) other than the intra mode and inter mode. However, if the third mode has functional characteristics similar to the intra mode or inter mode, it may be classified as intra mode or inter mode for convenience of explanation. In the present invention, the third mode will be classified and described separately only when a specific description is required.

[0064] When the intra mode is used as the prediction mode, the switch (115) can be switched to intra, and when the inter mode is used as the prediction mode, the switch (115) can be switched to inter. Here, the intra mode can mean the intra prediction mode, and the inter mode can mean the inter-screen prediction mode. The encoding device (100) can generate a prediction block for an input block of an input image. In addition, after the prediction block is generated, the encoding device (100) can encode a residual block using the residual of the input block and the prediction block. The input image can be referred to as a current image that is currently a target of encoding. The input block can be referred to as a current block that is currently a target of encoding or an encoding target block.

[0065] When the prediction mode is intra mode, the intra prediction unit (120) can use samples of blocks already encoded / decoded around the current block as reference samples. The intra prediction unit (120) can perform spatial prediction on the current block using the reference samples, and can generate prediction samples for the input block through spatial prediction. Here, intra prediction can mean prediction within the screen.

[0066] As an intra prediction method, non-directional prediction modes such as DC mode and Planar mode, as well as directional prediction modes (e.g., 65 directions) can be applied. Here, the intra prediction method can be expressed as an intra prediction mode or an intra prediction mode.

[0067] When the prediction mode is inter mode, the motion prediction unit (121) can search for an area that best matches the input block from the reference image during the motion prediction process and derive a motion vector using the searched area. At this time, the area can be used as a search area. The reference image can be stored in the reference picture buffer (190). Here, when encoding / decoding for the reference image is processed, it can be stored in the reference picture buffer (190).

[0068] The motion compensation unit (122) can generate a prediction block for the current block by performing motion compensation using a motion vector. Here, inter prediction may mean inter-screen prediction or motion compensation.

[0069] The above motion prediction unit (121) and motion compensation unit (122) can generate a prediction block by applying an interpolation filter to a portion of the reference image when the value of the motion vector does not have an integer value. In order to perform inter-screen prediction or motion compensation, it is possible to determine whether the motion prediction and motion compensation method of the prediction unit included in the corresponding encoding unit is one of Skip Mode, Merge Mode, Advanced Motion Vector Prediction (AMVP) mode, and Intra Block Copy (IBC) mode based on the encoding unit, and inter-screen prediction or motion compensation can be performed according to each mode.

[0070] In addition, based on the above inter-screen prediction method, the AFFINE mode of sub-PU based prediction, the SbTMVP (Subblock-based Temporal Motion Vector Prediction) mode, and the MMVD (Merge with MVD) mode and the GPM (Geometric Partitioning Mode) mode of PU based prediction can be applied. In addition, in order to improve the performance of each mode, the HMVP (History based MVP), the PAMVP (Pairwise Average MVP), the CIIP (Combined Intra / Inter Prediction), the AMVR (Adaptive Motion Vector Resolution), the BDOF (Bi-Directional Optical-Flow), the BCW (Bi-predictive with CU Weights), the LIC (Local Illumination Compensation), the TM (Template Matching), and the OBMC (Overlapped Block Motion Compensation) can be applied.

[0071] The subtractor (113) can generate a residual block using the difference between the input block and the predicted block. The residual block may also be referred to as a residual signal. The residual signal may refer to the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming, quantizing, or transforming and quantizing the difference between the original signal and the predicted signal. The residual block may be a residual signal in block units.

[0072] The transform unit (130) can perform a transform on the residual block to generate a transform coefficient and output the generated transform coefficient. Here, the transform coefficient may be a coefficient value generated by performing a transform on the residual block. When the transform skip mode is applied, the transform unit (130) may also skip the transform on the residual block.

[0073] Quantized levels can be generated by applying quantization to transform coefficients or residual signals. In the following embodiments, quantized levels may also be referred to as transform coefficients.

[0074] For example, a 4x4 luminance residual block generated through within-screen prediction can be transformed using a basis vector based on DST (Discrete Sine Transform), and the remaining residual blocks can be transformed using a basis vector based on DCT (Discrete Cosine Transform). In addition, through RQT (Residual Quad Tree) technology, the transform block is divided into a quad tree shape for one block, and after performing transformation and quantization on each transform block divided through RQT, a coded block flag (cbf) can be transmitted to increase encoding efficiency when all coefficients become 0.

[0075] Another alternative is to apply Multiple Transform Selection (MTS) technology, which selectively performs transformation using multiple transformation bases. That is, instead of dividing CUs into TUs via RQT, a Sub-block Transform (SBT) technology can perform a function similar to TU division. Specifically, SBT is applied only to inter-screen prediction blocks, and unlike RQT, it can divide the current block into ½ or ¼ blocks vertically or horizontally, and then perform transformation on only one of the blocks. For example, in a vertically divided block, the transformation can be performed on the leftmost or rightmost block, and in a horizontally divided block, the transformation can be performed on the topmost or bottommost block.

[0076] Additionally, LFNST (Low Frequency Non-Separable Transform), a secondary transform technique that further transforms the residual signal converted to the frequency domain through DCT or DST, can be applied. LFNST additionally performs a transform on the low-frequency region of 4x4 or 8x8 in the upper left, which allows the residual coefficients to be concentrated in the upper left.

[0077] The quantization unit (140) can generate a quantized level by quantizing a transform coefficient or residual signal according to a quantization parameter (QP), and can output the generated quantized level. At this time, the quantization unit (140) can quantize the transform coefficient using a quantization matrix.

[0078] For example, a quantizer with QP values ​​of 0 to 51 can be used. Alternatively, if the image size is larger and high encoding efficiency is required, a QP of 0 to 63 can be used. In addition, a Dependent Quantization (DQ) method that uses two quantizers instead of a single quantizer can be applied. DQ performs quantization using two quantizers (e.g., Q0 and Q1), but even without signaling information about the use of a specific quantizer, the quantizer to be used for the next transform coefficient can be selected based on the current state through a state transition model.

[0079] The entropy encoding unit (150) can generate a bitstream by performing entropy encoding according to a probability distribution on values ​​produced by the quantization unit (140) or coding parameter values ​​produced during the encoding process, and can output the bitstream. The entropy encoding unit (150) can perform entropy encoding on information about image samples and information for decoding the image. For example, the information for decoding the image can include syntax elements, etc.

[0080] When entropy encoding is applied, a small number of bits are allocated to symbols with a high occurrence probability, and a large number of bits are allocated to symbols with a low occurrence probability, thereby representing the symbols, whereby the size of the bit string for the symbols to be encoded can be reduced. The entropy encoding unit (150) can use an encoding method such as exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), or Context-Adaptive Binary Arithmetic Coding (CABAC) for entropy encoding. For example, the entropy encoding unit (150) can perform entropy encoding using a Variable Length Coding / Code (VLC) table. In addition, the entropy encoding unit (150) may perform arithmetic encoding using the binarization method, probability model, and context model derived from the binarization method of the target symbol and the probability model of the target symbol / bin.

[0081] In this regard, when applying CABAC, the table probability update method can be changed to a simple formula-based table update method to reduce the size of the probability table stored in the decryption device. Furthermore, two different probability models can be used to obtain more accurate symbol probability values.

[0082] The entropy encoding unit (150) can change a two-dimensional block form coefficient into a one-dimensional vector form through a transform coefficient scanning method to encode a transform coefficient level (quantized level).

[0083] Coding parameters may include not only information (flags, indexes, etc.) encoded in an encoding device (100) and signaled to a decoding device (200), such as syntax elements, but also information derived during an encoding process or a decoding process, and may mean information necessary when encoding or decoding an image.

[0084] Here, signaling a flag or index may mean that the encoder entropy encodes the flag or index and includes it in the bitstream, and that the decoder entropy decodes the flag or index from the bitstream.

[0085] The encoded current image can be used as a reference image for other images to be processed later. Accordingly, the encoding device (100) can reconstruct or decode the encoded current image again and store the reconstructed or decoded image as a reference image in the reference picture buffer (190).

[0086] The quantized level can be dequantized in the dequantization unit (160) and inversely transformed in the inverse transformation unit (170). The dequantized and / or inversely transformed coefficients can be combined with a prediction block through an adder (117), and a reconstructed block can be generated by combining the dequantized and / or inversely transformed coefficients and the prediction block. Here, the dequantized and / or inversely transformed coefficients refer to coefficients on which at least one of dequantization and inverse transformation has been performed, and may refer to a reconstructed residual block. The dequantization unit (160) and the inverse transformation unit (170) can be performed in the reverse process of the quantization unit (140) and the transformation unit (130).

[0087] The restoration block may pass through a filter unit (180). The filter unit (180) may apply a deblocking filter, a sample adaptive offset (SAO), an adaptive loop filter (ALF), a bilateral filter (BIF), a Luma Mapping with Chroma Scaling (LMCS), etc. as a filtering technique, in whole or in part, to the restoration sample, restoration block, or restoration image. The filter unit (180) may also be referred to as an in-loop filter. In this case, the in-loop filter is also used as a name excluding LMCS.

[0088] A deblocking filter can remove block distortion that occurs at the boundaries between blocks. Whether to apply a deblocking filter to the current block can be determined based on the samples contained in several columns or rows within the block. When applying a deblocking filter to a block, different filters can be applied depending on the required deblocking filtering strength.

[0089] Sample adaptive offset can be used to compensate for encoding errors by adding an appropriate offset value to sample values. Sample adaptive offset can compensate for the offset from the original image on a sample-by-sample basis for deblocked images. This can be done by dividing the samples contained in the image into a fixed number of regions, determining the regions to be offset, and applying the offset to those regions. Alternatively, the offset can be applied by considering the edge information of each sample.

[0090] Bilateral filter (BIF) can also compensate for the offset from the original image on a sample-by-sample basis for the deblocked image.

[0091] An adaptive loop filter can perform filtering based on a comparison between a reconstructed image and the original image. By dividing the samples contained in the image into predetermined groups and determining the filter to be applied to each group, filtering can be performed differentially for each group. Information regarding whether to apply an adaptive loop filter can be signaled for each coding unit (CU), and the shape and filter coefficients of the adaptive loop filter applied to each block can vary.

[0092] In LMCS (Luma Mapping with Chroma Scaling), luma mapping (LM) refers to remapping luminance values ​​through a piece-wise linear model, and chroma scaling (CS) refers to a technique that scales the residual values ​​of chrominance components according to the average luminance value of the prediction signal. In particular, LMCS can be utilized as an HDR correction technique that reflects the characteristics of HDR (High Dynamic Range) images.

[0093] The restored block or restored image that has passed through the filter unit (180) may be stored in the reference picture buffer (190). The restored block that has passed through the filter unit (180) may be a part of the reference image. In other words, the reference image may be a restored image composed of restored blocks that have passed through the filter unit (180). The stored reference image may be used for inter-screen prediction or motion compensation thereafter.

[0094] Figure 2 is a block diagram showing the configuration according to one embodiment of a decryption device to which the present invention is applied.

[0095] The decoding device (200) may be a decoder, a video decoding device, or an image decoding device.

[0096] Referring to FIG. 2, the decoding device (200) may include an entropy decoding unit (210), an inverse quantization unit (220), an inverse transformation unit (230), an intra prediction unit (240), a motion compensation unit (250), an adder (201), a switch (203), a filter unit (260), and a reference picture buffer (270).

[0097] The decoding device (200) can receive a bitstream output from the encoding device (100). The decoding device (200) can receive a bitstream stored in a computer-readable recording medium, or a bitstream streamed through a wired / wireless transmission medium. The decoding device (200) can perform decoding on the bitstream in intra mode or inter mode. In addition, the decoding device (200) can generate a restored image or a decoded image through decoding, and can output the restored image or the decoded image.

[0098] If the prediction mode used for decryption is intra mode, the switch (203) can be switched to intra. If the prediction mode used for decryption is inter mode, the switch (203) can be switched to inter.

[0099] The decoding device (200) can decode the input bitstream to obtain a reconstructed residual block and generate a prediction block. Once the reconstructed residual block and the prediction block are obtained, the decoding device (200) can generate a reconstructed block to be decoded by adding the reconstructed residual block and the prediction block. The block to be decoded may be referred to as a current block.

[0100] The entropy decoding unit (210) can generate symbols by performing entropy decoding according to a probability distribution for the bitstream. The generated symbols may include symbols in the form of quantized levels. Here, the entropy decoding method may be the reverse process of the entropy encoding method described above.

[0101] The entropy decoding unit (210) can change a one-dimensional vector-shaped coefficient into a two-dimensional block-shaped coefficient through a transform coefficient scanning method to decode a transform coefficient level (quantized level).

[0102] The quantized level can be inversely quantized in the inverse quantization unit (220) and inversely transformed in the inverse transformation unit (230). The quantized level can be generated as a restored residual block as a result of performing inverse quantization and / or inverse transformation. At this time, the inverse quantization unit (220) can apply a quantization matrix to the quantized level. The inverse quantization unit (220) and inverse transformation unit (230) applied to the decoding device can apply the same technology as the inverse quantization unit (160) and inverse transformation unit (170) applied to the encoding device described above.

[0103] When intra mode is used, the intra prediction unit (240) can generate a predicted block by performing spatial prediction on the current block using sample values ​​of already decoded blocks surrounding the block to be decoded. The intra prediction unit (240) applied to the decoding device can apply the same technology as the intra prediction unit (120) applied to the encoding device described above.

[0104] When the inter mode is used, the motion compensation unit (250) can generate a prediction block by performing motion compensation using a motion vector and a reference image stored in the reference picture buffer (270) on the current block. The motion compensation unit (250) can generate a prediction block by applying an interpolation filter to a portion of the reference image when the value of the motion vector does not have an integer value. In order to perform motion compensation, it is possible to determine whether the motion compensation method of the prediction unit included in the corresponding encoding unit is skip mode, merge mode, AMVP mode, or current picture reference mode based on the encoding unit, and motion compensation can be performed according to each mode. The motion compensation unit (250) applied to the decoding device can apply the same technology as the motion compensation unit (122) applied to the encoding device described above.

[0105] The adder (201) can add the restored residual block and the predicted block to generate a restored block. The filter unit (260) can apply at least one of an Inverse-LMCS, a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the restored block or restored image. The filter unit (260) applied to the decoding device can apply the same filtering technology as that applied to the filter unit (180) applied to the encoding device described above.

[0106] The filter unit (260) can output a restored image. The restored block or restored image can be stored in the reference picture buffer (270) and used for inter prediction. The restored block that has passed through the filter unit (260) can be a part of the reference image. In other words, the reference image can be a restored image composed of restored blocks that have passed through the filter unit (260). The stored reference image can be used for inter-screen prediction or motion compensation thereafter.

[0107] FIG. 3 is a diagram schematically showing a video coding system to which the present invention can be applied.

[0108] A video coding system according to one embodiment may include an encoding device (10) and a decoding device (20). The encoding device (10) may transmit encoded video and / or image information or data to the decoding device (20) in the form of a file or streaming via a digital storage medium or a network.

[0109] An encoding device (10) according to one embodiment may include a video source generation unit (11), an encoding unit (12), and a transmission unit (13). A decoding device (20) according to one embodiment may include a reception unit (21), a decoding unit (22), and a rendering unit (23). The encoding unit (12) may be referred to as a video / image encoding unit, and the decoding unit (22) may be referred to as a video / image decoding unit. The transmission unit (13) may be included in the encoding unit (12). The reception unit (21) may be included in the decoding unit (22). The rendering unit (23) may include a display unit, and the display unit may be configured as a separate device or an external component.

[0110] The video source generation unit (11) can obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source generation unit (11) can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced with a process of generating related data.

[0111] The encoding unit (12) can encode the input video / image. The encoding unit (12) can perform a series of procedures such as prediction, transformation, and quantization for compression and encoding efficiency. The encoding unit (12) can output encoded data (encoded video / image information) in the form of a bitstream. The detailed configuration of the encoding unit (12) can also be configured in the same manner as the encoding device (100) of FIG. 1 described above.

[0112] The transmission unit (13) can transmit encoded video / image information or data output in the form of a bitstream to the reception unit (21) of the decoding device (20) via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (13) can include an element for generating a media file through a predetermined file format and can include an element for transmission via a broadcasting / communication network. The reception unit (21) can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit (22).

[0113] The decoding unit (22) can decode video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding unit (12). The detailed configuration of the decoding unit (22) can also be configured in the same manner as the decoding device (200) of FIG. 2 described above.

[0114] The rendering unit (23) can render the decrypted video / image. The rendered video / image can be displayed through the display unit.

[0115]

[0116] The present disclosure describes various embodiments of a method for determining a sub-block-based temporal motion vector prediction (SbTMVP) candidate. The sub-block-based temporal motion vector prediction can be applied to sub-block-based encoding to improve overall video encoding / decoding efficiency.

[0117] According to sub-block-based coding, a block is divided into sub-blocks. Motion vector candidates for inter-screen prediction of a block include information about the motion of the sub-blocks. For example, a sub-block-based spatial motion vector candidate includes information about a control point motion vector to determine the motion vector of the sub-block according to an affine model. Furthermore, a sub-block-based temporal motion vector candidate may include position information about the corresponding sub-block referenced by each sub-block of the block.

[0118] When determining a sub-block-based temporal motion vector candidate, the efficiency of sub-block-based encoding can be improved by more accurately determining the motion vector indicating the position of the corresponding sub-block. Therefore, in the present disclosure, a method is proposed for more accurately determining the motion vector indicating the corresponding sub-block from an adjacent block of the current block and correcting the motion vector determined from the adjacent block.

[0119]

[0120] Figure 4 illustrates a method for determining a same-position reference block corresponding to a current block using a motion vector of a left adjacent block of the current block.

[0121] According to FIG. 4, sub-block-based temporal motion vector candidates of the current block (402) can be derived from a collocated reference block (412) of a collocated reference picture (410) of the current block (402). In FIG. 4, the collocated reference picture (410) is a reference picture that is referenced to derive sub-block-based temporal motion vector candidates of the current block (402) among a plurality of reference picture candidates of the current block (402).

[0122] In order to determine the same-position reference block (412), the adjacent block of the A1 position (404) at the lower left of the current block (402) may be referenced. If the reference picture referred to by the adjacent block of the A1 position (404) is the same as the same-position reference picture (410) of the current block (402), sub-block-based temporal motion vector candidates of the current block (402) may be derived using the motion vector of the same-position reference block (412). According to one embodiment, when the above condition is satisfied, the same-position reference block (412) may be determined based on the motion vector used for the block of the A1 position (404) of the current picture (400). A same-position reference block (412) can be determined based on a position derived by adding the motion vector MV (406) used for the block at position A1 (404) to a predetermined position (center, upper left, upper right, lower left, lower right, etc.) of the current block (402). Consequently, the same-position reference block (412) can be determined to be adjacent to the A1' position (414) derived by adding the motion vector used for the block at position A1 (404) to the A1 position (404) of the current picture (400).

[0123] If the reference picture referred to by the adjacent block of the current block is not the same as the same-position reference picture (410) of the current block (402), the motion vector MV (406) may be regarded as a zero motion vector (zero MV). Accordingly, in this case, a block corresponding to the position of the current block (402) within the same-position reference picture (410) may be determined as a reference block for deriving sub-block-based temporal motion vector candidates of the current block (402). Accordingly, a reference block for deriving sub-block-based temporal motion vector candidates of the current block (402) within the same-position reference picture (410) may be determined based on a predetermined position (center, upper left, upper right, lower left, lower right, etc.) of the current block (402).

[0124] Information indicating a same-position reference picture (410) can be transmitted from a picture header (PH) or a slice header (SH). According to FIG. 4, since the corresponding same-position reference block is determined using only the A1 position (404) at the lower left of the current block (402), there may be a limit to the encoding efficiency. Therefore, by using adjacent blocks at other positions as well as the A1 position (404), a same-position reference block including sub-block-based temporal motion vector candidates with higher prediction accuracy can be derived.

[0125]

[0126] Figure 5 illustrates a method for determining a same-position reference block corresponding to a current block using motion vectors of multiple adjacent blocks of the current block.

[0127] According to FIG. 5, the motion information of adjacent positions A0, A1, B0, B1, and B2 (504, 506, 508, 510, 512) of the current block (502) can be used to determine the same-position reference block (522) corresponding to the current block (502). According to the search order of A0, A1, B0, B1, and B2 (504, 506, 508, 510, 512) that has been predetermined, it is determined whether the motion vectors of A0, A1, B0, B1, and B2 (504, 506, 508, 510, 512) are available for determining the same-position reference block (522). The above adjacent locations A0, A1, B0, B1, and B2 (504, 506, 508, 510, 512) can be referenced in determining spatial motion candidates in merge mode.

[0128] For example, when searching in the order of A0, A1, B0, B1, and B2 (504, 506, 508, 510, 512), it is determined whether the reference picture of the adjacent block of A0 (504) is the same as the same-position reference picture (520). If the two pictures are the same, the motion vector MV (514) is determined based on the motion vector of the adjacent block of A0 (504). Then, the same-position reference block (522) is determined based on the determined motion vector MV (514). For example, the position of the same-position reference block (522) can be determined by adding a predetermined position of the current block (502) and the motion vector MV (514).

[0129] If the reference picture of the adjacent block of A0 (504) is not identical to the same-position reference picture (520), it is determined whether the reference picture of the adjacent block of A1 (506) is identical to the same-position reference picture (520). If the two pictures are identical, the motion vector MV (514) is determined based on the motion vector of the adjacent block of A1 (506). If the reference picture of the adjacent block of A1 (506) is not identical to the same-position reference picture (520), it is determined whether to use the motion vectors of the adjacent blocks of B0, B1, and B2 (508, 510, 512) in the order of B0, B1, and B2 (508, 510, 512).

[0130] If the reference pictures of all adjacent blocks of A0, A1, B0, B1, and B2 (504, 506, 508, 510, 512) are not identical to the same-position reference picture (520), the motion vector MV (514) may be regarded as a zero motion vector. Therefore, in this case, a block corresponding to the position of the current block (502) within the same-position reference picture (520) may be determined as a reference block for deriving sub-block-based temporal motion vector candidates of the current block (502). Accordingly, a reference block for deriving sub-block-based temporal motion vector candidates of the current block (502) within the same-position reference picture (520) may be determined based on a predetermined position (center, upper left, upper right, lower left, lower right, etc.) of the current block (502).

[0131] In Fig. 5, the search order of A0, A1, B0, B1, and B2 (504, 506, 508, 510, 512) can be changed. In addition, only a part of A0, A1, B0, B1, and B2 (504, 506, 508, 510, 512) can be referenced, or other locations adjacent to the current block (502) can also be further included as search targets. That is, blocks at arbitrary locations around the current block and any number of blocks can be used for sub-block-based temporal motion vector prediction.

[0132]

[0133] FIG. 6 illustrates a method for determining a same-position reference block corresponding to a current block based on the most suitable motion vector among the motion vectors of multiple adjacent blocks of the current block based on template matching (TM).

[0134] According to FIG. 6, in order to determine a reference block corresponding to the current block in sub-block-based temporal motion vector prediction, surrounding neighboring blocks A0, A1, B0, B1, and B2 (604, 606, 608, 610, 612) of the current block are used. The adjacent positions A0, A1, B0, B1, and B2 (604, 606, 608, 610, 612) can be referenced in determining a spatial motion candidate in the merge mode. Depending on the embodiment, only a part of A0, A1, B0, B1, and B2 (604, 606, 608, 610, 612) may be referenced, or other predetermined positions may be additionally used. That is, blocks at any position surrounding the current block and any number of blocks can be used in sub-block-based temporal motion vector prediction.

[0135] In Fig. 5, the same-position reference block (522) is determined based on the first searched available motion vector in a predetermined order. However, in Fig. 6, the same-position reference block (622) is determined based on the most appropriate motion vector determined according to template matching among all available motion vectors of all adjacent positions. Accordingly, for each of A0, A1, B0, B1, and B2 (604, 606, 608, 610, 612), it is determined whether the reference picture of the corresponding adjacent block is identical to the same-position reference picture (620) of the current block (602).

[0136] And the motion vector of the adjacent block having the same reference picture as the same position reference picture (620) of the current block (602) is determined as the template matching candidate motion vector. And when there are two or more template matching candidate motion vectors, according to template matching, the block indicated by the template matching candidate motion vector having the smallest distortion between the current template (614) and the reference template is determined as the same position reference block of the current block (602). The current template (614) is determined as an area adjacent to the current block, and the reference template is determined as an area adjacent to the block indicated by the template matching candidate motion vector. FIG. 6 illustrates an example of calculating the distortion of the reference template (624) of the A0 template matching candidate block (622) indicated by the motion vector MV (616) of the adjacent block of the A0 position (604) of the current block (602) and the current template (614) of the current block (602). According to an embodiment, a motion vector of an adjacent block having a different reference picture than the same position reference picture (620) of the current block (602) may also be determined as a template matching candidate motion vector.

[0137] Here, the size of the template area is L1 x height for the left area and width x L2 for the upper area. L1 and L2 are arbitrary positive integers. In addition, the shape of the template can also be arbitrarily determined. For example, only the upper area of ​​the block can be included in the template area, or only the left area of ​​the block can be included in the template area. In addition to the upper area of ​​the block, the upper right area where the upper area of ​​the block is extended to the right can be included in the template area. And in addition to the left area of ​​the block, the lower left area where the left area of ​​the block is extended to the bottom can be included in the template area.

[0138] The distortion between the current template (614) and the reference template can be determined by various distortion measures such as the sum of absolute difference (SAD), the sum of square error (SSE), or the sum of absolute transformed difference (SATD).

[0139]

[0140] Figure 7 illustrates a method for determining a same-position reference block corresponding to a current block based on the most suitable motion vector among the motion vectors of multiple adjacent blocks of the current block based on block matching (BM).

[0141] According to FIG. 7, in order to determine a reference block corresponding to the current block in sub-block-based temporal motion vector prediction, surrounding neighboring blocks A0, A1, B0, B1, and B2 (704, 706, 708, 710, 712) of the current block are used. The adjacent positions A0, A1, B0, B1, and B2 (704, 706, 708, 710, 712) can be referenced in determining a spatial motion candidate in the merge mode. Depending on the embodiment, only a part of A0, A1, B0, B1, and B2 (704, 706, 708, 710, 712) may be referenced, or other predetermined positions may be additionally used. That is, blocks at any position surrounding the current block and any number of blocks can be used in sub-block-based temporal motion vector prediction.

[0142] In Fig. 6, the same-position reference block (622) is determined based on the most appropriate motion vector determined by template matching among all available motion vectors of adjacent positions. However, in Fig. 7, the same-position reference block (722) is determined based on the most appropriate motion vector determined by block matching among all available motion vectors of adjacent positions. Accordingly, for each of A0, A1, B0, B1, and B2 (704, 706, 708, 710, 712), it is determined whether the reference picture of the corresponding adjacent block is identical to the same-position reference picture (720) of the current block (702).

[0143] And the motion vector of the adjacent block having the same reference picture as the same-position reference picture (720) of the current block (702) is determined as the block matching candidate motion vector. And when there are two or more block matching candidate motion vectors, according to block matching, the block indicated by the block matching candidate motion vector having the smallest distortion between the current block (702) and the reference block is determined as the same-position reference block of the current block (702). In addition, according to an embodiment, the motion vector of the adjacent block having a different reference picture from the same-position reference picture (720) of the current block (702) may also be determined as the block matching candidate motion vector.

[0144] In the encoder, a same-position reference block can be determined according to the block matching method described above. Then, block matching reference block information indicating an adjacent block having a motion vector required for deriving the same-position reference block is encoded. Then, in the decoder, a same-position reference block is derived based on the adjacent block indicated by the block matching reference block information. The block matching reference block information can be encoded according to various encoding methods such as a fixed length code (FLC) or a truncated unary code (TU) depending on the number of available adjacent blocks of the current block (702). When codewords of different lengths are used, such as in truncated unary encoding, the search order of adjacent blocks is important. Therefore, the search order of adjacent blocks can be determined so that an adjacent block with high efficiency among the adjacent blocks of the current block (702) has a high priority. At this time, the search order of adjacent blocks may use the search order of spatial merge candidates in merge mode or an arbitrarily determined search order. According to another embodiment, the search order may be determined using the reference template of a reference block determined based on the current template of the current block and the motion information of the adjacent block. In the above embodiment, the reference block refers to a block referenced to determine a replacement block for the current block.

[0145]

[0146] Based on FIGS. 4 to 7, a method for determining a co-located reference block for deriving a sub-block-based temporal motion vector candidate is described. According to FIGS. 4 to 7, a co-located reference block is derived based on a motion vector of an adjacent block of a current block. Here, if the motion vector of an adjacent block of the current block is corrected, a reference block more suitable for the current block can be derived. Therefore, hereinafter, a method for correcting a motion vector for deriving a co-located reference block is described based on FIGS. 8 to 10. Hereinafter, an initial motion vector means a motion vector of an adjacent block of the current block selected for deriving a co-located reference block.

[0147]

[0148] Figure 8 illustrates an embodiment of a method for correcting an initial motion vector using template matching.

[0149] In Fig. 8, the motion vector MV (808) of the adjacent block at position A1 (804) of the current block (802) is used as the initial motion vector of the current block (802), and the initial motion vector can be corrected according to the distortion evaluation based on template matching. Even if the initial motion vector is derived from an adjacent block at a position other than position A1, the initial motion vector can be corrected.

[0150] An optimal reference template (822) having the smallest distortion compared to the current template (806) can be searched within a predefined search range (816) centered around a same-position reference block (812) indicated by a motion vector MV (808), which is an initial motion vector. Then, based on the optimal reference template (822), a corrected motion vector MV' (818) and a corrected same-position reference block (820) can be determined.

[0151] Alternatively, if the distortion of the optimal reference template (822) searched within the search range (816) as described above is smaller than the distortion of the reference template (814) of the same-position reference block (812), a corrected motion vector MV' (818) and a corrected same-position reference block (820) can be determined based on the optimal reference template (822).

[0152]

[0153] Figure 9 illustrates an embodiment of a method for correcting an initial motion vector using block matching.

[0154] In Fig. 9, the motion vector MV (906) of the A1 block (904) at the A1 position of the current block (902) is used as the initial motion vector of the current block (902), and the initial motion vector can be corrected based on a distortion evaluation according to block matching of the A1 block (904) and the corresponding block of the A1 block (904). Even if the initial motion vector is derived from an adjacent block at a position other than the A1 position, the initial motion vector can be corrected.

[0155] According to FIG. 9, an optimal block matching reference block (918) having the smallest distortion with respect to the A1 block (904) can be searched for within a predefined search range (916) centered around the A1' block (914) derived by the motion vector MV (906), which is an initial motion vector. Then, based on the optimal block matching reference block (918), a corrected motion vector MV' (920) and a corrected same-position reference block (922) can be determined.

[0156] Alternatively, if the distortion of the optimal block matching reference block (918) searched within the search range (916) as described above is less than the distortion of the A1' block (914), a corrected motion vector MV' (920) and a corrected same-position reference block (922) can be determined based on the optimal block matching reference block (918).

[0157]

[0158] Fig. 10 illustrates another embodiment of a method for correcting an initial motion vector using block matching.

[0159] In Fig. 10, the motion vector MV (1006) of the adjacent block at position A1 (1004) of the current block (1002) is used as the initial motion vector of the current block (1002), and the initial motion vector can be corrected according to the distortion evaluation according to block matching. Even if the initial motion vector is derived from an adjacent block at a position other than position A1, the initial motion vector can be corrected.

[0160] An optimal block matching reference block with the smallest distortion from the current block (1002) can be searched for within a predefined search range (1014) of the same-position reference block (1012) indicated by the motion vector MV (1006), which is an initial motion vector. Then, based on the optimal block matching reference block, a corrected motion vector MV' (1016) and a corrected same-position reference block (1018) can be determined.

[0161] Alternatively, if the distortion of the optimal block matching reference block searched within the search range (1014) as described above is less than the distortion of the same-position reference block (1012), a corrected motion vector MV' (1016) and a corrected same-position reference block (1018) can be determined based on the optimal block matching reference block.

[0162] In the encoder, the initial motion vector can be corrected according to the above method. Then, differential motion vector information representing the difference vector between the initial motion vector and the corrected motion vector can be encoded. Then, in the decoder, the final motion vector can be derived by correcting the initial motion vector within a predetermined search range according to the differential motion vector information.

[0163] In FIGS. 8 to 10, the size of the search range (816, 916, 1014) is MxN, and M and N are arbitrary positive integers. In addition, the search for the correction of the initial motion vector can be performed with not only integer-pel precision but also fractional-pel precision. When performing the search for integer-pel precision motion vector correction, horizontal and vertical searches can be performed within arbitrary S and T pixel units, where S and T are each positive integers greater than or equal to 1. When performing the search for sub-pel precision motion vector correction, a sub-pel precision unit W can first be determined. For example, if W is 1 / 2, each sub-pel position that divides two integers into two equal parts is searched, and if W is 1 / 4, each sub-pel position that divides two integers into four equal parts is searched. In addition, horizontal and vertical searches can be performed within arbitrary O and P sub-pel units. Here, O and P are each positive integers greater than or equal to 1.

[0164] Fig. 11 illustrates an embodiment of a method for deriving a sub-block temporal motion vector candidate of a current block.

[0165] In step 1102, an initial motion vector used to derive sub-block temporal motion vector candidates of sub-blocks of the current block is derived from adjacent blocks of the current block.

[0166] In one embodiment, the adjacent block of the current block used to derive the initial motion vector is one of a plurality of candidate adjacent blocks of the current block.

[0167] According to one embodiment, the plurality of candidate adjacent blocks may include at least one of one or more adjacent blocks adjacent to the lower left of the current block, one or more adjacent blocks adjacent to the upper right of the current block, and one or more adjacent blocks adjacent to the upper left of the current block.

[0168] According to one embodiment, the adjacent block of the current block used to derive the initial motion vector may be determined as the candidate adjacent block that is first available according to a predetermined search order among a plurality of candidate adjacent blocks of the current block. In addition, if the same-position reference picture of the current block and the reference picture used for predicting the candidate adjacent block are the same, the candidate adjacent block may be determined to be available.

[0169] According to one embodiment, the adjacent block of the current block used for deriving the initial motion vector may be determined as the candidate adjacent block with the least distortion among a plurality of candidate adjacent blocks of the current block based on template matching. The template matching may be performed by determining the distortion of the reference template based on a comparison between a current template adjacent to the current block and a reference template adjacent to the candidate adjacent block.

[0170] According to one embodiment, the template matching is performed on a candidate adjacent block in which the same position reference picture of the current block and the reference picture used for prediction of the candidate adjacent block are the same, and the adjacent block of the current block may be determined as the candidate adjacent block with the least distortion among the candidate adjacent blocks in which template matching was performed. In addition, according to an embodiment, even in a case in which the same position reference picture of the current block and the reference picture used for prediction of the candidate adjacent block are different, the candidate adjacent block may be determined to be available.

[0171] According to one embodiment, the adjacent block of the current block used for deriving the initial motion vector may be determined as a candidate adjacent block with the least distortion among a plurality of candidate adjacent blocks of the current block based on block matching. The block matching may be performed by determining the distortion of a block matching reference block based on a comparison of a replacement block of the current block and a block matching reference block used for prediction of the candidate adjacent block in the encoder.

[0172] According to one embodiment, in the encoder, the block matching is performed on a candidate adjacent block in which the same position reference picture of the current block and the reference picture used for prediction of the candidate adjacent block are the same, and the adjacent block of the current block may be determined as the candidate adjacent block with the smallest distortion among the candidate adjacent blocks for which block matching is performed. In addition, according to an embodiment, the block matching may be performed even in a case in which the same position reference picture of the current block and the reference picture used for prediction of the candidate adjacent block are different.

[0173] According to one embodiment, the adjacent block of the current block may be determined by block matching reference block information indicating a candidate adjacent block used for deriving a same-position reference block among a plurality of candidate adjacent blocks of the current block. The block matching reference block information may be encoded into codewords of different lengths according to a predetermined order for the plurality of candidate adjacent blocks. The predetermined order for the plurality of candidate adjacent blocks may be determined according to a distortion degree determined by template matching of the plurality of candidate adjacent blocks.

[0174] In step 1104, the final motion vector is derived by correcting the initial motion vector within a predetermined search range.

[0175] According to one embodiment, a motion vector pointing to a reference template with the smallest distortion due to template matching within the predetermined search range may be determined as the final motion vector. The template matching may be performed by determining the distortion of the reference template based on a comparison between the current template adjacent to the current block and the reference template.

[0176] According to one embodiment, a motion vector pointing to a block matching reference block with the smallest distortion due to block matching within a predetermined search range may be determined as the final motion vector. The block matching may be performed by determining the distortion of the block matching reference block based on a comparison between an adjacent block of the current block and the block matching reference block. In addition, the block matching reference block may be a block located within a predetermined search range from a reference block referenced by an adjacent block of the current block.

[0177] According to one embodiment, a final motion vector can be derived by correcting an initial motion vector within a predetermined search range based on differential motion vector information representing a difference between an initial motion vector and a final motion vector.

[0178] In step 1106, based on the final motion vector, a same-position reference block referenced by the current block within the same-position reference picture of the current block is determined.

[0179] In step 1108, sub-block temporal motion vector candidates for sub-blocks of the current block are derived from the same-position reference block. The corresponding sub-blocks corresponding to the sub-blocks of the current block are determined within the same-position reference block. Then, the motion vectors applied to the sub-blocks of the current block are derived from the corresponding sub-blocks of the same-position reference block. Consequently, the motion vectors of all sub-blocks of the current block are derived from the sub-blocks of the same-position reference block.

[0180] Depending on the prediction method performed in steps 1102 to 1108, the current block may be encoded or decoded. In some embodiments, the first to third syntax elements are encoded by the encoder according to the prediction result. The information is then decoded by the decoder and used for block prediction.

[0181] And the bitstream generated by the encoder according to the prediction method performed in steps 1102 to 1108 can be stored in a recording medium or transmitted outside the encoder.

[0182] FIG. 12 is a drawing exemplifying a content streaming system to which an embodiment according to the present invention can be applied.

[0183] As illustrated in FIG. 12, a content streaming system to which an embodiment of the present invention is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0184] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and CCTVs into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and CCTVs directly generate bitstreams, the encoding server may be omitted.

[0185] The above bitstream can be generated by an image encoding method and / or an image encoding device to which an embodiment of the present invention is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0186] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server can act as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server may control commands / responses between each device within the content streaming system.

[0187] The streaming server can receive content from a media repository and / or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0188] Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.

[0189] Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.

[0190] The above embodiments can be performed in the same or corresponding manner in an encoding device and a decoding device. In addition, an image can be encoded / decoded using at least one or a combination of at least one of the above embodiments.

[0191] The order in which the above embodiments are applied may be different in the encoding device and the decoding device. Alternatively, the order in which the above embodiments are applied may be the same in the encoding device and the decoding device.

[0192] The above embodiments can be performed for each of the luminance and chrominance signals. Alternatively, the above embodiments can be performed identically for the luminance and chrominance signals.

[0193] In the above embodiments, the methods are described based on a flowchart as a series of steps or units. However, the present invention is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will understand that the steps depicted in the flowchart are not exclusive, and other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the present invention.

[0194] The above embodiments may be implemented in the form of program commands that can be executed by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program commands, data files, data structures, etc., either singly or in combination. The program commands recorded on the computer-readable recording medium may be those specifically designed and constructed for the present invention, or may be known and usable by those skilled in the art of computer software.

[0195] The bitstream generated by the encoding method according to the above embodiment can be stored in a non-transitory computer-readable recording medium. In addition, the bitstream stored in the non-transitory computer-readable recording medium can be decoded by the decoding method according to the above embodiment.

[0196] Here, examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions such as ROMs, RAMs, and flash memories. Examples of program instructions include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices may be configured to operate as one or more software modules to perform processing according to the present invention, and vice versa.

[0197] Although the present invention has been described above with specific details such as specific components and limited examples and drawings, these are provided only to help a more general understanding of the present invention, and the present invention is not limited to the above examples, and those with ordinary knowledge in the technical field to which the present invention pertains can make various modifications and variations from this description.

[0198] Therefore, the idea of ​​the present invention should not be limited to the embodiments described above, and all things that are modified equally or equivalently to the following claims as well as the claims are considered to fall within the scope of the idea of ​​the present invention.

[0199] The present invention can be used in a device for encoding / decoding an image and a recording medium storing a bitstream.

Claims

1. In the video decryption method, A step of deriving an initial motion vector used for deriving sub-block temporal motion vector candidates of sub-blocks of the current block from adjacent blocks of the current block; A step of deriving a final motion vector by correcting the initial motion vector within a predetermined search range; A step of determining a same-position reference block referenced by the current block within the same-position reference picture of the current block based on the final motion vector; An image decoding method comprising the step of deriving sub-block temporal motion vector candidates of the sub-blocks of the current block from the same position reference block.

2. In paragraph 1, An image decoding method, characterized in that the adjacent block of the current block used to derive the initial motion vector is one of a plurality of candidate adjacent blocks of the current block.

3. In paragraph 2, The above multiple candidate adjacent blocks are, An image decoding method, characterized in that it includes at least one of one or more adjacent blocks adjacent to the lower left of the current block, one or more adjacent blocks adjacent to the upper right of the current block, and one or more adjacent blocks adjacent to the upper left of the current block.

4. In paragraph 2, An image decoding method, characterized in that the adjacent block of the current block used for deriving the initial motion vector is determined as the candidate adjacent block that is available first according to a predetermined search order among a plurality of candidate adjacent blocks of the current block.

5. In paragraph 4, An image decoding method, characterized in that if the collocated reference picture of the current block and the reference picture used for predicting the candidate adjacent block are the same, the candidate adjacent block is determined to be available.

6. In paragraph 4, An image decoding method, characterized in that even if the same-position reference picture of the current block and the reference picture used for predicting the candidate adjacent block are different, the candidate adjacent block is determined to be available.

7. In paragraph 2, The adjacent block of the current block used to derive the initial motion vector is characterized in that it is determined as the candidate adjacent block with the least distortion according to template matching among a plurality of candidate adjacent blocks of the current block. An image decoding method, characterized in that the template matching is performed by determining distortion of the reference template based on a comparison between a current template adjacent to the current block and a reference template adjacent to a candidate adjacent block.

8. In paragraph 7, The above template matching is, The prediction is performed on the same candidate adjacent block as the reference picture at the same position of the current block and the reference picture used for prediction of the candidate adjacent block, An image decoding method, characterized in that the adjacent block of the current block is determined as the candidate adjacent block with the least distortion among the candidate adjacent blocks for which template matching is performed.

9. In paragraph 2, An image decoding method, characterized in that the adjacent block of the current block used to derive the initial motion vector is determined as the candidate adjacent block with the least distortion based on block matching among a plurality of candidate adjacent blocks of the current block.

10. In paragraph 9, An image decoding method, characterized in that the adjacent block of the current block is determined by block matching reference block information indicating a candidate adjacent block used for deriving a same-position reference block among a plurality of candidate adjacent blocks of the current block.

11. In paragraph 10, The above block matching reference block information is, An image decoding method characterized in that the plurality of candidate adjacent blocks are encoded with codewords of different lengths according to a predetermined order.

12. In paragraph 11, The predetermined order for the above multiple candidate adjacent blocks is: An image decoding method characterized in that the distortion is determined according to the template matching of the plurality of candidate adjacent blocks.

13. In paragraph 1, In the step of deriving the final motion vector, A motion vector pointing to a reference template with the smallest distortion due to template matching within the above-mentioned search range is determined as the final motion vector, An image decoding method, characterized in that the template matching is performed by determining distortion of the reference template based on a comparison between the current template adjacent to the current block and the reference template.

14. In paragraph 1, In the step of deriving the final motion vector, A motion vector pointing to a block matching reference block with the smallest distortion due to block matching within the above-mentioned predetermined search range is determined as the final motion vector, The above block matching is performed by determining the distortion of the block matching reference block based on a comparison between the adjacent blocks of the current block and the block matching reference block, An image decoding method, characterized in that the above block matching reference block is a block located within a predetermined search range from a reference block referenced by an adjacent block of the current block.

15. In paragraph 1, In the step of deriving the final motion vector, An image decoding method, characterized in that a final motion vector is derived by correcting the initial motion vector within a predetermined search range based on differential motion vector information indicating the difference between the initial motion vector and the final motion vector.

16. In a video encoding method, A step of deriving an initial motion vector used for deriving sub-block temporal motion vector candidates of sub-blocks of the current block from adjacent blocks of the current block; A step of deriving a final motion vector by correcting the initial motion vector within a predetermined search range; A step of determining a same-position reference block referenced by the current block within the same-position reference picture of the current block based on the final motion vector; A video encoding method comprising the step of deriving sub-block temporal motion vector candidates of the sub-blocks of the current block from the same position reference block.

17. In a computer-readable recording medium storing a bitstream generated by a video encoding method, In a video encoding method, A step of deriving an initial motion vector used for deriving sub-block temporal motion vector candidates of sub-blocks of the current block from adjacent blocks of the current block; A step of deriving a final motion vector by correcting the initial motion vector within a predetermined search range; A step of determining a same-position reference block referenced by the current block within the same-position reference picture of the current block based on the final motion vector; A recording medium comprising a step of deriving sub-block temporal motion vector candidates of the sub-blocks of the current block from the same position reference block.

18. A bitstream transmission method for transmitting a bitstream generated by a video encoding method, A step of encoding an image based on the above image encoding method; and Comprising a step of transmitting a bitstream including the encoded image, The above image encoding method is, A step of deriving an initial motion vector used for deriving sub-block temporal motion vector candidates of sub-blocks of the current block from adjacent blocks of the current block; A step of deriving a final motion vector by correcting the initial motion vector within a predetermined search range; A step of determining a same-position reference block referenced by the current block within the same-position reference picture of the current block based on the final motion vector; A bitstream transmission method comprising the step of deriving sub-block temporal motion vector candidates of the sub-blocks of the current block from the same position reference block.

Citation Information

Patent Citations

  • O2o common-house domestic service providing server and method thereof

    KR1020200141684A

  • Electromagnetic induction heating apparatus for heating an aerosol-forming article of an electronic cigarette and driving method thereof

    KR1020240057162A

  • 2 free stop hinge systemomitted

    KR1020240145614A

  • Product demand forecasting method and device using decomposition technique and machine learning hybrid model

    KR102687917B1