Method and apparatus for encoding / decoding image, and recording medium storing bitstream

By deriving and correcting candidate temporal motion vectors from neighboring blocks of the current block, the encoding/decoding efficiency of high-resolution, high-quality images is improved, and the transmission and storage costs caused by the increase in data volume are resolved.

CN121970340APending Publication Date: 2026-05-01HYUNDAI MOTOR CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HYUNDAI MOTOR CO LTD
Filing Date
2024-11-22
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In the process of encoding/decoding high-resolution, high-quality images, the increased data volume in existing technologies leads to high transmission and storage costs, necessitating improvements in the accuracy of motion vector prediction based on sub-blocks to enhance encoding/decoding efficiency.

Method used

By deriving the initial motion vectors of the sub-block temporal motion vector candidates from the adjacent blocks of the current block, correcting the initial motion vectors within a predetermined search range, determining the final motion vectors, and deriving the sub-block temporal motion vector candidates from the co-located reference image.

Benefits of technology

This improves the accuracy of temporal motion vector prediction based on sub-blocks, thereby enhancing overall coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121970340A_ABST
    Figure CN121970340A_ABST
Patent Text Reader

Abstract

An image decoding method may include: deriving, from a neighboring block of a current block, an initial motion vector for deriving a sub-block temporal motion vector candidate of a sub-block of the current block; deriving a final motion vector by correcting the initial motion vector within a predetermined search range; based on the final motion vector, a co-located reference block referenced by the current block is determined in a co-located reference picture of the current block, and a sub-block temporal motion vector candidate of a sub-block of the current block is derived from the co-located reference block.
Need to check novelty before this filing date? Find Prior Art

Description

Methods and devices for encoding / decoding images and recording media for storing bit streams Technical Field

[0001] This invention relates to image encoding / decoding methods and apparatus, and recording media for storing bitstreams. More specifically, this invention relates to improved image encoding / decoding methods and apparatus using sub-block-based temporal motion vector prediction, and recording media for storing bitstreams. Background Technology

[0002] Recently, the demand for high-resolution, high-quality images, such as ultra-high-definition (UHD) images, has been increasing across various application areas. As image data becomes higher in resolution and quality, the relative volume of data increases compared to existing image data. Therefore, transmission and storage costs increase when using media such as existing wired and wireless broadband lines to transmit image data or when using existing storage media to store image data. To address these issues arising from the increasing resolution and quality of image data, efficient image encoding / decoding techniques for images with higher resolution and quality are needed.

[0003] By performing inter-frame prediction on a sub-block basis, more accurate block predictions can be achieved. To improve the prediction accuracy of sub-block-based block encoding / decoding, it is necessary to improve the prediction accuracy of sub-block-based motion vector candidates. Different methods are being discussed to improve the prediction accuracy of sub-block-based temporal motion vector candidates within the sub-block-based motion vector candidate pool. Summary of the Invention

[0004] Technical issues

[0005] The purpose of this invention is to provide a method and apparatus for encoding / decoding images with improved encoding / decoding efficiency.

[0006] Another object of the present invention is to provide a recording medium for storing a bitstream generated by the method or apparatus for decoding images provided by the present invention.

[0007] Technical solution

[0008] An image decoding method according to an embodiment of the present invention may include: deriving an initial motion vector from neighboring blocks of the current block for deriving candidate sub-block temporal motion vectors of sub-blocks of the current block; deriving a final motion vector by correcting the initial motion vector within a predetermined search range; determining a co-location reference block referenced by the current block in a co-location reference image of the current block based on the final motion vector, and deriving candidate sub-block temporal motion vectors of sub-blocks of the current block from the co-location reference block.

[0009] According to the implementation method, the neighboring block of the current block used to derive the initial motion vector can be one of a plurality of candidate neighboring blocks of the current block.

[0010] According to the implementation, the plurality of candidate neighboring blocks may include at least one of the following: one or more neighboring blocks adjacent to the lower left of the current block, one or more neighboring blocks adjacent to the upper right of the current block, and one or more neighboring blocks adjacent to the upper left of the current block.

[0011] According to the implementation method, the neighboring blocks of the current block used to derive the initial motion vector can be determined as the first available candidate neighboring block among a plurality of candidate neighboring blocks of the current block according to a predetermined search order.

[0012] According to the implementation method, if the co-location reference image of the current block is the same as the reference image used to predict the candidate neighboring block, then the candidate neighboring block can be determined to be available.

[0013] According to the implementation method, even if the co-located reference image of the current block is different from the reference image used to predict the candidate neighboring block, it can be determined that the candidate neighboring block is available.

[0014] According to the implementation method, the neighboring blocks of the current block used to derive the initial motion vector can be determined as the candidate neighboring blocks with the smallest distortion according to the template matching among a plurality of candidate neighboring blocks of the current block, and the template matching can be performed by determining the distortion of the reference template based on a comparison between the current template adjacent to the current block and the reference template adjacent to the candidate neighboring block.

[0015] According to the implementation method, template matching can be performed on candidate neighboring blocks such that the co-location reference image of the current block is the same as the reference image used to predict the candidate neighboring blocks, and the neighboring blocks of the current block can be determined as the candidate neighboring blocks with the least distortion among the candidate neighboring blocks for which template matching has been performed.

[0016] According to the implementation method, the neighboring blocks of the current block used to derive the initial motion vector can be determined as the candidate neighboring block with the smallest distortion based on block matching among a plurality of candidate neighboring blocks of the current block.

[0017] According to the implementation method, the neighboring blocks of the current block can be determined by block matching reference block information, which indicates the candidate neighboring blocks used to derive the co-position reference block among a plurality of candidate neighboring blocks of the current block.

[0018] According to the implementation method, the block matching reference block information can be encoded into codewords of different lengths according to a predetermined order for multiple candidate adjacent blocks.

[0019] According to the implementation method, the predetermined order of multiple candidate adjacent blocks can be determined based on the distortion determined by template matching of the multiple candidate adjacent blocks.

[0020] According to the implementation method, in the process of deriving the final motion vector, the motion vector of the reference template that indicates the minimum distortion according to the template matching within the predetermined search range can be determined as the final motion vector, and the template matching can be performed by determining the distortion of the reference template based on the comparison between the current template and the reference template adjacent to the current block.

[0021] According to the implementation method, in the process of deriving the final motion vector, the motion vector of the block matching reference block that indicates the minimum distortion according to block matching within a predetermined search range can be determined as the final motion vector. Block matching can be performed in the following way: the distortion of the block matching reference block is determined based on the comparison between the current block's neighboring blocks and the block matching reference block, and the block matching reference block can be a block whose distance from the reference block referenced by the current block's neighboring blocks is within the predetermined search range.

[0022] According to the implementation method, in the process of deriving the final motion vector, the final motion vector can be derived by correcting the initial motion vector within a predetermined search range based on differential motion vector information indicating the difference between the initial motion vector and the final motion vector.

[0023] The image encoding method according to an embodiment of the present invention may include: deriving an initial motion vector from neighboring blocks of the current block for deriving candidate sub-block temporal motion vectors for sub-blocks of the current block; deriving a final motion vector by correcting the initial motion vector within a predetermined search range; determining a co-location reference block referenced by the current block in a co-location reference image of the current block based on the final motion vector; and deriving candidate sub-block temporal motion vectors for sub-blocks of the current block from the co-location reference block.

[0024] According to embodiments of the present invention, a non-transient computer-readable recording medium can store a bitstream generated by an image encoding method.

[0025] The transmission method according to an embodiment of the present invention transmits a bit stream generated by an image encoding method.

[0026] The features briefly summarized above are merely exemplary aspects of the following specific embodiments of the present disclosure and do not limit the scope of the present disclosure.

[0027] Beneficial effects

[0028] This invention proposes various implementations of a method for determining candidates for time motion vector prediction based on sub-blocks.

[0029] Furthermore, this invention proposes various embodiments of a method for correcting candidates for time motion vector prediction based on sub-blocks.

[0030] According to various implementation methods, as the prediction accuracy of sub-block-based temporal motion vector prediction candidates is improved, the overall coding efficiency can be improved. Attached Figure Description

[0031] Figure 1 is a block diagram illustrating the configuration of an encoding device according to an embodiment of the present invention.

[0032] Figure 2 is a block diagram illustrating the configuration of a decoding device according to an embodiment of the present invention.

[0033] Figure 3 is a schematic diagram illustrating a video encoding system to which the present invention can be applied.

[0034] Figure 4 illustrates a method for determining the corresponding reference block using the motion vector of the left neighboring block of the current block.

[0035] Figure 5 illustrates a method for determining the corresponding reference block using the motion vectors of multiple neighboring blocks of the current block.

[0036] Figure 6 illustrates a method based on template matching to determine the corresponding reference block for the current block by selecting the most suitable motion vector from among the motion vectors of multiple neighboring blocks.

[0037] Figure 7 illustrates a method based on block matching, which determines the corresponding reference block to the current block based on the most suitable motion vector among the motion vectors of multiple neighboring blocks of the current block.

[0038] Figure 8 illustrates an implementation of the method for correcting the initial motion vector using template matching.

[0039] Figure 9 illustrates an implementation of the method for correcting the initial motion vector using block matching.

[0040] Figure 10 illustrates another implementation of the method for correcting the initial motion vector using block matching.

[0041] Figure 11 illustrates an implementation of a method for determining candidates for time motion vector prediction based on sub-blocks.

[0042] Figure 12 illustrates an exemplary content streaming system that can be applied according to an embodiment of the present invention.

[0043] Best mode

[0044] The image decoding method according to an embodiment of the present invention may include: deriving an initial motion vector from neighboring blocks of the current block for deriving candidate sub-block temporal motion vectors for sub-blocks of the current block; deriving a final motion vector by correcting the initial motion vector within a predetermined search range; determining a co-location reference block referenced by the current block within a co-location reference image of the current block based on the final motion vector; and deriving candidate sub-block temporal motion vectors for sub-blocks of the current block from the co-location reference block. Detailed Implementation

[0045] This disclosure can have various modifications and implementations, and specific implementations are illustrated in the accompanying drawings and described in detail in the specific implementations. However, this is not intended to limit this disclosure to the specific implementations, but should be understood to include all modifications, equivalents, or substitutions included within the spirit and scope of this disclosure. In various aspects, similar reference numerals in the drawings indicate the same or similar functions. The shapes and dimensions of the elements in the drawings are provided by way of example for clearer description. The detailed description of the exemplary implementations described below refers to the accompanying drawings, which illustrate specific implementations by way of example. These implementations are described in sufficient detail to enable those skilled in the art to practice the implementations. It should be understood that the various implementations differ from each other, but are not necessarily mutually exclusive. For example, the particular shapes, structures, and characteristics described herein may be implemented in other implementations without departing from the spirit and scope of this disclosure with respect to one implementation. It should also be understood that the position or arrangement of the various components in each disclosed implementation may be changed without departing from the spirit and scope of the implementations. Therefore, the specific implementations set forth below are not intended to be limiting, and the scope of the exemplary implementations is defined only by the full scope of the appended claims and their equivalents (if appropriately described).

[0046] In this disclosure, the terms first, second, etc., may be used to describe various components, but these components should not be limited by these terms. These terms are used only for the purpose of distinguishing one component from another. For example, without departing from the scope of this disclosure, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component. The term is and / or includes a combination of or any one of a plurality of related descriptive terms.

[0047] The components shown in the embodiments of this disclosure are depicted independently to indicate different functional characteristics, and this does not mean that each component is formed as a separate hardware or software configuration unit. That is, for ease of explanation, each component is listed and included as a separate component, and at least two components may be combined to form a single component, or a component may be divided into multiple components to perform functions, and embodiments that integrate components and embodiments that divide each component are also included within the scope of this disclosure, provided that they do not depart from the essence of this disclosure.

[0048] The terminology used in this disclosure is for descriptive purposes only and is not intended to limit the scope of the disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, some components of this disclosure are not essential for performing the necessary functions described herein and may be optional components used only to improve performance. This disclosure can be implemented by including only the essential components necessary to achieve the essence of this disclosure (excluding components used only to improve performance), and structures that include only the essential components (excluding optional components used only to improve performance) are also included within the scope of this disclosure.

[0049] In implementations, the term "at least one" may refer to one of a number greater than or equal to 1, such as 1, 2, 3, and 4. In implementations, the term "a plurality of" may refer to one of a number greater than or equal to 2, for example, 2, 3, and 4.

[0050] In the following, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In describing embodiments of this specification, detailed descriptions will be omitted if it is determined that a detailed description of a related known configuration or function might obscure the subject matter of this specification, and the same reference numerals will be used for the same components in the drawings, with repeated descriptions of the same components omitted.

[0051] Terminology Description

[0052] In the following text, “image” can refer to a picture that constitutes a video, or it can refer to the video itself. For example, “encoding and / or decoding of an image” can mean “encoding and / or decoding of a video,” and it can also mean “encoding and / or decoding of one of the images that constitutes a video.”

[0053] In the following text, "moving image" and "video" can be used with the same meaning and can be used interchangeably. Additionally, the target image can be an encoded target image that is the target of encoding and / or a decoded target image that is the target of decoding. Furthermore, the target image can be an input image to an encoding device and can also be an input image to a decoding device. Here, the target image can have the same meaning as the current image.

[0054] Hereinafter, "image", "picture", "frame", and "picture" may be used with the same meaning and may be used interchangeably.

[0055] Hereinafter, a "target block" may be an encoding target block that is an encoding target and / or a decoding target block that is a decoding target. Additionally, the target block may be the current block that is the target of current encoding and / or decoding. For example, "target block" and "current block" may be used with the same meaning and may be used interchangeably.

[0056] Hereinafter, "block" and "unit" may be used with the same meaning and may be used interchangeably. In addition, "unit" may represent including a luminance component block and its corresponding chrominance component block in order to distinguish it from a block. For example, a coding tree unit (CTU) may be composed of one luminance component (Y) coding tree block (CTB) and two chrominance component (Cb, Cr) coding tree blocks associated therewith.

[0057] Hereinafter, "sample", "image element", and "pixel" may be used with the same meaning and may be used interchangeably. In this document, a sample may represent the basic unit that constitutes a block.

[0058] Hereinafter, "inter" and "inter-picture" may be used with the same meaning and may be used interchangeably.

[0059] Hereinafter, "intra" and "intra-picture" may be used with the same meaning and may be used interchangeably.

[0060] FIG. 1 is a block diagram showing the configuration of an encoding device according to an embodiment of the present disclosure.

[0061] The encoding device 100 may be an encoder, a video encoding device, or an image encoding device. A video may include one or more images. The encoding device 100 may sequentially encode one or more images.

[0062] Referring to FIG. 1, the encoding device 100 may include an image segmentation unit 110, an intra prediction unit 120, a motion prediction unit 121, a motion compensation unit 122, a switch 115, a subtractor 113, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, a dequantization unit 160, an inverse transform unit 170, an adder 117, a filter unit 180, and a reference picture buffer 190.

[0063] In addition, the encoding device 100 may generate a bitstream including information encoded by encoding an input image, and output the generated bitstream. The generated bitstream may be stored in a computer-readable recording medium, or may be streamed through a wired / wireless transmission medium.

[0064] Image segmentation unit 110 can segment the input image into various forms to improve the efficiency of video encoding / decoding. That is, the input video consists of multiple images, and an image can be segmented and processed hierarchically for compression efficiency, parallel processing, etc. For example, an image can be segmented into one or more slices or pieces, and then further segmented into multiple coding tree units (CTUs). Alternatively, an image can first be segmented into multiple sub-images defined as groups of rectangular slices, and each sub-image can be segmented into slices / pieces. Here, sub-images can be used to support partially independent encoding / decoding and transmission of images. Since multiple sub-images can be reconstructed independently, it has the advantage of easy editing in applications that configure multi-channel input into a single image. Furthermore, slices can be horizontally divided to generate bricks. Here, bricks can be used as basic units for parallel processing within an image. Additionally, a CTU can be recursively segmented into a quadtree (QT), and the segmented terminal nodes can be defined as CUs (coding units). CUs can be segmented into PUs (prediction units) as prediction units and TUs (transform units) as transform units to perform prediction and segmentation. Simultaneously, the CU can be used as a prediction unit and / or a transform unit itself. Here, for flexible segmentation, each CTU can be recursively segmented into a multi-type tree (MTT) and a quadtree (QT). CTU segmentation into a multi-type tree can start from the terminal node of the QT, and the MTT can consist of a binary tree (BT) and a ternary tree (TT). For example, the MTT structure can be classified into a vertical binary tree segmentation pattern (SPLIT_BT_VER), a horizontal binary tree segmentation pattern (SPLIT_BT_HOR), a vertical ternary tree segmentation pattern (SPLIT_TT_VER), and a horizontal ternary tree segmentation pattern (SPLIT_TT_HOR). Furthermore, during segmentation, the minimum block size (MinQTSize) of the quadtree for the luma block can be set to 16×16, the maximum block size (MaxBtSize) of the binary tree can be set to 128×128, and the maximum block size (MaxTtSize) of the ternary tree can be set to 64×64. Furthermore, the minimum block size (MinBtSize) of the binary tree and the minimum block size (MinTtSize) of the ternary tree can be specified as 4×4, and the maximum depth (MaxMttDepth) of the multi-type tree can be specified as 4. Additionally, to improve the coding efficiency of I slices, a dual tree structure using different CTU partitioning structures for luma and chroma components can be applied. On the other hand, in P and B slices, the luma and chroma CTBs (coding tree blocks) within the CTU can be partitioned into single trees sharing a coding tree structure.

[0065] Encoding device 100 can encode the input image in intra-frame mode and / or inter-frame mode. Alternatively, encoding device 100 can encode the input image in a third mode other than intra-frame mode and inter-frame mode (e.g., IBC mode, palette mode, etc.). However, if the third mode has similar functional characteristics to the intra-frame mode or inter-frame mode, it can be classified as an intra-frame mode or inter-frame mode for ease of interpretation. In this disclosure, the third mode is classified and described separately only when a specific description is required.

[0066] When the intra-frame mode is used as the prediction mode, switch 115 can be switched to intra-frame, and when the inter-frame mode is used as the prediction mode, switch 115 can be switched to inter-frame. Here, intra-frame mode can refer to intra-frame prediction mode, and inter-frame mode can refer to inter-frame prediction mode. Encoding device 100 can generate prediction blocks for input blocks of the input image. Furthermore, encoding device 100 can encode residual blocks using the residuals of the input blocks and prediction blocks after generating the prediction blocks. The input image can be referred to as the current image of the current coding target. The input block can be referred to as the current block of the current coding target or the coding target block.

[0067] When the prediction mode is intra-frame mode, the intra-frame prediction unit 120 can use samples from blocks that have already been encoded / decoded around the current block as reference samples. The intra-frame prediction unit 120 can perform spatial prediction on the current block using the reference samples, or generate prediction samples for the input block through spatial prediction. Here, intra-frame prediction can refer to intra-frame prediction.

[0068] As an intra-frame prediction method, non-directional prediction modes and directional prediction modes (e.g., 65 directions) such as DC mode and Planar mode can be applied. Here, the intra-frame prediction method can be represented as an intra-frame prediction mode or an intra-frame prediction mode.

[0069] When the prediction mode is inter-frame mode, the motion prediction unit 121 can retrieve the region that best matches the input block from the reference image during motion prediction processing and derive the motion vector using the retrieved region. In this case, the search region can be used as the region. The reference image can be stored in the reference image buffer 190. Here, it can be stored in the reference image buffer 190 when encoding / decoding for the reference image is performed.

[0070] The motion compensation unit 122 can generate a prediction block for the current block by performing motion compensation using motion vectors. Here, inter-frame prediction can mean inter-picture prediction or motion compensation.

[0071] When the value of the motion vector is not an integer, the motion prediction unit 121 and the motion compensation unit 122 can generate prediction blocks by applying an interpolation filter to a portion of the reference image. To perform inter-frame prediction or motion compensation, it can be determined whether the motion prediction and motion compensation mode of the prediction unit contained in the coding unit is one of skip mode, merge mode, advanced motion vector prediction (AMVP) mode, or coding unit-based intra-block copy (IBC) mode, and inter-frame prediction or motion compensation can be performed according to each mode.

[0072] Furthermore, based on the above inter-frame prediction methods, the following modes can be applied: AFFINE mode based on sub-PU prediction, SbTMVP (sub-block temporal motion vector prediction), MMVD (merged with MVD) mode based on PU prediction, and GPM (geometric segmentation mode). In addition, to improve the performance of each mode, HMVP (history-based MVP), PAMVP (pairwise average MVP), CIIP (combined intra / inter-frame prediction), AMVR (adaptive motion vector resolution), BDOF (bidirectional optical flow), BCW (bidirectional prediction with CU weights), LIC (local illumination compensation), TM (template matching), and OBMC (overlapping block motion compensation).

[0073] Subtractor 113 can generate a residual block by using the difference between the input block and the prediction block. The residual block can be referred to as a residual signal. The residual signal can represent the difference between the original signal and the prediction signal. Alternatively, the residual signal can be a signal generated by transforming or quantizing, or transforming and quantizing the difference between the original signal and the prediction signal. The residual block can be the residual signal of a block cell.

[0074] Transform unit 130 can generate transform coefficients by performing a transform on the residual block and output the generated transform coefficients. Here, the transform coefficients can be coefficient values ​​generated by performing a transform on the residual block. When a transform skip mode is applied, transform unit 130 can skip the transform on the residual block.

[0075] Quantization levels can be generated by applying quantization to the transform coefficients or the residual signal. In the following implementation, quantization levels may also be referred to as transform coefficients.

[0076] For example, a 4×4 lumen residual block generated by intra-frame prediction is transformed using DST (Discrete Sine Transform)-based basis vectors, and the remaining residual blocks can be transformed using DCT (Discrete Cosine Transform)-based basis vectors. Furthermore, the Residual Quad Tree (RQT) technique is used to partition the transformed block into a quadtree shape for each block, and after performing transformation and quantization on each transformed block partitioned by RQT, the coded block flag (cbf) can be transmitted when all coefficients become 0 to increase coding efficiency.

[0077] As an alternative, a Multiple Transform Selection (MTS) technique can be applied, which selectively uses multiple transform bases to perform the transform. That is, instead of segmenting the CU into TUs via RQT, a sub-block transform (SBT) technique can be used to perform a function similar to TU segmentation. Specifically, SBT is applied only to inter-frame prediction blocks, and unlike RQT, the current block can be segmented into 1 / 2 or 1 / 4 of its size vertically or horizontally, and then the transform can be performed on only one block. For example, if the current block is segmented vertically, the transform can be performed on the leftmost or rightmost block, and if the current block is segmented horizontally, the transform can be performed on the topmost or bottommost block.

[0078] Alternatively, LFNST (Low Frequency Non-Separable Transform) can be applied. LFNST is a secondary transform technique that additionally transforms the residual signal transformed to the frequency domain by DCT or DST. LFNST additionally performs a transform on the upper left 4×4 or 8×8 low-frequency region, allowing the residual coefficients to be concentrated in the upper left.

[0079] The quantization unit 140 can generate a quantization level by quantizing the transform coefficients or residual signal according to the quantization parameters (QP), and output the generated quantization level. Here, the quantization unit 140 can quantize the transform coefficients using a quantization matrix.

[0080] For example, quantizers with QP values ​​from 0 to 51 can be applied. Alternatively, if the image size is large and high coding efficiency is required, QP values ​​from 0 to 63 can be used. Furthermore, the DQ (Dependency Quantization) method, which uses two quantizers instead of one, can be applied. DQ performs quantization using two quantizers (e.g., Q0 and Q1), but even without signaling information about the use of a particular quantizer, a state transition model can be used to select the quantizer to be used for the next transform coefficient based on the current state.

[0081] Entropy coding unit 150 can generate and output a bitstream by performing entropy coding on values ​​calculated by quantization unit 140 or on coding parameter values ​​calculated during encoding according to a probability distribution. Entropy coding unit 150 can perform entropy coding on information about samples of the image and information used for decoding the image. For example, information used for decoding the image may include syntax elements.

[0082] When entropy coding is applied, symbols are represented such that fewer bits are allocated to symbols with high occurrence probabilities and more bits are allocated to symbols with low occurrence probabilities, and therefore, the size of the bitstream used to encode the symbols can be reduced. The entropy coding unit 150 can perform entropy coding using coding methods such as Exponential Golomb, Context Adaptive Variable-Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), etc. For example, the entropy coding unit 150 can perform entropy coding by using a variable-length code / code (VLC) table. Furthermore, the entropy coding unit 150 can derive a binaryization method for the target symbol and a probability model for the target symbol / binary bit (bin), and perform arithmetic coding by using the derived binaryization method and context model.

[0083] Relatedly, when applying CABAC, to reduce the size of the probability table stored in the decoding device, the table probability update method can be changed to a table update method using simple equations and applied. Furthermore, two different probability models can be used to obtain more accurate symbol probability values.

[0084] In order to encode the transform coefficient level (quantization level), the entropy coding unit 150 can change the two-dimensional block form coefficients into a one-dimensional vector form by the transform coefficient scanning method.

[0085] The encoding parameters may include information such as syntax elements encoded in the encoding device 100 and signaled to the decoding device 200 (tags, indexes, etc.) and information derived during the encoding or decoding process, and may represent the information required when encoding or decoding an image.

[0086] Here, passing a tag or index signal can represent entropy encoding of the corresponding tag or index in the encoder and including it in the bitstream, and can also represent entropy decoding of the corresponding tag or index from the bitstream in the decoder.

[0087] The encoded current image can be used as a reference image for another image that will be processed later. Therefore, the encoding device 100 can reconstruct or decode the encoded current image again and store the reconstructed or decoded image as a reference image in the reference image buffer 190.

[0088] The quantization level can be dequantized in dequantization unit 160 or inverse transformed in inverse transform unit 170. The coefficients of the dequantization and / or inverse transform can be added to the prediction block via adder 117. Here, the coefficients of the dequantization and / or inverse transform can refer to the coefficients to which at least one of dequantization and inverse transform is performed, and can refer to the reconstructed residual block. Dequantization unit 160 and inverse transform unit 170 can be performed as the inverse process of quantization unit 140 and transform unit 130.

[0089] The reconstructed blocks can be processed by filter unit 180. Filter unit 180 can apply deblocking filters, sample adaptive offset (SAO), adaptive loop filters (ALF), bilateral filters (BIF), luminance mapping and chrominance scaling (LMCS), etc., to apply all or some of these filtering techniques to reconstruct samples, reconstructed blocks, or reconstructed images. Filter unit 180 can be referred to as an in-loop filter. In this case, in-loop filter is also used as a name to exclude LMCS.

[0090] Deblocking filters remove block distortion that occurs at the boundaries between blocks. To determine whether to apply a deblocking filter, the selection can be based on samples included in several rows or columns contained within the block. When applying a deblocking filter to a block, different filters can be applied depending on the desired deblocking intensity.

[0091] To compensate for coding errors using sample-adaptive offsets, appropriate offset values ​​can be added to the sample values. Sample-adaptive offsets can correct the offset between the deblocked image and the original image on a sample-by-sample basis. Methods include segmenting the samples contained in the image into a predetermined number of regions, determining the regions to which the offset will be applied, and then applying the offset to the determined regions; or considering edge information about each sample when applying the offset.

[0092] Bilateral filters (BIF) can also correct the offset from the original image on a sample-by-sample basis for images that have undergone deblocking.

[0093] An adaptive loop filter (ALF) can perform filtering based on a comparison between the reconstructed image and the original image. Samples contained in the image can be segmented into predetermined groups, the filter to be applied to each group can be determined, and differential filtering can be performed for each group. Information regarding whether to apply the ALF can be transmitted via coding unit (CU) signals, and the form and coefficients of the adaptive loop filter to be applied to each block can vary.

[0094] In LMCS (Luminosity Mapping and Chroma Scaling), luminosity mapping (LM) refers to remapping luminosity values ​​using a piecewise linear model, and chromosity scaling (CS) refers to a technique used to scale the residual values ​​of the chromosity components based on the average luminosity values ​​of the predicted signal. Specifically, LMCS can be used as an HDR correction technique that reflects the characteristics of HDR (High Dynamic Range) images.

[0095] The reconstructed blocks or reconstructed image that have passed through filter unit 180 can be stored in reference image buffer 190. The reconstructed blocks that have passed through filter unit 180 can be a part of the reference image. That is, the reference image is a reconstructed image composed of the reconstructed blocks that have passed through filter unit 180. The stored reference image can be used subsequently in inter-frame prediction or motion compensation.

[0096] Figure 2 is a block diagram illustrating the configuration of a decoding device according to an embodiment of the present disclosure.

[0097] Decoding device 200 can be a decoder, video decoding device, or image decoding device.

[0098] Referring to Figure 2, the decoding device 200 may include an entropy decoding unit 210, a dequantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, a motion compensation unit 250, an adder 201, a switch 203, a filter unit 260, and a reference image buffer 270.

[0099] Decoding device 200 can receive bitstreams output from encoding device 100. Decoding device 200 can receive bitstreams stored on a computer-readable recording medium, or bitstreams streamed via wired / wireless transmission media. Decoding device 200 can decode bitstreams in intra-frame mode or inter-frame mode. Furthermore, decoding device 200 can generate and output reconstructed or decoded images generated through decoding.

[0100] When the prediction mode used for decoding is intra-frame mode, switch 203 can be switched to intra-frame mode. Alternatively, when the prediction mode used for decoding is inter-frame mode, switch 203 can be switched to inter-frame mode.

[0101] Decoding device 200 can obtain reconstructed residual blocks and generate prediction blocks by decoding the input bitstream. When the reconstructed residual blocks and prediction blocks are obtained, decoding device 200 can generate a reconstructed block that becomes the decoding target by adding the reconstructed residual blocks and prediction blocks together. The decoding target block can be referred to as the current block.

[0102] The entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream according to a probability distribution. The generated symbols may include symbols in the form of quantization levels. Here, the entropy decoding method can be the inverse process of the entropy encoding method described above.

[0103] The entropy decoding unit 210 can transform the coefficients of a one-dimensional vector shape into coefficients of a two-dimensional block shape using a transform coefficient scanning method, in order to decode the transform coefficient level (quantization level).

[0104] The quantization level can be dequantized in dequantization unit 220 or inverse transformed in inverse transform unit 230. The quantization level can be the result of dequantization and / or inverse transform, and can be generated as a reconstructed residual block. Here, dequantization unit 220 can apply a quantization matrix to the quantization level. The dequantization unit 220 and inverse transform unit 230 applied to the decoding device can employ the same techniques as those applied to the dequantization unit 160 and inverse transform unit 170 of the aforementioned encoding device.

[0105] When using intra-frame mode, intra-frame prediction unit 240 can generate a prediction block by performing spatial prediction on the current block using sample values ​​from blocks already decoded around the target block. Intra-frame prediction unit 240 applied to the decoding device can employ the same techniques as intra-frame prediction unit 120 applied to the encoding device described above.

[0106] When using inter-frame mode, motion compensation unit 250 can generate a prediction block by performing motion compensation on the current block using motion vectors and a reference image stored in reference image buffer 270. When the value of the motion vector is not an integer, motion compensation unit 250 can generate a prediction block by applying an interpolation filter to a portion of the reference image. To perform motion compensation, the motion compensation method of the prediction unit contained in the corresponding coding unit can be determined based on the coding unit: skip mode, merge mode, AMVP mode, or current image reference mode, and motion compensation can be performed according to each mode. The motion compensation unit 250 applied to the decoding device can apply the same techniques as the motion compensation unit 122 applied to the coding device described above.

[0107] Adder 201 generates a reconstructed block by adding the reconstructed residual block and the predicted block. Filter unit 260 can apply at least one of inverse LMCS, deblocking filter, sample adaptive offset, and adaptive loop filter to the reconstructed block or the reconstructed image. Filter unit 260 applied to the decoding device can apply the same filtering technique as that applied to filter unit 180 applied to the aforementioned encoding device.

[0108] Filter unit 260 can output a reconstructed image. The reconstructed blocks or reconstructed image can be stored in reference image buffer 270 and used for inter-frame prediction. The reconstructed blocks that have passed through filter unit 260 can be part of the reference image. That is, the reference image can be a reconstructed image composed of reconstructed blocks that have passed through filter unit 260. The stored reference image can be used subsequently in inter-frame prediction or motion compensation.

[0109] Figure 3 is a schematic diagram illustrating a video coding system to which this disclosure can be applied.

[0110] The video encoding system according to the embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may transmit encoded video and / or image information or data to the decoding device 20 in the form of a file or stream via a digital storage medium or network.

[0111] The encoding device 10 according to the embodiments may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. The decoding device 20 according to the embodiments may include a receiving unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The receiving unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, and the display unit may be configured as a separate device or an external component.

[0112] The video source generation unit 11 can obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, tablet computer, and smartphone, etc., and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc., in which case the video / image capture process can be replaced by a process of generating relevant data.

[0113] Encoding unit 12 can encode the input video / image. Encoding unit 12 can perform a series of processes for compression and encoding efficiency, such as prediction, transformation, and quantization. Encoding unit 12 can output encoded data (encoded video / image information) in the form of a bitstream. The detailed configuration of encoding unit 12 can also be configured in the same way as the encoding device 100 of Figure 1 above.

[0114] The transmission unit 13 can transmit encoded video / image information or data, output in bitstream form, to the receiving unit 21 of the decoding device 20 via a digital storage medium or network, either as a file or a stream. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit 13 can include elements for generating media files according to a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and send it to the decoding unit 22.

[0115] The decoding unit 22 can decode the video / image by performing a series of processes (such as dequantization, inverse transform, and prediction) corresponding to the operations of the encoding unit 12. The detailed configuration of the decoding unit 22 can also be configured in the same way as the decoding device 200 of FIG2 above.

[0116] Rendering unit 23 can render decoded video / images. The rendered video / images can be displayed through the display unit.

[0117] This disclosure describes different implementations of a method for determining sub-block-based temporal motion vector prediction (SbTMVP) candidates. Sub-block-based temporal motion vector prediction can be applied to sub-block-based coding to improve overall video coding / decoding efficiency.

[0118] The block is divided into sub-blocks based on sub-block coding. Furthermore, motion vector candidates for inter-frame prediction of the block include information about the motion of the sub-blocks. For example, sub-block-based spatial motion vector candidates include information about control point motion vectors to determine the motion vectors of the sub-blocks according to an affine model. Additionally, sub-block-based temporal motion vector candidates may include positional information of the corresponding sub-blocks referenced by each sub-block of the block.

[0119] When determining candidates for temporal motion vectors based on sub-blocks, the efficiency of sub-block-based coding can be improved by more accurately determining the motion vector indicating the position of the corresponding sub-block. Therefore, this disclosure proposes a method for more accurately determining the motion vector indicating the corresponding sub-block from the neighboring blocks of the current block and correcting the motion vector determined from the neighboring blocks.

[0120] Figure 4 illustrates a method for determining the corresponding reference block using the motion vector of the left neighboring block of the current block.

[0121] According to Figure 4, the sub-block-based temporal motion vector candidates of the current block 402 can be derived from the co-position reference block 412 of the co-position reference image 410 of the current block 402. In Figure 4, the co-position reference image 410 is a reference image that is referenced to derive the sub-block-based temporal motion vector candidates of the current block 402 from multiple reference image candidates of the current block 402.

[0122] To determine the corresponding reference block 412, the adjacent block at position A1 404 on the lower left side of the current block 402 can be referenced. If the reference image referenced by the adjacent block at position A1 404 is the same as the corresponding reference image 410 of the current block 402, then the motion vector of the corresponding reference block 412 can be used to derive the candidate time motion vector of the current block 402 based on sub-blocks. According to one embodiment, when this condition is met, the corresponding reference block 412 can be determined based on the motion vector of the block at position A1 404 for the current image 400. The corresponding reference block 412 can be determined based on the position derived by adding the motion vector MV 406 of the block at position A1 404 to a predetermined position (center, upper left, upper right, lower left, lower right, etc.) of the current block 402. As a result, the corresponding reference block 412 can be determined to be adjacent to position A1' 414 derived by adding the motion vector of the block at position A1 404 to position A1 404 of the current image 400.

[0123] If the reference image referenced by the adjacent blocks of the current block is different from the co-located reference image 410 of the current block 402, then the motion vector MV 406 can be considered a zero motion vector (zero MV). Therefore, in this case, the block corresponding to the position of the current block 402 within the co-located reference image 410 can be identified as a reference block for deriving the sub-block-based temporal motion vector candidate of the current block 402. Thus, a reference block for deriving the sub-block-based temporal motion vector candidate of the current block 402 within the co-located reference image 410 can be determined based on a predetermined position of the current block 402 (center, top left, top right, bottom left, bottom right, etc.).

[0124] Information indicating the corresponding reference image 410 can be sent from the image header (PH) or slice header (SH). According to Figure 4, since only the A1 position 404 on the lower left side of the current block 402 is used to determine the corresponding corresponding reference block, there may be limitations in coding efficiency. Therefore, by using adjacent blocks at other positions and the A1 position 404, a corresponding reference block including sub-block-based temporal motion vector candidates with higher prediction accuracy can be derived.

[0125] Figure 5 illustrates a method for determining the corresponding reference block using the motion vectors of multiple neighboring blocks of the current block.

[0126] According to Figure 5, the motion information at adjacent positions A0 504, A1 506, B0 508, B1 510, and B2 512 of the current block 502 can be used to determine the corresponding reference block 522. Based on the pre-determined search order of A0 504, A1 506, B0 508, B1 510, and B2 512, it is determined whether the motion vectors of A0 504, A1 506, B0 508, B1 510, and B2 512 can be used to determine the corresponding reference block 522. When determining spatial motion candidates in the merging mode, adjacent positions A0 504, A1 506, B0 508, B1 510, and B2 512 can be referenced.

[0127] For example, when searching in the order of A0 504, A1 506, B0 508, B1 510, and B2 512, it is determined whether the reference image of the adjacent block A0 504 is the same as the corresponding reference image 520. If the two images are the same, a motion vector MV 514 is determined based on the motion vector of the adjacent block A0 504. Then, the corresponding reference block 522 is determined based on the determined motion vector MV 514. For example, the position of the corresponding reference block 522 can be determined by adding the predetermined position of the current block 502 to the motion vector MV 514.

[0128] If the reference image of adjacent block A0 504 is different from the corresponding reference image 520, then it is determined whether the reference image of adjacent block A1 506 is the same as the corresponding reference image 520. If the two images are the same, then motion vector MV 514 is determined based on the motion vector of adjacent block A1 506. If the reference image of adjacent block A1 506 is different from the corresponding reference image 520, then it is determined whether the motion vectors of adjacent blocks B0 508, B1 510, and B2 512 should be used in the order of B0 508, B1 510, and B2 512.

[0129] If the reference images of all adjacent blocks A0 504, A1 506, B0 508, B1 510, and B2 512 are different from the co-position reference image 520, then motion vector MV 514 can be considered a zero motion vector. Therefore, in this case, the block corresponding to the position of the current block 502 within the co-position reference image 520 can be identified as a reference block for deriving the sub-block-based temporal motion vector candidate for the current block 502. Thus, reference blocks for deriving the sub-block-based temporal motion vector candidate for the current block 502 within the co-position reference image 520 can be determined based on the predetermined position (center, top left, top right, bottom left, bottom right, etc.) of the current block 502.

[0130] In Figure 5, the search order of A0 504, A1 506, B0 508, B1 510, and B2 512 can be changed. Furthermore, only some of A0 504, A1 506, B0 508, B1 510, and B2 512 can be referenced, or other locations adjacent to the current block 502 can be included as search targets. That is, blocks at any location around the current block and any number of blocks can be used for sub-block-based temporal motion vector prediction.

[0131] Figure 6 illustrates a method based on template matching (TM) to determine the corresponding reference block for the current block by selecting the most suitable motion vector from among the motion vectors of multiple neighboring blocks.

[0132] According to Figure 6, in sub-block-based temporal motion vector prediction, to determine the reference block corresponding to the current block, the surrounding neighboring blocks A0 604, A1 606, B0 608, B1 610, and B2 612 of the current block are used. When determining spatial motion candidates in the merging mode, neighboring positions A0 604, A1 606, B0 608, B1 610, and B2 612 can be referenced. Depending on the implementation, only some of A0 604, A1 606, B0 608, B1 610, and B2 612 can be referenced, or other predetermined positions can be used additionally. That is, blocks at any location around the current block and any number of blocks can be used in sub-block-based temporal motion vector prediction.

[0133] In Figure 5, the corresponding reference block 522 is determined based on the first available motion vector found in a predetermined order. However, in Figure 6, the corresponding reference block 622 is determined based on the most suitable motion vector determined from all available motion vectors at all adjacent positions according to template matching. Therefore, for each of A0 604, A1 606, B0 608, B1 610, and B2 612, it is determined whether the reference image of the corresponding adjacent block is the same as the corresponding reference image 620 of the current block 602.

[0134] Furthermore, motion vectors of adjacent blocks having the same reference image as the co-located reference image 620 of the current block 602 are determined as template matching candidate motion vectors. Additionally, when two or more template matching candidate motion vectors exist, the block indicated by the template matching candidate motion vector with the minimum distortion between the current template 614 and the reference template is determined as the co-located reference block of the current block 602, according to template matching. The current template 614 is determined as the region adjacent to the current block, and the reference template is determined as the region adjacent to the block indicated by the template matching candidate motion vector. Figure 6 shows an example of the distortion calculated between the reference template 624 of the A0 template matching candidate block 622 indicated by the motion vector MV 616 of the adjacent block at position A0 604 of the current block 602 and the current template 614 of the current block 602. According to an embodiment, motion vectors of adjacent blocks having a reference image different from the co-located reference image 620 of the current block 602 can also be determined as template matching candidate motion vectors.

[0135] Here, the size of the left side of the template area is L1 × height, and the size of the top area is width × L2. L1 and L2 are arbitrary positive integers. Furthermore, the shape of the template can be arbitrarily determined. For example, the template area may include only the top area of ​​the block, or only the left side of the block. In addition to the top area of ​​the block, the upper right region extending to the right from the top area can be included in the template area. Furthermore, in addition to the left area of ​​the block, the lower left region extending from the left side of the block to the bottom can be included in the template area.

[0136] The distortion between the current template 614 and the reference template can be determined by different distortion measurement methods, such as sum of absolute differences (SAD), sum of squared errors (SSE), or sum of absolute transformation differences (SATD).

[0137] Figure 7 illustrates a method based on block matching (BM) to determine the corresponding reference block for the current block by selecting the most suitable motion vector from among the motion vectors of multiple neighboring blocks.

[0138] According to Figure 7, in sub-block-based temporal motion vector prediction, to determine the reference block corresponding to the current block, the surrounding neighboring blocks A0 704, A1 706, B0 708, B1 710, and B2 712 of the current block are used. When determining spatial motion candidates in merging mode, neighboring positions A0 704, A1 706, B0 708, B1 710, and B2 712 can be referenced. Depending on the implementation, only some of A0 704, A1 706, B0 708, B1 710, and B2 712 can be referenced, or other predetermined positions can be used additionally. That is, blocks at any location around the current block and any number of blocks can be used in sub-block-based temporal motion vector prediction.

[0139] In Figure 6, a corresponding reference block 622 is determined based on the most suitable motion vector identified through template matching among all available motion vectors at adjacent positions. However, in Figure 7, a corresponding reference block 722 is determined based on the most suitable motion vector identified through block matching among all available motion vectors at adjacent positions. Therefore, for each of A0 704, A1 706, B0 708, B1 710, and B2 712, it is determined whether the reference image of the corresponding adjacent block is the same as the corresponding reference image 720 of the current block 702.

[0140] Furthermore, motion vectors of adjacent blocks having the same reference image as the co-located reference image 720 of the current block 702 are determined as block matching candidate motion vectors. Additionally, when two or more block matching candidate motion vectors exist, the block indicated by the block matching candidate motion vector with the least distortion between the current block 702 and the reference block is determined as the co-located reference block of the current block 702, according to block matching. Furthermore, according to an embodiment, motion vectors of adjacent blocks having a reference image different from the co-located reference image 720 of the current block 702 can also be determined as block matching candidate motion vectors.

[0141] In the encoder, a corresponding reference block can be determined according to the block matching method described above. Furthermore, block matching reference block information is encoded, indicating neighboring blocks with the motion vectors required to derive the corresponding reference block. In the decoder, the corresponding reference block is derived based on the neighboring blocks indicated by the block matching reference block information. The block matching reference block information can be encoded according to various encoding methods (such as fixed-length coding (FLC) or truncated unary coding (TU)) depending on the number of available neighboring blocks in the current block 702. When using codewords of different lengths (such as truncated unary coding), the search order of neighboring blocks is important. Therefore, the search order of neighboring blocks can be determined such that neighboring blocks with high efficiency among the neighboring blocks of the current block 702 have high priority. In this case, the search order of neighboring blocks can use the search order of spatial merging candidates in merging mode or an arbitrarily determined search order. According to another embodiment, the search order can be determined using a reference template of a reference block determined based on the current template of the current block and the motion information of neighboring blocks. In this embodiment, a reference block refers to a block that is referenced to determine a replacement block for the current block.

[0142] Based on Figures 4 to 7, a method for determining a co-location reference block for deriving candidate time motion vectors based on sub-blocks has been described. According to Figures 4 to 7, the co-location reference block is derived based on the motion vectors of neighboring blocks of the current block. Here, when the motion vectors of neighboring blocks of the current block are corrected, a reference block more suitable for the current block can be derived. Therefore, in the following, a method for correcting the motion vectors used to derive the co-location reference block will be described based on Figures 8 to 10. In the following, the initial motion vector refers to the motion vector of the neighboring blocks of the current block selected for deriving the co-location reference block.

[0143] Figure 8 illustrates an implementation of the method for correcting the initial motion vector using template matching.

[0144] In Figure 8, the motion vector MV 808 of the adjacent block at position A1 804 of the current block 802 is used as the initial motion vector of the current block 802, and the initial motion vector can be corrected based on the distortion evaluation according to template matching. The initial motion vector can be corrected even if it is derived from the adjacent block at a position other than A1.

[0145] An optimal reference template 822 with minimal distortion relative to the current template 806 can be searched within a predefined search range 816 centered on a co-position reference block 812 indicated by motion vector MV 808 (which is the initial motion vector). Furthermore, based on the optimal reference template 822, the corrected motion vector MV' 818 and the corrected co-position reference block 820 can be determined.

[0146] Alternatively, if the distortion of the optimal reference template 822 searched within the search range 816 as described above is less than the distortion of the reference template 814 of the co-position reference block 812, the corrected motion vector MV'818 and the corrected co-position reference block 820 can be determined based on the optimal reference template 822.

[0147] Figure 9 illustrates an implementation of the method for correcting the initial motion vector using block matching.

[0148] In Figure 9, the motion vector MV 906 of block A1 904 at position A1 of current block 902 is used as the initial motion vector of current block 902, and the initial motion vector can be corrected based on the distortion evaluation of block matching according to block A1 904 and its corresponding block. The initial motion vector can be corrected even if it is derived from an adjacent block at a position other than A1.

[0149] According to Figure 9, an optimal block-matching reference block 918 with minimal distortion between itself and block A1 904 can be searched within a predefined search range 916 centered on block A1' 914 derived from motion vector MV 906 (which is the initial motion vector). Furthermore, based on the optimal block-matching reference block 918, the corrected motion vector MV' 920 and the corrected co-position reference block 922 can be determined.

[0150] Alternatively, if the distortion of the optimal block matching reference block 918 searched within the search range 916 as described above is less than the distortion of block 914 A1', the corrected motion vector MV'920 and the corrected co-position reference block 922 can be determined based on the optimal block matching reference block 918.

[0151] Figure 10 illustrates another implementation of the method for correcting the initial motion vector using block matching.

[0152] In Figure 10, the motion vector MV 1006 of the adjacent block at position A1 1004 of the current block 1002 is used as the initial motion vector of the current block 1002, and the initial motion vector can be corrected based on the distortion evaluation according to block matching. The initial motion vector can be corrected even if it is derived from the adjacent block at a position other than A1.

[0153] The optimal block-matching reference block with the minimum distortion between the current block 1002 and the reference block 1012 can be searched within a predefined search range 1014 indicated by the motion vector MV 1006 (which is the initial motion vector). Furthermore, based on the optimal block-matching reference block, the corrected motion vector MV' 1016 and the corrected reference block 1018 can be determined.

[0154] Alternatively, if the distortion of the optimal block-matching reference block searched within the search range 1014 described above is less than the distortion of the co-position reference block 1012, the corrected motion vector MV'1016 and the corrected co-position reference block 1018 can be determined based on the optimal block-matching reference block.

[0155] In the encoder, the initial motion vector can be corrected according to the method described above. Furthermore, differential motion vector information, representing the difference between the initial and corrected motion vectors, can be encoded. Then, in the decoder, the final motion vector can be derived by correcting the initial motion vector within a predetermined search range based on the differential motion vector information.

[0156] In Figures 8 to 10, the search range 816, 916, or 1014 is M × N, where M and N are arbitrary positive integers. Furthermore, the search for initial motion vector correction can be performed not only with integer pixel precision but also with fractional pixel precision. When performing a search for motion vector correction with integer pixel precision, horizontal and vertical searches can be performed within arbitrary S and T pixel units. Here, S and T are positive integers greater than or equal to 1. When performing a search for motion vector correction with fractional pixel precision, the fractional pixel precision unit W can be determined first. For example, when W is 1 / 2, the search divides two integer pixels into two equal parts at each fractional pixel position; when W is 1 / 4, the search divides two integer pixels into four equal parts at each fractional pixel position. Furthermore, horizontal and vertical searches can be performed within arbitrary O and P fractional pixel units. Here, O and P are positive integers greater than or equal to 1.

[0157] Figure 11 illustrates an implementation of a method for deriving sub-block-based temporal motion vector prediction candidates for the current block.

[0158] In step 1102, initial motion vectors for deriving time motion vector candidates for the sub-blocks of the current block are derived from the adjacent blocks of the current block.

[0159] According to one implementation, the neighboring blocks of the current block used to derive the initial motion vector are one of a plurality of candidate neighboring blocks of the current block.

[0160] According to one implementation, the plurality of candidate neighboring blocks may include at least one of the following: one or more neighboring blocks adjacent to the lower left of the current block, one or more neighboring blocks adjacent to the upper right of the current block, and one or more neighboring blocks adjacent to the upper left of the current block.

[0161] According to one implementation, the neighboring blocks of the current block used to derive the initial motion vector can be determined as the first available candidate neighboring block among a plurality of candidate neighboring blocks of the current block according to a predetermined search order. Furthermore, if the co-located reference image of the current block is the same as the reference image used to predict the candidate neighboring block, then the candidate neighboring block can be determined to be available.

[0162] According to one implementation, the neighboring blocks of the current block used to derive the initial motion vector can be determined as the candidate neighboring blocks with the smallest distortion according to template matching among a plurality of candidate neighboring blocks of the current block. Template matching can be performed by determining the distortion of the reference template based on a comparison between the current template adjacent to the current block and the reference template adjacent to the candidate neighboring block.

[0163] According to one implementation, template matching is performed on candidate neighbor blocks where the co-location reference image of the current block is the same as the reference image used to predict the candidate neighbor blocks, and the neighbor blocks of the current block can be determined as the candidate neighbor blocks with the least distortion among the candidate neighbor blocks for which template matching has been performed. Furthermore, according to one implementation, even if the co-location reference image of the current block is different from the reference image used to predict the candidate neighbor blocks, it can still be determined that the candidate neighbor block is usable.

[0164] According to one implementation, the neighboring blocks of the current block used to derive the initial motion vector can be determined as the candidate neighboring blocks with the smallest distortion based on block matching among a plurality of candidate neighboring blocks of the current block. Block matching can be performed by determining the distortion of the block matching reference block in the encoder based on a comparison between the replacement block of the current block and the block matching reference block used to predict the candidate neighboring blocks.

[0165] According to one embodiment, in the encoder, block matching is performed on candidate neighboring blocks in which the co-location reference image of the current block is the same as the reference image used to predict candidate neighboring blocks, and the neighboring blocks of the current block can be determined as the candidate neighboring blocks with the least distortion among the candidate neighboring blocks for which block matching is performed. Furthermore, according to the embodiment, block matching can be performed even when the co-location reference image of the current block is different from the reference image used to predict candidate neighboring blocks.

[0166] According to one implementation, neighboring blocks of the current block can be determined using block matching reference block information, which indicates candidate neighboring blocks among a plurality of candidate neighboring blocks of the current block for deriving a co-located reference block. The block matching reference block information can be encoded into codewords of different lengths according to a predetermined order for the plurality of candidate neighboring blocks. The predetermined order of the plurality of candidate neighboring blocks can be determined based on the distortion determined by template matching of the plurality of candidate neighboring blocks.

[0167] In step 1104, the final motion vector is derived by correcting the initial motion vector within a predetermined search range.

[0168] According to one implementation, the motion vector of the reference template that indicates the minimum distortion according to the template matching within a predetermined search range can be determined as the final motion vector. Template matching can be performed by determining the distortion of the reference template based on a comparison between the current template and the reference template adjacent to the current block.

[0169] According to one implementation, the motion vector of the block-matching reference block that indicates the minimum distortion according to block matching within a predetermined search range can be determined as the final motion vector. Block matching can be performed by determining the distortion of the block-matching reference block based on a comparison between the current block's neighboring blocks and the block-matching reference block. Furthermore, the block-matching reference block can be a block whose distance from the reference block referenced by the current block's neighboring blocks is within the predetermined search range.

[0170] According to one implementation, the final motion vector can be derived by correcting the initial motion vector within a predetermined search range based on differential motion vector information indicating the difference between the initial motion vector and the final motion vector.

[0171] In step 1106, based on the final motion vector, the co-location reference block of the current block is determined within the co-location reference image of the current block.

[0172] In step 1108, candidate sub-block time motion vectors for the current block's sub-blocks are derived from the co-location reference block. A corresponding sub-block within the co-location reference block is determined that corresponds to a sub-block of the current block. Furthermore, the motion vectors applied to the sub-blocks of the current block are derived from the corresponding sub-blocks of the co-location reference block. As a result, the motion vectors for all sub-blocks of the current block are derived from the sub-blocks of the co-location reference block.

[0173] Based on the prediction method performed in steps 1102 to 1108, the current block can be encoded or decoded. In some implementations, first to third syntax elements are encoded within the encoder based on the prediction results. The information can be decoded in the decoder and used for block prediction.

[0174] Furthermore, the bitstream generated by the encoder according to the prediction method executed in steps 1102 to 1108 can be stored in the recording medium or sent to the outside of the encoder.

[0175] Figure 12 exemplarily illustrates a content streaming system that can be applied according to an embodiment of the present disclosure.

[0176] As shown in Figure 12, a content streaming media system using the embodiments of this disclosure may mainly include an encoding server, a streaming media server, a network server, a media storage device, a user device, and a multimedia input device.

[0177] The encoding server compresses content received from multimedia input devices (such as smartphones, cameras, CCTV, etc.) into digital data to generate a bitstream, and then transmits it to the streaming media server. As another example, if the multimedia input device (e.g., smartphone, camera, CCTV, etc.) generates the bitstream directly, then the encoding server can be omitted.

[0178] The bitstream can be generated by applying the image encoding method and / or image encoding device of the present disclosure, and the streaming media server can temporarily store the bitstream during the sending or receiving of the bitstream.

[0179] A streaming media server sends multimedia data to a user's device via a web server based on a user's request, and the web server can act as an intermediary to notify the user of any available services. When a user requests a desired service from the web server, the web server sends it to the streaming media server, which then sends the multimedia data to the user. At this point, the content streaming system may include a separate control server, which in this case can control the commands / responses between devices within the content streaming system.

[0180] A streaming media server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming media server can store the bitstream for a specific time period.

[0181] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet computers, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.

[0182] In the above content streaming media system, each server can operate as a distributed server, in which case the data received from each server can be distributed and processed.

[0183] The above embodiments can be implemented in the same or corresponding manner in encoding and decoding devices. Furthermore, at least one of the above embodiments or a combination of at least one of the above embodiments can be used to encode / decode images.

[0184] The order in which the above-described embodiments are applied in the encoding and decoding devices may differ. Alternatively, the order in which the above-described embodiments are applied in the encoding and decoding devices may be the same.

[0185] The above implementation methods can be performed for each of the luminance and chrominance signals. Alternatively, the above implementation methods for luminance and chrominance signals can be performed identically.

[0186] In the above embodiments, the method is based on a flowchart description having a series of steps or units. However, this disclosure is not limited to the order of the steps, but some steps may be performed simultaneously with other steps or in a different order. Furthermore, those skilled in the art should understand that the steps in the flowchart are not exclusive to each other, and other steps may be added to the flowchart or some steps may be deleted from the flowchart without affecting the scope of this disclosure.

[0187] The implementation can be carried out in the form of program instructions executable by various computer components and recorded in a computer-readable recording medium. The computer-readable recording medium may include individual program instructions, data files, data structures, etc., or combinations thereof. The program instructions recorded in the computer-readable recording medium may be specifically designed and constructed for this disclosure, or may be well known to those skilled in the art of computer software.

[0188] The bitstream generated by the encoding method according to the above embodiments can be stored in a non-transient computer-readable recording medium. Furthermore, the bitstream stored in the non-transient computer-readable recording medium can be decoded using the decoding method according to the above embodiments.

[0189] Examples of computer-readable recording media include magnetic recording media such as hard disks, floppy disks, and magnetic tapes; optical data storage media such as CD-ROMs or DVD-ROMs; magneto-optical media such as floppy disks; and hardware devices specifically configured to store and implement program instructions, such as read-only memory (ROM), random access memory (RAM), flash memory, etc. Examples of program instructions include not only machine language code formatted by a compiler but also high-level language code that can be implemented by a computer using an interpreter. The hardware device may be configured to operate by one or more software modules to perform process processing according to this disclosure, and vice versa.

[0190] Although this disclosure has been described with reference to specific items such as detailed elements and limited embodiments and drawings, these embodiments and drawings are provided only to aid in a more complete understanding of the invention, and this disclosure is not limited to the above embodiments. Those skilled in the art to which this disclosure pertains will understand that various modifications and changes can be made from the above description.

[0191] Therefore, the spirit of this disclosure should not be limited to the above-described embodiments, and the entire scope of the appended claims and their equivalents shall fall within the scope and spirit of the invention.

[0192] Industrial applicability

[0193] This disclosure can be used in devices for encoding / decoding images and in recording media for storing bit streams.

Claims

1. An image decoding method, comprising: Derive initial motion vectors from the adjacent blocks of the current block to derive candidate sub-block time motion vectors for deriving the sub-blocks of the current block; The initial motion vector is corrected within a predetermined search range to derive the final motion vector; based on the final motion vector, a co-position reference block referenced by the current block is determined in the co-position reference image of the current block. Furthermore, the candidate time motion vectors of the sub-blocks of the current block are derived from the co-located reference block.

2. The image decoding method according to claim 1, wherein, The neighboring block of the current block used to derive the initial motion vector is one of a plurality of candidate neighboring blocks of the current block.

3. The image decoding method according to claim 2, wherein, The plurality of candidate neighboring blocks includes at least one of the following: one or more neighboring blocks adjacent to the lower left portion of the current block, one or more neighboring blocks adjacent to the upper right portion of the current block, and one or more neighboring blocks adjacent to the upper left portion of the current block.

4. The image decoding method according to claim 2, wherein, The neighboring blocks of the current block used to derive the initial motion vector are determined as the first available candidate neighboring block among the plurality of candidate neighboring blocks of the current block according to a predetermined search order.

5. The image decoding method according to claim 4, wherein, If the co-position reference image of the current block is the same as the reference image used to predict the candidate neighboring block, then the candidate neighboring block is determined to be available.

6. The image decoding method according to claim 4, wherein, Even if the co-located reference image of the current block is different from the reference image used to predict the candidate neighboring block, the candidate neighboring block is determined to be available.

7. The image decoding method according to claim 2, wherein, The neighboring blocks of the current block used to derive the initial motion vector are determined as the candidate neighboring blocks with the smallest distortion according to template matching among the plurality of candidate neighboring blocks of the current block, wherein the template matching is performed by determining the distortion of the reference template based on a comparison between a current template adjacent to the current block and a reference template adjacent to the candidate neighboring block.

8. The image decoding method according to claim 7, wherein, Template matching is performed on candidate neighbor blocks such that the co-location reference image of the current block is the same as the reference image used to predict the candidate neighbor blocks, and wherein the neighbor blocks of the current block are determined to be the candidate neighbor blocks with the least distortion among the candidate neighbor blocks for which template matching has been performed.

9. The image decoding method according to claim 2, wherein, The neighboring blocks of the current block used to derive the initial motion vector are determined as the candidate neighboring blocks with the smallest distortion based on block matching among the plurality of candidate neighboring blocks of the current block.

10. The image decoding method according to claim 9, wherein, The neighboring blocks of the current block are determined by block matching reference block information, which indicates the candidate neighboring blocks among the plurality of candidate neighboring blocks of the current block for deriving the co-position reference block.

11. The image decoding method according to claim 10, wherein, The block matching reference block information is encoded into codewords of different lengths according to a predetermined order for the plurality of candidate adjacent blocks.

12. The image decoding method according to claim 11, wherein, The predetermined order of the plurality of candidate adjacent blocks is determined based on the distortion determined by template matching of the plurality of candidate adjacent blocks.

13. The image decoding method according to claim 1, wherein, During the process of deriving the final motion vector, the motion vector of the reference template with the minimum distortion according to the template matching within the predetermined search range is determined as the final motion vector, and the template matching is performed by determining the distortion of the reference template based on a comparison between the current template adjacent to the current block and the reference template.

14. The image decoding method according to claim 1, wherein, In the process of deriving the final motion vector, the motion vector of the block matching reference block that indicates the minimum distortion according to the block matching within the predetermined search range is determined as the final motion vector. The block matching is performed by determining the distortion of the block matching reference block based on a comparison between the neighboring blocks of the current block and the block matching reference block, and the block matching reference block is a block whose distance from the reference block referenced by the neighboring blocks of the current block is within the predetermined search range.

15. The image decoding method according to claim 1, wherein, In the process of deriving the final motion vector, the final motion vector is derived by correcting the initial motion vector within the predetermined search range based on differential motion vector information indicating the difference between the initial motion vector and the final motion vector.

16. An image encoding method, comprising: Derive initial motion vectors from the adjacent blocks of the current block to derive candidate sub-block time motion vectors for deriving the sub-blocks of the current block; The initial motion vector is corrected within a predetermined search range to derive the final motion vector; based on the final motion vector, a co-position reference block referenced by the current block is determined in the co-position reference image of the current block. Furthermore, the candidate time motion vectors of the sub-blocks of the current block are derived from the co-located reference block.

17. A computer-readable recording medium for storing a bitstream generated by an image encoding method, the image encoding method comprising: Derive initial motion vectors from the adjacent blocks of the current block to derive candidate sub-block time motion vectors for deriving the sub-blocks of the current block; The initial motion vector is corrected within a predetermined search range to derive the final motion vector; based on the final motion vector, a co-position reference block referenced by the current block is determined in the co-position reference image of the current block. Furthermore, the candidate time motion vectors of the sub-blocks of the current block are derived from the co-located reference block.

18. A method for transmitting a bitstream generated by an image encoding method, the method comprising: The image is encoded based on the image encoding method described above; The bitstream containing the encoded image is transmitted, wherein the image encoding method includes: deriving an initial motion vector from neighboring blocks of the current block for deriving sub-block temporal motion vector candidates for sub-blocks of the current block; deriving a final motion vector by correcting the initial motion vector within a predetermined search range; determining a co-location reference block referenced by the current block in a co-location reference image of the current block based on the final motion vector; and deriving the sub-block temporal motion vector candidates for the sub-blocks of the current block from the co-location reference block.