Image encoding / decoding method and apparatus, and recording medium for storing bit stream

By deriving and refining block vectors during image encoding/decoding, the image encoding/decoding method is optimized, solving the problem of increased data volume in high-resolution, high-quality images and achieving more efficient encoding/decoding and prediction.

CN120958822APending Publication Date: 2025-11-14HYUNDAI MOTOR CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480025089.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-09
Filing Date
2024-05-30
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing image encoding/decoding technologies face increased data volume and high transmission and storage costs when processing high-resolution, high-quality images, necessitating improvements in encoding/decoding efficiency.

Method used

By deriving the initial block vector of the current block, deriving the search range based on the initial block vector, and refining the block vector within the search range, the image encoding/decoding process is optimized by utilizing intra-prediction mode and template intra-mode derivation methods.

Benefits of technology

It improves the efficiency and prediction accuracy of image encoding/decoding, and reduces transmission and storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120958822A_ABST
    Figure CN120958822A_ABST
Patent Text Reader

Abstract

The invention provides an image encoding / decoding method and apparatus, a recording medium for storing a bit stream, and a transmission method. The image decoding method may include: deriving an initial block vector of a current block; deriving a search range based on the initial block vector; and deriving a refined block vector based on the search range. The refined block vector is derived from a difference between a first prediction signal generated based on candidate block vectors within the search range and a second prediction signal generated based on an intra prediction mode of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to image encoding / decoding methods, apparatus, and recording media for storing bitstreams. Specifically, this invention relates to image encoding / decoding methods based on decoder-based block vector refinement methods and apparatus, and recording media for storing bitstreams. Background Technology

[0002] In recent years, the demand for high-resolution, high-quality images (such as ultra-high-definition (UHD) images) has been increasing across various application fields. As image data resolution and quality improve, the data volume also increases relative to existing image data. Therefore, both transmission and storage costs increase when using existing wired and wireless broadband lines or other media to transmit image data, or when using existing storage media to store image data. To address these issues arising from improved image data resolution and quality, it is necessary to develop efficient image encoding / decoding technologies to achieve higher resolution and higher image quality.

[0003] The mode that predicts based on the image reconstruction region containing the current block and block vector is called the intra-block copy mode. Summary of the Invention

[0004] Technical issues The purpose of this invention is to provide an image encoding / decoding method and apparatus that has improved encoding / decoding efficiency.

[0005] Furthermore, the object of the present invention is to provide a recording medium for storing a bitstream generated by a method or apparatus for decoding an image according to the present invention.

[0006] Furthermore, the present invention aims to provide a decoder-based block vector refinement method.

[0007] Technical solution An image decoding method according to an embodiment of the present invention may include: deriving an initial block vector for the current block; deriving a search range based on the initial block vector; and deriving a refined block vector based on the search range. The refined block vector can be derived from the difference between a first prediction signal generated based on candidate block vectors within the search range and a second prediction signal generated based on the intra-frame prediction mode of the current block.

[0008] In this image decoding method, the intra-frame prediction mode can be derived using the template-based intra-frame mode derivation (TIMD) method.

[0009] In this image decoding method, the TIMD method can be executed by selecting any candidate mode from the intra-prediction candidate list of the current block based on the template of the current block.

[0010] In this image decoding method, the refined block vector can be a candidate block vector that minimizes the distortion values ​​of the first and second predicted signals.

[0011] In this image decoding method, the search range can be a region with a predetermined shape.

[0012] In this image decoding method, the predetermined shape can be a quadrilateral.

[0013] In this image decoding method, the predetermined shape can be a cross.

[0014] The image decoding method may further include determining whether to refine the block vector of the current block based on at least one of whether the current block is in intra-block copy (IBC) merging mode, the number of samples of the current block, or the size of the current block.

[0015] In this image decoding method, the initial block vector can be a block vector in IBC merging mode.

[0016] An image coding method according to an embodiment of the present invention may include: deriving an initial block vector for the current block; deriving a search range based on the initial block vector; and deriving a refined block vector based on the search range. The refined block vector can be derived from the difference between a first prediction signal generated based on candidate block vectors within the search range and a second prediction signal generated based on the intra-frame prediction mode of the current block.

[0017] A non-transitory computer-readable recording medium according to embodiments of the present invention can store a bitstream generated by an image coding method, the method comprising: deriving an initial block vector for the current block; deriving a search range based on the initial block vector; and deriving a refined block vector based on the search range. The refined block vector can be derived from the difference between a first prediction signal generated based on candidate block vectors within the search range and a second prediction signal generated based on the intra-frame prediction mode of the current block.

[0018] A transmission method according to an embodiment of the present invention may include transmitting a bitstream generated by an image coding method, the image coding method comprising: deriving an initial block vector for the current block; deriving a search range based on the initial block vector; and deriving a refined block vector based on the search range. The refined block vector may be derived from the difference between a first prediction signal generated based on candidate block vectors within the search range and a second prediction signal generated based on the intra-frame prediction mode of the current block.

[0019] The features briefly outlined above are merely illustrative aspects described in detail below and do not limit the scope of this disclosure.

[0020] Beneficial effects According to the present invention, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0021] Furthermore, according to the present invention, a decoder-based block vector refinement method can be provided.

[0022] Furthermore, according to the present invention, prediction accuracy can be improved.

[0023] Those skilled in the art will understand that the effects achievable by the present invention are not limited to the specific description above, and other advantages of the present invention will become clearer from the following detailed description. Attached Figure Description

[0024] Figure 1 This is a block diagram illustrating the configuration of an encoding apparatus according to an embodiment of the present invention.

[0025] Figure 2 This is a block diagram illustrating the configuration of a decoding apparatus according to an embodiment of the present invention.

[0026] Figure 3 This is a schematic diagram illustrating the video encoding system to which the present invention is applicable.

[0027] Figure 4 This is a schematic diagram illustrating the IBC mode according to an embodiment of the present invention.

[0028] Figure 5 This is a schematic diagram illustrating a decoder-based block vector refinement method according to an embodiment of the present invention.

[0029] Figure 6 This is a flowchart illustrating an image decoding method according to an embodiment of the present invention.

[0030] Figure 7 An exemplary content streaming system applicable to embodiments of the present invention is shown. Detailed Implementation

[0031] Invention Model This invention can have various modifications and implementations, which are illustrated in the accompanying drawings and described in detail in the specification. However, this does not mean that the invention is limited to the specific implementations, but should be understood to include all modifications, equivalents, or alternatives within the spirit and technical scope of the invention. Similar reference numerals in the drawings indicate the same or similar functions in various aspects. For clarity of description, the shapes and dimensions of elements in the drawings may be provided by way of example. The detailed description of exemplary embodiments described below refers to the accompanying drawings, which illustrate specific embodiments by way of example. These embodiments are described in sufficient detail to enable those skilled in the art to practice them. It should be understood that the various embodiments differ from one another, but are not necessarily mutually exclusive. For example, the specific shapes, structures, and features described herein may be implemented in other embodiments without departing from the spirit and scope of the invention with respect to one embodiment. It should also be understood that the position or arrangement of the various components in each disclosed embodiment may be changed without departing from the spirit and scope of the embodiments. Therefore, the detailed description described below is not intended to be limiting, and the scope of the exemplary embodiments is defined only by the full scope of the appended claims and the equivalents conferred by such claims (if properly described).

[0032] In this invention, the terms "first," "second," etc., may be used to describe various components, but the components should not be limited by these terms. These terms are only used to distinguish one component from another. For example, without departing from the scope of this invention, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component. The term is and / or includes a combination of multiple related descriptive items or any item from multiple related descriptive items.

[0033] The components shown in the embodiments of the present invention are depicted independently to indicate different functional characteristics, which does not mean that each component constitutes a separate hardware or software configuration unit. That is, for ease of explanation, each component is listed and treated as a separate component, and at least two components can be combined to form a single component, or a component can be broken down into multiple components to perform functions. Furthermore, implementations of component integration and implementations of each component being broken down are also included within the scope of the present invention, as long as they do not depart from the essence of the present invention.

[0034] The terminology used in this invention is for descriptive purposes only and is not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, some components of this invention are not essential for performing the basic functions of the invention, but may be optional components used only to improve performance. This invention can be implemented by excluding components used only to improve performance and including only the essential components for implementing the essence of the invention, and structures that include only the essential components and exclude optional components used only to improve performance are also included within the scope of this invention.

[0035] In one implementation, the term "at least one" may refer to a numerical value greater than or equal to 1, such as 1, 2, 3, and 4. In one implementation, the term "a plurality of" may refer to a numerical value greater than or equal to 2, such as 2, 3, and 4.

[0036] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings. In describing the embodiments of this specification, if it is determined that a detailed description of a related known configuration or function would obscure the subject matter of this specification, such a detailed description will be omitted, and the same reference numerals will be used for the same components in the drawings, and repeated descriptions of the same components will be omitted.

[0037] Terminology Explanation In the following text, "image" can refer to a single frame that makes up a video, or it can refer to the video itself. For example, "encoding and / or decoding of an image" can refer to "encoding and / or decoding of a video," or it can refer to "encoding and / or decoding of one of the images that make up a video."

[0038] In the following text, "moving image" and "video" have the same meaning and can be used interchangeably. Furthermore, the target image can be an encoded target image that serves as the encoding target and / or a decoded target image that serves as the decoding target. Additionally, the target image can be an input image input to the encoding device or an input image input to the decoding device. Here, the meaning of target image can be the same as the current image.

[0039] In the following text, "encoder" and "image encoding device" have the same meaning and can be used interchangeably.

[0040] In the following text, "decoder" and "image decoding device" have the same meaning and can be used interchangeably.

[0041] In the following text, "image," "picture," "frame," and "screen" have the same meaning and can be used interchangeably.

[0042] In the following text, "target block" can refer to an encoding target block that is the target of encoding and / or a decoding target block that is the target of decoding. Furthermore, a target block can be the current block, i.e., the target block for the current encoding and / or decoding. For example, "target block" and "current block" have the same meaning and can be used interchangeably.

[0043] In the following text, "block" and "unit" have the same meaning and can be used interchangeably. Furthermore, "unit" can refer to a block containing a luma component block and a corresponding chroma component block, thus distinguishing it from a block. For example, a coding tree unit (CTU) can consist of a luma component (Y) coding tree block (CTB) and two associated chroma component (Cb, Cr) coding tree blocks.

[0044] In the following text, "sample," "image element," and "pixel" have the same meaning and can be used interchangeably. Here, a sample can represent the basic unit that makes up a block.

[0045] In the following text, "inter-frame" and "inter-screen" have the same meaning and can be used interchangeably.

[0046] In the following text, "intra-frame" and "intra-screen" have the same meaning and can be used interchangeably.

[0047] Figure 1 This is a block diagram illustrating the configuration of an encoding apparatus according to an embodiment of the present invention.

[0048] The encoding device 100 can be an encoder, a video encoding device, or an image encoding device. The video can include one or more images. The encoding device 100 can encode one or more images sequentially.

[0049] refer to Figure 1 The encoding device 100 may include an image segmentation unit 110, an intra-frame prediction unit 120, a motion prediction unit 121, a motion compensation unit 122, a switch 115, a subtractor 113, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 117, a filtering unit 180, and a reference image buffer 190.

[0050] Furthermore, the encoding device 100 can generate a bitstream containing encoded information by encoding the input image and output the generated bitstream. The generated bitstream can be stored in a computer-readable recording medium or streamed via a wired / wireless transmission medium.

[0051] Image segmentation unit 110 can segment the input image into various forms to improve the efficiency of video encoding / decoding. That is, the input video consists of multiple images, and individual images can be segmented and processed hierarchically to improve compression efficiency, achieve parallel processing, etc. For example, a single image can be segmented into one or more tiles or slices, and then further segmented into multiple CTUs (Coding Tree Units). Alternatively, a single image can first be segmented into multiple sub-images defined as groups of rectangular slices, and then each sub-image can be segmented into tiles / slices. In this case, the sub-images can be used to support partially independent encoding / decoding and transmission of images. Since multiple sub-images can be reconstructed individually, it has the advantage of ease of editing in applications that configure multi-channel input as a single image. Furthermore, tiles can be horizontally segmented into blocks. In this case, blocks can serve as the basic unit for parallel processing within an image. Additionally, a CTU can be recursively segmented into a quadtree (QT), and the segmented terminal node can be defined as a coding unit (CU). A CU can be segmented into prediction units (PUs) and transform units (TUs) to perform prediction and segmentation. Simultaneously, the CU can be used as the prediction unit and / or the transform unit itself. Here, for flexible segmentation, each CTU can be recursively segmented into a multi-type tree (MTT) and a quadtree (QT). Segmenting a CTU into a multi-type tree can start from the terminal node of the QT. The MTT can consist of a binary tree (BT) and a ternary tree (TT). For example, the MTT structure can be divided into a vertical binary segmentation mode (SPLIT_BT_VER), a horizontal binary segmentation mode (SPLIT_BT_HOR), a vertical ternary segmentation mode (SPLIT_TT_VER), and a horizontal ternary segmentation mode (SPLIT_TT_HOR). Furthermore, during segmentation, the minimum block size (MinQTSize) of the quadtree for the luma block can be set to 16×16, the maximum block size (MaxBtSize) of the binary tree can be set to 128×128, and the maximum block size (MaxTtSize) of the ternary tree can be set to 64××64. Furthermore, the minimum block size (MinBtSize) of the binary tree and the minimum block size (MinTtSize) of the ternary tree can be specified as 4×4, and the maximum depth (MaxMttDepth) of the multi-type tree can be specified as 4. Additionally, to improve the coding efficiency of I-slices, a dual-tree approach can be adopted, which uses a CTU partitioning structure for both luma and chroma components. On the other hand, in P-slices and B-slices, the luma and chroma CTBs (coding tree blocks) within the CTU can be partitioned into single trees sharing a coding tree structure.

[0052] The encoding device 100 can encode the input image in intra-frame mode and / or inter-frame mode. Alternatively, the encoding device 100 can encode the input image in a third mode other than intra-frame mode and inter-frame mode (e.g., IBC mode, palette mode, etc.). However, if the third mode has similar functional characteristics to the intra-frame mode or inter-frame mode, it can be classified as an intra-frame mode or inter-frame mode for ease of explanation. In this invention, the third mode will only be classified and described separately when it is necessary to specifically describe it.

[0053] When using intra-frame mode as the prediction mode, switch 115 can switch to intra-frame mode; when using inter-frame mode as the prediction mode, switch 115 can switch to inter-frame mode. Here, intra-frame mode can refer to intra-frame prediction mode, and inter-frame mode can refer to inter-frame prediction mode. The encoding device 100 can generate prediction blocks for input blocks of the input image. Furthermore, after generating prediction blocks, the encoding device 100 can encode residual blocks using the residuals between the input blocks and the prediction blocks. The input image can be called the current image, i.e., the current encoding target. The input block can be called the current block, i.e., the current encoding target or encoding target block.

[0054] When the prediction mode is intra-frame mode, the intra-frame prediction unit 120 can use samples from encoded / decoded blocks surrounding the current block as reference samples. The intra-frame prediction unit 120 can use the reference samples to perform spatial prediction on the current block, or generate prediction samples for the input block through spatial prediction. Here, intra-frame prediction can refer to in-screen prediction.

[0055] Intra-frame prediction methods can apply non-directional prediction modes (such as DC mode and planar mode) as well as directional prediction modes (such as 65 directions). Here, intra-frame prediction methods can be represented as intra-frame prediction modes or in-screen prediction modes.

[0056] When the prediction mode is inter-frame mode, the motion prediction unit 121 can retrieve the region in the reference image that best matches the input block during motion prediction and derive the motion vector using the retrieved region. In this case, the search region can be used as the region. The reference image can be stored in the reference image buffer 190. Here, it can be stored in the reference image buffer 190 when encoding / decoding the reference image.

[0057] The motion compensation unit 122 can generate a prediction block for the current block by using motion vectors for motion compensation. Here, inter-frame prediction can refer to inter-screen prediction or motion compensation.

[0058] When the value of the motion vector is not an integer, the motion prediction unit 121 and the motion compensation unit 122 can generate prediction blocks by applying an interpolation filter to a portion of the reference image. To perform inter-frame prediction or motion compensation, the motion prediction and motion compensation modes of the prediction units included in the coding unit can be determined based on the coding unit, such as skip mode, merge mode, advanced motion vector prediction (AMVP) mode, or intra-block copy (IBC) mode, and inter-frame prediction or motion compensation can be performed according to each mode.

[0059] Furthermore, based on the aforementioned inter-frame prediction methods, other modes can be applied, including AFFINE mode based on sub-PU prediction, SbTMVP (sub-block temporal motion vector prediction), MMVD (merged with MVD) mode based on PU prediction, and Geometric Partitioning (GPM). Additionally, to improve the performance of each mode, other modes can be applied, such as HMVP (history-based MVP), PAMVP (pairwise average MVP), CIIP (intra-frame / inter-frame combined prediction), AMVR (adaptive motion vector resolution), BDOF (bidirectional optical flow), BCW (bidirectional prediction with CU weights), LIC (local illumination compensation), TM (template matching), and OBMC (overlapping block motion compensation).

[0060] Among them, the AFFINE mode is used in both AMVP and MERGE modes and is a highly efficient coding technique. In existing video coding standards, because MC (motion compensation) only considers the parallel motion of blocks, its drawback is that it cannot correctly compensate for actual motions such as zooming in / out and rotation. To compensate for this, a four-parameter affine motion model with two control point motion vectors (CPMV) and a six-parameter affine motion model with three control point motion vectors can be used and applied to inter-frame prediction. Here, CPMV is a vector representing the affine motion model of one of the top-left, top-right, and bottom-left positions of the current block.

[0061] Subtractor 113 can generate a residual block using the difference between the input block and the prediction block. The residual block can be called a residual signal. The residual signal can represent the difference between the original signal and the prediction signal. Alternatively, the residual signal can be a signal generated by transforming or quantizing the difference between the original signal and the prediction signal, or a signal generated by transforming and quantizing the difference between the original signal and the prediction signal. The residual block can be a residual signal on a block-by-block basis.

[0062] Transform unit 130 can generate transform coefficients by transforming the residual block and output the generated transform coefficients. Here, the transform coefficients can be coefficient values ​​generated by transforming the residual block. When a transform skip mode is applied, transform unit 130 can skip the transformation of the residual block.

[0063] Quantization levels can be generated by applying quantization to the transform coefficients or residual signals. In the following implementation, quantization levels may also be referred to as transform coefficients.

[0064] For example, a 4×4 lumen residual block generated by intra-frame prediction is transformed using DST (Discrete Sine Transform)-based basis vectors, and then the remaining residual blocks are transformed using DCT (Discrete Cosine Transform)-based basis vectors. Furthermore, RQT (Residual Quadtree) technology is used to divide the transformed blocks into quadtree shapes. After transforming and quantizing each transformed block segmented by RQT, a Coded Block Flag (CBF) can be sent when all coefficients become 0 to improve coding efficiency.

[0065] Another alternative is to apply the Multiple Transform Selection (MTS) technique, which selectively uses multiple transform bases for transformation. That is, instead of segmenting the CU into TUs using RQT, a function similar to TU segmentation can be performed using Sub-Block Transform (SBT). Specifically, SBT applies only to inter-frame prediction blocks. Unlike RQT, the current block can be segmented into 1 / 2 or 1 / 4 blocks vertically or horizontally, and then the transformation is performed on only one of these blocks. For example, in a vertical segmentation, the transformation can be performed on the leftmost or rightmost block; in a horizontal segmentation, the transformation can be performed on the topmost or bottommost block.

[0066] In addition, LFNST (Low Frequency Non-Separable Transform) can be applied. This is a secondary transform technique that performs an additional transform on the residual signal transformed to the frequency domain by DCT or DST. LFNST also performs an additional transform on the 4×4 or 8×8 low-frequency region in the upper left corner, thereby concentrating the residual coefficients in the upper left corner.

[0067] The quantization unit 140 can quantize the transform coefficients or residual signal according to the quantization parameters (QP) to generate a quantization level and output the generated quantization level. Here, the quantization unit 140 can use a quantization matrix to quantize the transform coefficients.

[0068] For example, quantizers with QP values ​​from 0 to 51 can be used. Alternatively, if the image size is large and high coding efficiency is required, QP values ​​from 0 to 63 can be used. Furthermore, the DQ (correlated quantization) method, which uses two quantizers (instead of one), can also be employed. DQ uses two quantizers (e.g., Q0 and Q1) for quantization, but even without conveying information about using a specific quantizer, the quantizer used for the next transform coefficient can be selected based on the current state through a state transition model.

[0069] The entropy coding unit 150 can generate a bitstream by entropy coding based on the value calculated by the quantization unit 140 or the probability distribution of the coding parameter values ​​calculated during encoding, and then output the bitstream. The entropy coding unit 150 can entropy code image sample information and information used for decoding the image. For example, the information used for decoding the image may include syntax elements.

[0070] When applying entropy coding, symbols are represented by allocating fewer bits to symbols with high occurrence probabilities and more bits to symbols with low occurrence probabilities, thereby reducing the size of the bitstream to be encoded. The entropy coding unit 150 can perform entropy coding using methods such as exponential Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). For example, the entropy coding unit 150 can perform entropy coding using a variable-length coding (VLC) table. Furthermore, the entropy coding unit 150 can derive a binarization method for the target symbol and a probability model for the target symbol / bit, and perform arithmetic coding using the derived binarization method and context model.

[0071] Relatedly, when applying CABAC, to reduce the size of the probability table stored in the decoding device, the table probability update method can be changed to a table update method using simple equations. Furthermore, two different probability models can be used to obtain more accurate symbol probability values.

[0072] In order to encode the transform coefficient levels (quantization levels), the entropy coding unit 150 can convert the two-dimensional block form coefficients into one-dimensional vector form by the transform coefficient scanning method.

[0073] The encoding parameters may include information (flags, indexes, etc.) encoded in the encoding device 100 and transmitted to the decoding device 200 via signals, such as syntax elements, as well as information derived during the encoding or decoding process, and may represent the information required when encoding or decoding an image.

[0074] In this paper, sending a signal flag or index can indicate that the corresponding flag or index is entropy encoded in the encoder and included in the bitstream, or it can indicate that the corresponding flag or index is entropy decoded from the bitstream in the decoder.

[0075] The encoded current image can be used as a reference image for another image to be processed later. Therefore, the encoding device 100 can reconstruct or decode the encoded current image again and store the reconstructed or decoded image as a reference image in the reference image buffer 190.

[0076] The quantization level can be dequantized in dequantization unit 160 or inverse transformed in inverse transform unit 170. The coefficients after dequantization and / or inverse transform can be added to the prediction block via adder 117. Here, the coefficients after dequantization and / or inverse transform can refer to the coefficients of at least one of the dequantization and / or inverse transform operations, or they can refer to the reconstructed residual block. Dequantization unit 160 and inverse transform unit 170 can be performed as the inverse process of quantization unit 140 and transform unit 130.

[0077] The reconstructed blocks can be processed by filtering unit 180. Filtering unit 180 can use all or some filtering techniques, applying deblocking filters, sample adaptive offset (SAO), adaptive loop filter (ALF), bilateral filter (BIF), luma mapping and chroma scaling (LMCS), etc., to the reconstructed samples, reconstructed blocks, or reconstructed images. Filtering unit 180 can be referred to as a loop filter. In this case, loop filter is also used as a name, but LMCS is not included.

[0078] Deblocking filters can eliminate block distortion generated at the boundaries between blocks. To determine whether to apply a deblocking filter, samples from several rows or columns contained in the current block can be used. When applying a deblocking filter to a block, different filters can be applied depending on the desired deblocking intensity.

[0079] To use sample-adaptive offsets to compensate for coding errors, appropriate offset values ​​can be added to the sample values. Sample-adaptive offsets can refine the offset between the deblocked image and the original image on a sample-by-sample basis. This can be achieved by dividing the samples in the image into a predetermined number of regions, determining the regions to which the offset will be applied, and then applying the offset to the determined regions; alternatively, the edge information of each sample can be considered when applying the offset.

[0080] For images that have already undergone deblocking, the bilateral filter (BIF) can also refine the offset from the original image sample by sample.

[0081] Adaptive loop filters (ALFs) can perform filtering based on a comparison between the reconstructed image and the original image. The samples contained in the image can be divided into predetermined groups, the filter to be applied to each group can be determined, and differential filtering can be performed on each group. Information regarding whether to apply the ALF can be signaled by the coding unit (CU), and the form and coefficients of the adaptive loop filter applied to each block can vary.

[0082] In LMCS (Luminosity Mapping and Chroma Scaling), Luminosity Mapping (LM) refers to remapping luminance values ​​using a piecewise linear model, while Chroma Scaling (CS) refers to scaling the residual chroma components based on the average luminance value of the predicted signal. Specifically, LMCS can be used as an HDR enhancement technique that reflects the characteristics of HDR (High Dynamic Range) images.

[0083] The reconstructed blocks or reconstructed image processed by filtering unit 180 can be stored in reference image buffer 190. The reconstructed blocks processed by filtering unit 180 can be a portion of the reference image. That is, the reference image is a reconstructed image composed of the reconstructed blocks processed by filtering unit 180. The stored reference image can later be used for inter-frame prediction or motion compensation.

[0084] Figure 2 This is a block diagram illustrating the configuration of a decoding apparatus according to an embodiment of the present invention.

[0085] The decoding device 200 can be a decoder, a video decoding device, or an image decoding device.

[0086] refer to Figure 2 The decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, a motion compensation unit 250, an adder 201, a switch 203, a filtering unit 260, and a reference image buffer 270.

[0087] The decoding device 200 can receive a bitstream output from the encoding device 100. The decoding device 200 can receive a bitstream stored in a computer-readable recording medium, or it can receive a bitstream streamed via a wired / wireless transmission medium. The decoding device 200 can decode the bitstream in intra-frame mode or inter-frame mode. Furthermore, the decoding device 200 can generate a reconstructed image or a decoded image through decoding, and output the reconstructed image or the decoded image.

[0088] When the prediction mode used for decoding is intra-frame mode, switch 203 can switch to intra-frame mode. Alternatively, when the prediction mode used for decoding is inter-frame mode, switch 203 can switch to inter-frame mode.

[0089] Decoding device 200 can obtain a reconstructed residual block and generate a prediction block by decoding the input bitstream. When the reconstructed residual block and prediction block are obtained, decoding device 200 can generate a reconstructed block by adding the reconstructed residual block and prediction block; this reconstructed block serves as the decoding target. The decoding target block can be referred to as the current block.

[0090] Entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream according to the probability distribution. The generated symbols may include symbols in the form of quantization levels. Here, the entropy decoding method can be the inverse process of the entropy encoding method described above.

[0091] The entropy decoding unit 210 can convert one-dimensional vector coefficients into two-dimensional block coefficients by the transform coefficient scanning method to decode the transform coefficient level (quantization level).

[0092] The quantization level can be dequantized in the dequantization unit 220 or inverse transformed in the inverse transform unit 230. The quantization level can be the result of dequantization and / or inverse transform, and can be generated as a reconstruction residual block. Here, the dequantization unit 220 can apply the quantization matrix to the quantization level. The dequantization unit 220 and inverse transform unit 230 applied to the decoding device can employ the same techniques as the dequantization unit 160 and inverse transform unit 170 applied to the encoding device described above.

[0093] When using intra-frame mode, intra-frame prediction unit 240 can generate a prediction block by performing spatial prediction on the current block, which uses sample values ​​from decoded blocks surrounding the target block. Intra-frame prediction unit 240 applied to the decoding apparatus can employ the same techniques as intra-frame prediction unit 120 applied to the encoding apparatus described above.

[0094] When using inter-frame mode, motion compensation unit 250 can generate a prediction block by performing motion compensation on the current block, using motion vectors and a reference image stored in reference image buffer 270. When the value of the motion vector is not an integer, motion compensation unit 250 can generate a prediction block by applying an interpolation filter to a portion of the reference image. To perform motion compensation, the motion compensation method of the prediction unit included in the corresponding coding unit can be determined based on the coding unit: skip mode, merge mode, AMVP mode, or current image reference mode, and motion compensation can be performed according to each mode. The motion compensation unit 250 applied to the decoding device can employ the same technique as the motion compensation unit 122 applied to the coding device described above.

[0095] Adder 201 generates a reconstructed block by adding the reconstructed residual block and the predicted block. Filtering unit 260 can apply at least one of inverse LMCS, deblocking filter, sample adaptive offset, and adaptive loop filter to the reconstructed block or reconstructed image. Filtering unit 260 applied to the decoding device can employ the same filtering technique as that used by filtering unit 180 applied to the encoding device described above.

[0096] The filtering unit 260 can output a reconstructed image. The reconstructed blocks or reconstructed image can be stored in the reference image buffer 270 and used for inter-frame prediction. The reconstructed blocks processed by the filtering unit 260 can be used as part of the reference image. That is, the reference image can be a reconstructed image composed of the reconstructed blocks processed by the filtering unit 260. The stored reference image can later be used for inter-frame prediction or motion compensation.

[0097] Figure 3 This is a schematic diagram illustrating the video encoding system to which the present invention is applicable.

[0098] According to one embodiment, a video encoding system may include an encoding device 10 and a decoding device 20. The encoding device 10 may transmit encoded video and / or image information or data to the decoding device 20 in the form of a file or stream via a digital storage medium or network.

[0099] The encoding apparatus 10 according to the embodiments may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. The decoding apparatus 20 according to the embodiments may include a receiving unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The receiving unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, and the display unit may be configured as a separate device or an external component.

[0100] The video source generation unit 11 can acquire video / images through a process of capturing, synthesizing, or generating video / images. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include, for example, a computer, tablet computer, and smartphone, etc., and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc., in which case the video / image capture process can be replaced by a process of generating related data.

[0101] Encoding unit 12 can encode the input video / image. Encoding unit 12 can perform a series of processes, such as prediction, transformation, and quantization, to improve compression and encoding efficiency. Encoding unit 12 can output encoded data (encoded video / image information) in the form of a bitstream. Detailed configuration of encoding unit 12 can also be as described above. Figure 1 The encoding device 100 is configured in the same way.

[0102] The transmitting unit 13 can output encoded video / image information or data as a bitstream, and transmit it to the receiving unit 21 of the decoding device 20 in the form of a file or stream via a digital storage medium or network. The digital storage medium can include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit 13 can include elements for generating media files according to a predetermined file format, and can also include elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and send it to the decoding unit 22.

[0103] Decoding unit 22 can decode video / images by performing a series of steps (e.g., dequantization, inverse transform, and prediction) corresponding to the operations of encoding unit 12. The detailed configuration of decoding unit 22 can also be as described above. Figure 2 The decoding device 200 is configured in the same way.

[0104] Rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed through the display unit.

[0105] The following will refer to Figures 4 to 6 This describes a decoder-based block vector refinement (DBVR) method according to an embodiment of the present invention.

[0106] In this specification, block vector refers to the motion vector used in IBC (Intra-Block Copy) mode.

[0107] Here, IBC mode is a prediction technique that performs block matching (BM) within the reconstructed region to search for the block most similar to the current block (e.g., a coding unit) and uses it as the prediction block. In other words, the block vector (BV) can represent the spatial displacement relative to the spatial location of the searched block.

[0108] Figure 4 This is a schematic diagram illustrating the IBC mode according to an embodiment of the present invention.

[0109] refer to Figure 4 Based on the block vector 420 of the current block 410, the matching block 430 can be derived within a predefined search range R1, R2, R3, and R4 of the reconstructed region of the current image 400. Here, information about the block vector 420 can be sent from the encoder to the decoder via a bitstream.

[0110] Meanwhile, the predefined search range R1 can be defined as the CTU (Code Tree Unit) containing the current block, R2 as the top-left CTU, R3 as the top CTU, and R4 as the left-side CTU. However, this is just an example, and the search range can be set to any size and position other than those described above.

[0111] IBC mode can be divided into IBC merging (intra-block copy merging) mode and IBC motion vector prediction mode.

[0112] IBC merging mode can directly derive block vector information from a candidate list generated from the block vector information of adjacent blocks based on IBC mode or IntraTMP mode.

[0113] On the other hand, the IBC motion vector prediction mode can derive block vector information by adding block vector difference (BVD) information to the block vector information derived from the candidate list mentioned above. Here, the block vector difference information can be sent from the encoder to the decoder via a bitstream.

[0114] The decoder-based block vector refinement method according to an embodiment of the present invention is a method for refining the block vector (BV) derived from the IBC merging pattern through a block matching (BM) based block vector search process without sending additional information during the decoding process.

[0115] Here, block matching-based block vector search is a method for searching candidate block vectors that minimizes the distortion between the predicted signal (or predicted block) generated from candidate block vectors within a search range around the initial block vector derived using the IBC merging mode and the predicted signal (or predicted block) generated by template-based intra-mode derivation (TIMD). The searched candidate block vectors can be set as refined block vectors.

[0116] According to the TIMD method of the present invention, the directionality of all candidate modes in the Most Probable Mode (MPM) list is applied to the reference pixels of the current template to generate each prediction template, and the sum of absolute transformation differences (SATD) between the pixels of the generated prediction template and the pixels of the previously reconstructed template is calculated. The mode with the smallest sum of absolute transformation differences among all candidate modes in the MPM list can be determined as the intra-prediction mode.

[0117] Meanwhile, the TIMD method can be applied to P candidate modes in the MPM list. Here, P is any positive integer greater than or equal to 1 and less than or equal to the number of MPM candidate modes (1 ≤ P ≤ the number of MPM candidate modes). In this case, the mode with the smallest sum of absolute transform differences among the P candidate modes in the MPM list can be determined as the intra-prediction mode.

[0118] If P is 1, the first mode in the MPM list is the intra-frame mode for the current block. The intra-frame prediction mode determined by TIMD can be used to generate the prediction signal.

[0119] Meanwhile, the template used when deriving the intra-prediction mode of the current block using the TIMD method can be set to the same as... Figure 5The templates used in intra-frame template matching (the current template) have the same area, or they can be set to different areas. Alternatively, the size and shape of the templates used in the TIMD method can be arbitrarily determined.

[0120] Meanwhile, the above-mentioned TIMD method uses the sum of absolute transform differences (SATD) when deriving intra-predictive modes based on templates using candidate modes in the MPM list. However, according to another embodiment of the present invention, any of various error measurement methods can be selected and used, such as the sum of absolute differences (SAD) or the sum of squared errors (SSE).

[0121] Meanwhile, the TIMD method according to the present invention can perform template matching-based reordering of candidate modes in the MPM list of the current block, and then use the TIMD method to derive intra-prediction modes from the candidate modes in the reordered MPM list.

[0122] Meanwhile, in the TIMD method according to the present invention, the best mode is selected from the candidate modes in the MPM list, but this is only an example, and the best mode can be selected from any selected candidate mode.

[0123] Figure 5 This is a schematic diagram illustrating a decoder-based block vector refinement method according to an embodiment of the present invention.

[0124] refer to Figure 5 If the current block 510 is in IBC merge mode, the block vector (BV) 530 can be derived from the IBC merge mode. The derived block vector 530 is used as the initial block vector, and the surrounding search range (S×T) 540 of the initial block vector can be derived. At this time, the width S and height T of the search range are both any positive integers greater than or equal to 1.

[0125] Candidate block vectors (BV') 550 can be searched within the derived search range 540. Specifically, the candidate block vector (BV') 550 that minimizes the distortion between the predicted signal (matching block') 570 generated by the candidate block vectors (BV') 550 and the predicted signal generated by the intra-prediction mode derived from template-based intra-mode derivation (TIMD) within the search range 540 can be determined as the final block vector.

[0126] Formula 1 represents the final block vector refined using a decoder-based block vector refinement method.

[0127] [Formula 1] BV' = BV + BV offset As shown in Equation 1, the final block vector (BV') can be obtained by offsetting the block vector (BV). offsetThe information added to the initial block vector (BV) derived in the IBC merge mode is used to determine the value, where BV is... offset The information is selected by performing a block vector search process based on block matching within the search scope.

[0128] like Figure 5 As shown, the decoder-based block vector refinement method can select the corresponding block (matching block') indicated by the refined block vector (BV') instead of the corresponding block (matching block) indicated by the initial block vector (BV) derived from the IBC merge mode as the predicted block of the current block (current decoded block).

[0129] Meanwhile, the decoder-based block vector refinement method according to the present invention can be executed when at least one of the following conditions is met.

[0130] - When the current block (e.g., coding unit) is in IBC merge mode - When the number of samples in the current block is M or more (where M is a positive integer greater than or equal to 1). - When the width and / or height of the current block is N or higher (where N is a positive integer greater than or equal to 1) Meanwhile, the decoder-based block vector refinement method according to the present invention can perform integer sample cell search only within a given search range, or it can perform a two-step search process of integer sample cell search and non-integer sample cell search to determine the final block vector.

[0131] For example, decoder-based block vector refinement methods can set the search scope and search process of block matching-based block vector search in different ways.

[0132] Meanwhile, the TIMD method has been described as a method for deriving the intra-prediction mode based on block matching block vector search used in this invention. However, the intra-prediction mode can also be derived using either of the following methods: 1) a method for deriving the intra-prediction mode based on intra-prediction mode information transmitted via the bitstream signal; 2) a decoder-side intra-prediction mode derivation (DIMD) method. Here, the method for deriving the intra-prediction mode based on intra-prediction mode information transmitted via the bitstream signal can be called a general intra-prediction method.

[0133] Figure 6 This is a flowchart illustrating an image decoding method according to an embodiment of the present invention. Figure 6 The image decoding method can be executed by an image decoding device.

[0134] The image decoding device can derive the initial block vector of the current block (S610). Specifically, the initial block vector can be a block vector of the IBC merging mode. That is, it can be a block vector derived through the IBC merging mode.

[0135] Then, the image decoding device can derive the search range based on the initial block vector (S620). Specifically, the search range can be a region with a predetermined shape. Here, the predetermined shape can be a quadrilateral or a cross.

[0136] Then, the image decoding device can derive the refined block vector based on the search range (S630).

[0137] The refined block vector can be derived from the difference between a first prediction signal generated based on candidate block vectors within the search range and a second prediction signal generated based on the intra-prediction mode of the current block. Specifically, the refined block vector can be a candidate block vector that minimizes the distortion values ​​of the first and second prediction signals.

[0138] The intra-prediction mode used to generate the second predicted signal can be derived using the TIMD (Template-Based Intra-Mode Derivation) method. Specifically, the TIMD method can select a candidate mode from the intra-prediction candidate list of the current block based on the template of the current block.

[0139] Then, the image decoding device can determine whether to refine the block vector of the current block based on at least one of whether the current block is in IBC (Intra-Block Copy) merging mode, the number of samples of the current block, or the size of the current block, and refine the block vector only when it is determined that block vector refinement should be performed.

[0140] at the same time, Figure 6 The steps described herein can be performed in the same manner in image encoding methods. Furthermore, this can be achieved by including... Figure 6 The image encoding method described in the steps generates a bitstream. This bitstream can be stored in a non-transitory computer-readable recording medium or transmitted (or streamed).

[0141] Figure 7 An exemplary content streaming system applicable to embodiments of the present invention is shown.

[0142] like Figure 7 As shown, the content streaming media system according to the embodiments of the present invention mainly includes an encoding server, a streaming media server, a web server, a media storage device, a user device, and a multimedia input device.

[0143] The encoding server compresses content received from multimedia input devices such as smartphones, cameras, and CCTV into digital data to generate a bitstream, which is then transmitted to the streaming media server. Alternatively, if the multimedia input devices such as smartphones, cameras, and CCTV directly generate the bitstream, the encoding server can be omitted.

[0144] The bitstream can be generated by applying an image encoding method and / or image encoding device according to an embodiment of the present invention, and the streaming media server can temporarily store the bitstream during the transmission or reception of the bitstream.

[0145] The streaming media server transmits multimedia data to the user's device via a web server based on the user's request. The web server can also act as an intermediary, informing the user of any available services. When a user requests a service from the web server, the web server transmits it to the streaming media server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which can control the commands / responses between devices within the content streaming system.

[0146] A streaming media server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming media server can store the bitstream for a period of time.

[0147] Examples of user devices may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, board PCs, tablets, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.

[0148] Each server in the aforementioned content streaming system can operate as a distributed server, in which case data received from each server can be distributed and processed.

[0149] The above embodiments can be performed in the same or corresponding manner in the encoding and decoding devices. Furthermore, at least one or a combination of at least one of the above embodiments can be used to encode / decode images.

[0150] The order in which the above embodiments are applied in the encoding and decoding devices may be different. Alternatively, the order in which the above embodiments are applied in the encoding and decoding devices may be the same.

[0151] The above implementation methods can be performed separately for luminance signals and chrominance signals. Alternatively, the above implementation methods for luminance signals and chrominance signals can also be performed in the same way.

[0152] In the above embodiments, the method is described based on a flowchart containing a series of steps or units. However, the present invention is not limited to the order of the steps; on the contrary, some steps may be performed simultaneously or in a different order than other steps. Furthermore, those skilled in the art should understand that the steps in the flowchart are not mutually exclusive, and other steps can be added to or some steps can be deleted from the flowchart without affecting the scope of the present invention.

[0153] These implementations can be implemented in the form of program instructions, which can be executed by various computer components and recorded in a computer-readable recording medium. The computer-readable recording medium may include individual program instructions, data files, data structures, etc., or combinations of such program instructions, data files, data structures, etc. The program instructions recorded in the computer-readable recording medium may be specifically designed and constructed for this invention, or may be well known to those skilled in the art of computer software.

[0154] The bitstream generated by the encoding method according to the above embodiments can be stored in a non-transitory computer-readable recording medium. Furthermore, the bitstream stored in the non-transitory computer-readable recording medium can be decoded by the decoding method according to the above embodiments.

[0155] Examples of computer-readable recording media include: magnetic recording media, such as hard disks, floppy disks, and magnetic tapes; optical data storage media, such as CD-ROMs or DVD-ROMs; magnetically optimized media, such as optical-magnetic floppy disks; and hardware devices specifically designed for storing and executing program instructions, such as read-only memory (ROM), random access memory (RAM), flash memory, etc. Examples of program instructions include not only machine language code formatted by a compiler but also high-level language code executable by a computer using an interpreter. The hardware device can be configured to operate by one or more software modules, and vice versa, to execute the processes according to the invention.

[0156] Although the present invention has been described with reference to specific items (e.g., detailed elements) and exemplary embodiments and drawings, these descriptions are provided only to help to more fully understand the invention, which is not limited to the embodiments described above. Those skilled in the art will understand that various modifications and changes can be made based on the above description.

[0157] Therefore, the spirit of the present invention is not limited to the above-described embodiments, and the entire scope of the appended claims and their equivalents should fall within the scope and spirit of the present invention.

[0158] Industrial applicability This invention can be used in apparatuses for encoding / decoding images and in recording media for storing bit streams.

Claims

1. An image decoding method, comprising the following steps: Generate inter-frame prediction blocks for the current block; Derive the initial block vector for the current block; The search range is derived based on the initial block vector; and Based on the search range, derive the refined block vector. The refined block vector is derived from the difference between a first prediction signal generated based on candidate block vectors within the search range and a second prediction signal generated based on the intra-frame prediction mode of the current block.

2. The image decoding method according to claim 1, wherein, The intra-frame prediction mode is derived using the template-based intra-frame mode derivation TIMD method.

3. The image decoding method according to claim 2, wherein, The TIMD method is executed by selecting any candidate mode from the intra-prediction candidate list of the current block based on the template of the current block.

4. The image decoding method according to claim 1, wherein, The refined block vector is a candidate block vector that minimizes the distortion values ​​of the first and second predicted signals.

5. The image decoding method according to claim 1, wherein, The search range is a region with a predetermined shape.

6. The image decoding method according to claim 5, wherein, The predetermined shape is a quadrilateral.

7. The image decoding method according to claim 5, wherein, The predetermined shape is a cross shape.

8. The image decoding method according to claim 1 further includes the following steps: Whether to refine the block vector of the current block is determined based on at least one of the following: whether the current block is in intra-block copy (IBC) merge mode, the number of samples of the current block, or the size of the current block.

9. The image decoding method according to claim 1, wherein, The initial block vector is the block vector in IBC merge mode.

10. An image encoding method, comprising the following steps: Derive the initial block vector for the current block; The search range is derived based on the initial block vector; and Based on the search range, a refined block vector is derived. The refined block vector is derived from the difference between a first prediction signal generated based on candidate block vectors within the search range and a second prediction signal generated based on the intra-frame prediction mode of the current block.

11. A non-transitory computer-readable recording medium for storing a bitstream generated by an image encoding method, the image encoding method comprising the following steps: Derive the initial block vector for the current block; The search range is derived based on the initial block vector; and Based on the search range, a refined block vector is derived. The refined block vector is derived from the difference between a first prediction signal generated based on candidate block vectors within the search range and a second prediction signal generated based on the intra-frame prediction mode of the current block.

12. A method for transmitting a bitstream generated by an image encoding method, the transmission method comprising the following steps: Transmit the bit stream, The image encoding method includes the following steps: Derive the initial block vector for the current block; The search range is derived based on the initial block vector; and Based on the search range, a refined block vector is derived. The refined block vector is derived from the difference between a first prediction signal generated based on candidate block vectors within the search range and a second prediction signal generated based on the intra-frame prediction mode of the current block.