Image encoding / decoding method and apparatus, and recording medium storing a bitstream.
By inducing multiple prediction blocks using advanced prediction modes and correction methods, the method enhances image encoding/decoding efficiency for high-resolution images, addressing the inefficiencies in existing technologies.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2024-04-22
- Publication Date
- 2026-05-13
AI Technical Summary
Existing image compression technologies face challenges in efficiently encoding and decoding high-resolution, high-quality images, particularly in deriving accurate prediction blocks for improved compression efficiency.
The method and apparatus induce multiple prediction blocks for a current block, using various prediction modes and correction methods based on motion vector predictors, such as AMVP and merge modes, and apply bidirectional or template matching corrections to enhance prediction accuracy and efficiency.
This approach improves compression efficiency by deriving a final prediction block through weighted summation of multiple prediction blocks, enhancing prediction performance and overall image encoding/decoding efficiency.
Smart Images

Figure 2026514902000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image encoding / decoding method and apparatus, and a recording medium storing a bitstream. [Background technology]
[0002] In recent years, the demand for high-resolution, high-quality images such as HD (High Definition) and UHD (Ultra High Definition) images has increased in various application fields, leading to discussions about highly efficient image compression technologies.
[0003] Various image compression techniques exist, such as inter-prediction techniques that predict pixel values contained in the current picture from previous or subsequent pictures; intra-prediction techniques that predict pixel values contained in the current picture using pixel information within the current picture; and entropy coding techniques that assign short codes to frequently occurring values and long codes to less frequently occurring values. By using such image compression techniques, image data can be effectively compressed and transmitted or stored. [Overview of the Initiative] [Problems that the invention aims to solve]
[0004] This disclosure aims to provide a method and apparatus for inducing multiple predicted blocks relative to current blocks.
[0005] This disclosure aims to provide a method and apparatus for inducing multiple prediction blocks using a prediction mode. [Means for solving the problem]
[0006] The image decoding method and apparatus according to this disclosure can induce a plurality of prediction blocks, including a first prediction block and a second prediction block, from the current block, induce a final prediction block of the current block based on the first prediction block and the second prediction block, and restore the current block based on the final prediction block of the current block.
[0007] In the image decoding method and apparatus according to this disclosure, the first prediction block may be guided based on at least one of a first L0 motion vector predictor or a second L1 motion vector predictor, and the second prediction block may be guided based on at least one of a first L1 motion vector predictor or a second L0 motion vector predictor. Here, at least one of the first L0 motion vector predictor or the first L1 motion vector predictor may be guided based on a first prediction mode, and at least one of the second L0 motion vector predictor or the second L1 motion vector predictor may be guided based on a second prediction mode.
[0008] In the image decoding method and apparatus relating to this disclosure, the first prediction mode may be AMVP mode, and the second prediction mode may be merge mode.
[0009] In the image decoding method and apparatus relating to this disclosure, the first L0 motion vector predictor may be guided based on the first prediction mode, and the first L1 motion vector predictor may be guided based on the first L0 motion vector predictor which has been guided in advance.
[0010] In the image decoding method and apparatus relating to this disclosure, at least one of the first L0 motion vector predictor or the second L1 motion vector predictor may be corrected based on either a bidirectional matching-based correction method or a template matching-based correction method.
[0011] In the image decoding method and apparatus relating to this disclosure, one of the two-way matching-based correction method or the template matching-based correction method may be selected based on a predefined first condition.
[0012] In the image decoding method and apparatus relating to this disclosure, at least one of the first L1 motion vector predictor or the second L0 motion vector predictor may be corrected based on either a bidirectional matching-based correction method or a template matching-based correction method.
[0013] In the image decoding method and apparatus relating to this disclosure, one of the two-way matching-based correction method or the template matching-based correction method may be selected based on a predefined second condition.
[0014] In the image decoding method and apparatus relating to this disclosure, whether or not the second condition is met may be determined based on whether or not the first condition is met.
[0015] In the image decoding method and apparatus relating to this disclosure, the first prediction block may be guided based on an intermode, and the second prediction block may be guided based on one or more intra-prediction modes for the current block. Here, the one or more intra-prediction modes for the current block may include at least one of a planar mode, MPM, MIP mode, DIMD-based intra-prediction mode, or TIMD-based intra-prediction mode.
[0016] In the image decoding method and apparatus relating to this disclosure, the first prediction block is guided based on the inter-prediction block and intra-prediction block of the current block, and the second prediction block may be guided based on a further intra-prediction mode for the current block.
[0017] In the image decoding method and apparatus according to the present disclosure, the first prediction block is derived based on one or more block vectors derived from the IBC candidate list of the current block, and the second prediction block may be derived based on one or more intra prediction modes for the current block.
[0018] The image encoding method and apparatus according to the present disclosure derive a plurality of prediction blocks including a first prediction block and a second prediction block for a current block, derive a final prediction block of the current block based on the first prediction block and the second prediction block, derive a residual block of the current block based on the final prediction block of the current block, and can encode the residual block of the current block.
[0019] There is provided a computer-readable digital storage medium storing encoded video / image information for performing an image decoding method by a decoding apparatus according to the present disclosure.
[0020] There is provided a computer-readable digital storage medium storing video / image information generated by an image encoding method according to the present disclosure.
[0021] There are provided a method and an apparatus for transmitting video / image information generated by an image encoding method according to the present disclosure.
Advantages of the Invention
[0022] According to the present disclosure, by deriving multiple prediction blocks for a current block and deriving a final prediction block by weighted summation thereof, the compression efficiency can be improved.
[0023] According to the present disclosure, by proposing various methods for deriving multiple prediction blocks according to the prediction mode of the current block, the prediction performance and the compression efficiency can be improved.
Brief Description of the Drawings
[0024] [Figure 1] A diagram showing a video / image coding system according to the present disclosure. [Figure 2] A schematic block diagram of an encoding device to which an embodiment of the present disclosure is applicable and in which video / image signals are encoded. [Figure 3] A schematic block diagram of a decoding device to which an embodiment of the present disclosure is applicable and in which video / image signals are decoded. [Figure 4] An example of an embodiment according to the present disclosure, showing an inter prediction method performed by a decoding device 300. [Figure 5] A diagram showing a schematic configuration of an inter prediction unit 332 that performs an inter prediction method according to the present disclosure. [Figure 6] An example of an embodiment according to the present disclosure, showing an inter prediction method performed by an encoding device (200). [Figure 7] A diagram showing a schematic configuration of an inter prediction unit (221) that performs an inter prediction method according to the present disclosure. [Figure 8] A diagram showing an example of a content streaming system to which an embodiment of the present disclosure is applicable.
Embodiments for Carrying Out the Invention
[0025] The present disclosure can be modified in various ways and can have various embodiments. Specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood to include all modifications, equivalents, or alternatives included in the spirit and technical scope of the present disclosure. In the description of each figure, similar reference numerals are used for similar components.
[0026] Terms such as "First," "Second," etc., may be used to describe various components, but these components should not be limited by such terms. These terms are used solely for the purpose of distinguishing one component from another. For example, without exceeding the scope of rights of this disclosure, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of multiple related descriptions or any one of multiple related descriptions.
[0027] When it is stated that one component is "connected" or "linked" to another component, it should be understood that it may be directly connected or linked to the other component, and that there may be other components in between. On the other hand, when it is stated that one component is "directly connected" or "linked" to another component, it should be understood that there are no other components in between.
[0028] The terminology used in this application is solely for the purpose of describing specific embodiments and is not intended to limit the disclosure. Singular expressions include plural expressions unless otherwise specified in the context. In this application, terms such as “includes” or “having” are intended to specify the existence of features, figures, stages, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preemptively exclude the possibility of the existence or addition of one or more other features, figures, stages, operations, components, parts, or combinations thereof.
[0029] This disclosure relates to video / image coding. For example, the methods / examples disclosed herein may be applied to methods disclosed in the VVC (versatile video coding) standard. Also, the methods / examples disclosed herein may be applied to methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (e.g., H.267 or H.268).
[0030] This specification presents various embodiments relating to video / image coding, and unless otherwise specified, the above embodiments may be combined with each other.
[0031] In this specification, video can mean a collection of images over time. Picture generally means a unit representing a single image at a specific time point, and slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may contain one or more CTUs (coding tree units). A picture may consist of one or more slices / tiles. A tile is a rectangular area consisting of multiple CTUs in a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of CTUs having the same height as the picture height and the width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs having the same height as the picture parameter set and the width as the picture width. CTUs within a tile may be arranged consecutively by a CTU raster scan, while tiles within a picture may be arranged consecutively by a tile raster scan. A slice may contain an integer number of complete tiles or an integer number of consecutive complete CTU rows within a picture tile that can exclusively be contained within a single NAL unit. On the other hand, a picture may be divided into two or more subpictures. A subpicture may be a rectangular region of one or more slices in the picture.
[0032] A pixel, or pel, can refer to the smallest unit that makes up a picture (or image). The term "sample" may also be used as a counterpart to "pixel." A sample can generally represent a pixel or a pixel value, and may represent only the pixel / pixel value for the lumen component, or only the pixel / pixel value for the chroma component.
[0033] A unit can refer to a basic unit of image processing. A unit may contain at least one of a specific region of a picture and information associated with that region. A unit may contain one luma block and two chroma (e.g., cb, cr) blocks. The term unit may, as in some cases, be used interchangeably with terms such as block or area. Generally, an MxN block may contain a sample (or sample array) or a set (or array) of transform coefficients consisting of M columns and N rows.
[0034] In this specification, "A or B" can mean "A only," "B only," or "both A and B." In other words, in this specification, "A or B" may be interpreted as "A and / or B." For example, in this specification, "A, B or C" can mean "A only," "B only," "C only," or "any combination of A, B and C."
[0035] In this specification, a slash ( / ) or comma can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B or C".
[0036] In this specification, "at least one of A and B" can mean "A only," "B only," or "both A and B." Furthermore, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted as equivalent to "at least one of A and B."
[0037] Furthermore, in this specification, "at least one of A, B and C" can mean "A only," "B only," "C only," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" can mean "at least one of A, B and C."
[0038] Furthermore, parentheses used in this specification can mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra prediction," and "intra prediction" may be proposed as an example of "prediction." Also, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction."
[0039] In this specification, technical features described individually in the same drawings may be embodied individually or simultaneously.
[0040] Figure 1 shows the video / image coding system related to this disclosure.
[0041] Referring to Figure 1, the video / image coding system may include a first device (source device) and a second device (receiving device).
[0042] A source device can transmit encoded video / image information or data to a receiving device in the form of a file or streaming via a digital storage medium or network. The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may also be called a video / image encoding device, and the decoding device may also be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may consist of a separate device or external component.
[0043] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include a computer, tablet, and smartphone, and can generate video / images (electronically). For example, virtual video / images may be generated through a computer, in which case the video / image capture process may be replaced by the process of generating the associated data.
[0044] An encoding device can encode input video / images. The encoding device can perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) may be output in the form of a bitstream.
[0045] The transmission unit can transmit encoded video / image information or data output in bitstream format to the receiving unit of a receiving device via a digital storage medium or network in file or streaming format. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmission unit may include elements for generating media files in a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0046] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of an encoding device.
[0047] The renderer can render the decoded video / image. The rendered video / image may be displayed on the display unit.
[0048] Figure 2 is a schematic block diagram of an encoding device to which the embodiments of this disclosure can be applied, in which video / image signals are encoded.
[0049] Referring to Figure 2, the encoding device 200 may include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-prediction unit (221) and an intra-prediction unit (222). The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit 250 may be called a reconstructor or a reconstructed block generator. The image segmentation unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 described above may be composed of one or more hardware components (e.g., an encoding device chipset or processor) depending on the embodiment. The memory 270 may also include a DPB (decoded picture buffer) and may be composed of a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0050] The image splitting unit 210 can split an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively split from a coding tree unit (CTU) or the largest coding unit (LCU) using a QTBTTT (Quad-tree binary-tree ternary-tree) structure.
[0051] As an example, a single coding unit may be divided into multiple coding units having deeper depths based on a quad-tree structure, a binary tree structure, and / or a tertiary structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary tree structure and / or tertiary structure. Alternatively, the binary tree structure may be applied before the quad-tree structure. The coding procedure according to this specification may be performed based on a final coding unit that is not further divided. In this case, based on coding efficiency due to image characteristics, the largest coding unit may be immediately used as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units of lower depths, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and reconstruction, which will be described later.
[0052] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be separated or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit that derives a transform coefficient and / or a unit that derives a residual signal from the transform coefficient.
[0053] The term "unit" may be used interchangeably with terms such as "block" or "area." Generally, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the lumen component, or only the pixel / pixel value of the chroma component. The term "sample" may be used in conjunction with a single picture (or image), pixel, or pel.
[0054] The encoding device 200 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from the inter-prediction unit 221 or intra-prediction unit 222 from the input image signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within the encoding device 200 may be called the subtraction unit 231.
[0055] The prediction unit 220 can make predictions for the block to be processed (hereinafter referred to as the current block) and generate a predicted block containing prediction samples for the current block. The prediction unit 220 can determine whether intra-prediction or inter-prediction is applied to the current block or on a CU basis. As will be described later in the explanation of each prediction mode, the prediction unit 220 can generate various information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240. The information related to prediction may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0056] The intra-prediction unit 222 can predict the current block by referring to a sample in the current picture. The referenced sample may be located in the vicinity (neighbor) of the current block, or it may be located at a certain distance from the current block, depending on the prediction mode. In intra-prediction, the prediction mode may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one of DC mode or planar mode. The directional mode may include 33 or 65 directional modes, depending on the degree of fineness of the prediction direction. However, this is an example, and more or fewer directional modes may be used depending on the settings. The intra-prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction modes applied to the surrounding blocks.
[0057] The interprediction unit 221 can guide a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the surrounding block and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, the surrounding block may include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, colCU, etc., and the reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the interpretation unit 221 can construct a motion information candidate list based on surrounding blocks and generate information indicating which candidates are used to derive the motion vector and / or reference picture index of the current block. Interpretation may be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 221 can use the motion information of surrounding blocks as the motion information of the current block. In skip mode, unlike merge mode, the residual signal does not need to be transmitted.In motion vector prediction (MVP) mode, the motion vectors of surrounding blocks are used as motion vector predictors, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0058] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for predictions for a single block, and can also apply intra-prediction and inter-prediction simultaneously. This may be called CIIP (combined inter and intra prediction) mode. The prediction unit may also be based on intra-block copy (IBC) prediction mode or palette mode for predictions for blocks. The IBC prediction mode or palette mode may be used for coding content images / videos such as games, as in SCC (screen content coding). IBC basically performs predictions within the current picture, but can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can use at least one of the inter-prediction methods described herein. Palette mode may be considered an example of intra-coding or intra-prediction. When palette mode is applied, in-picture sample values can be signaled based on information about the palette table and palette index. The prediction signal generated by the prediction unit 220 may be used to generate a restoration signal or to generate a residual signal.
[0059] The transformation unit 232 can generate transformation coefficients by applying a transformation method to the residual signal. For example, the transformation method may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to the transformation obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to the transformation obtained by generating a prediction signal using all previously restored pixels and obtaining a transformation based on it. Furthermore, the transformation process may be applied to pixel blocks of the same size that are square, or to blocks of variable size that are not square.
[0060] The quantization unit 233 quantizes the conversion coefficients and transmits them to the entropy encoding unit 240, which can encode the quantized signal (information regarding the quantized conversion coefficients) and output it as a bitstream. The information regarding the quantized conversion coefficients may be called residual information. The quantization unit 233 can rearrange the block-shaped quantized conversion coefficients into a one-dimensional vector based on the coefficient scan order, and can generate information regarding the quantized conversion coefficients based on the one-dimensional vector-shaped quantized conversion coefficients.
[0061] The entropy encoding unit 240 can perform various encoding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy encoding unit 240 can also encode information necessary for video / image restoration (e.g., the values of syntax elements) together with or separately from the quantized conversion coefficients.
[0062] The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as an adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. In this specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded by the encoding procedure described above and included in the bitstream. The bitstream may be transmitted over a network or stored on a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 may be transmitted by a transmission unit (not shown) and / or stored by a storage unit (not shown) which are configured as internal / external elements of the encoding device 200, or the transmission unit may be included in the entropy encoding unit 240.
[0063] The quantized conversion coefficients output from the quantization unit 233 may be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized conversion coefficients in the inverse quantization unit 234 and the inverse transformation unit 235, the residual signal (residual block or residual sample) can be reconstructed. The adder unit 250 may add the reconstructed residual signal to the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block may be used as the reconstructed block. The adder unit 250 may be called the reconstruction unit or the reconstructed block generation unit. The generated reconstructed signal may be used for intra-prediction of the next block to be processed in the current picture, and may be used for inter-prediction of the next picture after filtering, as described later. On the other hand, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or reconstruction process.
[0064] The filtering unit 260 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 260 can generate various filtering-related information and transmit it to the entropy encoding unit 240. The filtering-related information may be encoded by the entropy encoding unit 240 and output in bitstream format.
[0065] The corrected restored picture transmitted to memory 270 may be used as a reference picture in the interpretation unit 221. This allows the encoding device to avoid prediction mismatches between the encoding device 200 and the decoding device when interpretation is applied, and also improves encoding efficiency.
[0066] The DPB in memory 270 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 221. Memory 270 can store motion information of blocks from which motion information in the current picture has been derived (or encoded), and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 221 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 270 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 222.
[0067] Figure 3 is a schematic block diagram of a decoding device to which the embodiments of this disclosure can be applied, in which video / image signals are decoded.
[0068] Referring to Figure 3, the decoding device 300 may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-prediction unit (332) and an intra-prediction unit (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321).
[0069] The entropy decoding unit 310, the residual processing unit 320, the prediction unit 330, the addition unit 340, and the filtering unit 350 described above may be configured by a single hardware component (e.g., a decoding device chipset or processor) depending on the embodiment. The memory 360 may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.
[0070] When a bitstream containing video / image information is input, the decoding device 300 can reconstruct the image in accordance with the process by which the video / image information was processed in the encoding device shown in Figure 2. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the decoding processing units may be coding units, which may be divided from a coding tree unit or a maximum coding unit according to a quad-tree structure, a binary tree structure, and / or a tertiary tree structure. One or more conversion units may be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 may then be reproduced by a playback device.
[0071] The decoding device 300 can receive the signal output from the encoding device shown in Figure 2 in the form of a bitstream, and the received signal may be decoded by the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream and derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can decode the picture based on the parameter set information and / or the general constraint information. The signal / received information and / or syntax elements described later in this specification may be decoded by the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements necessary for image reconstruction and the quantized values of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded and the decoding information of the surrounding and decoded blocks or symbol / bin information decoded in a previous stage, predicts the probability of bin occurrence based on the determined context model, and generates symbols corresponding to the values of each syntax element by performing arithmetic decoding of the bins. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the symbol / bin information decoded for the context model of the next symbol / bin.Information related to prediction from the information decoded by the entropy decoding unit 310 is provided to the prediction unit (inter-prediction unit 332 and intra-prediction unit 331), and residual values that have been entropy decoded by the entropy decoding unit 310, i.e., quantized conversion coefficients and related parameter information, may be input to the residual processing unit 320. The residual processing unit 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, information related to filtering from the information decoded by the entropy decoding unit 310 may be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives signals output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310.
[0072] On the other hand, the decoding device according to this specification may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoding device (video / image / picture information decoding device) and a sample decoding device (video / image / picture sample decoding device). The information decoding device may include the entropy decoding unit 310, and the sample decoding device may include at least one of the inverse quantization unit 321, inverse transformation unit 322, addition unit 340, filtering unit 350, memory 360, inter-prediction unit 332, and intra-prediction unit 331.
[0073] The inverse quantization unit 321 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 321 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients.
[0074] The inverse conversion unit 322 performs an inverse conversion on the conversion coefficients to obtain a resistive signal (residual block, resistive sample array).
[0075] The prediction unit 320 can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 310, the prediction unit 320 can determine whether intra-prediction or inter-prediction is applied to the current block and determine a specific intra / inter-prediction mode.
[0076] The prediction unit 320 can generate prediction signals based on various prediction methods described later. For example, the prediction unit 320 can apply intra-prediction or inter-prediction for prediction of a single block, or it can apply intra-prediction and inter-prediction simultaneously. This may be called CIIP (combined inter and intra prediction) mode. The prediction unit may also be based on intra-block copy (IBC) prediction mode or palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content image / video coding such as SCC (screen content coding) for games. IBC basically performs prediction within the current picture, but can be performed similarly to inter-prediction in that it derives a reference block within the current picture. That is, IBC can use at least one of the inter-prediction methods described herein. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, information about the palette table and palette index may be included in the video / image information and signaled.
[0077] The intra-prediction unit 331 can predict the current block by referring to a sample in the current picture. The referenced sample may be located in the vicinity (neighbor) of the current block, or it may be located at a certain distance from the current block, depending on the prediction mode. In intra-prediction, the prediction mode may include one or more non-directional modes and multiple directional modes. The intra-prediction unit 331 can determine the prediction mode to be applied to the current block using the prediction modes applied to the surrounding blocks.
[0078] The interprediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between surrounding blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the interprediction unit 332 can construct a motion information candidate list based on surrounding blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction may be performed based on various prediction modes, and the prediction information may include information indicating the interprediction mode for the current block.
[0079] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-prediction unit 332 and / or intra-prediction unit 331). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block may be used as the restored block.
[0080] The summing unit 340 may be called the restoration unit or the restoration block generation unit. The generated restoration signal may be used for intra-prediction of the next block to be processed in the current picture, and may be output after filtering as described later, or may be used for intra-prediction of the next picture. On the other hand, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0081] The filtering unit 350 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and transmit the modified restored picture to the memory 360, specifically to the DPB of the memory 360. The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like.
[0082] The restored picture stored (modified) in the DPB of memory 360 may be used as a reference picture in the inter-prediction unit 332. Memory 360 can store motion information of blocks from which motion information in the current picture has been derived (or decoded), and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 332 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 360 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 331.
[0083] In this specification, the embodiments described for the filtering unit 260, the inter-prediction unit 221, and the intra-prediction unit 222 of the encoding device 200 may also be applied identically or in a corresponding manner to the filtering unit 350, the inter-prediction unit 332, and the intra-prediction unit 331 of the decoding device 300, respectively.
[0084] Figure 4 is a diagram illustrating an embodiment of the present disclosure, showing an inter prediction method performed by a decoding device 300.
[0085] Referring to Figure 4, multiple predicted blocks can be induced for the current block (S400).
[0086] Multiple prediction blocks for the current block may include a first prediction block and a second prediction block. The first prediction block may be induced based on a first prediction mode, and the second prediction block may be induced based on a second prediction mode. Here, the first prediction mode and the second prediction mode may be the same prediction mode or may be different prediction modes. The method for inducing the first and second prediction blocks will be described below.
[0087] Example 1
[0088] This disclosure relates to the case where the first and second prediction modes are AMVP-MERGE modes, and more particularly to the case where motion information for unidirectional prediction is acquired for each of the AMVP mode and the merge mode.
[0089] In AMVP-MERGE mode, motion information in both directions may be acquired for the current block. Here, motion information in both directions may include motion information in a first predicted direction acquired based on AMVP mode and motion information in a second predicted direction acquired based on merge mode. The first predicted direction may be expressed as the LX direction, and the second predicted direction may be expressed as the L(1-X) direction. X may have a value of 0 or 1. The same meaning may be interpreted in the embodiments described later.
[0090] Specifically, for a first prediction direction, at least one of the following may be signaled: an MVP index, a reference picture index, or MVD information. The MVP index can identify one of several MVP candidates belonging to the MVP candidate list. The reference picture index can identify one of the reference pictures belonging to the reference picture list for the first prediction direction. MVD information can represent information about the motion vector difference. Based on the MVP candidate identified by the MVP index, a motion vector predictor (MVP[0][X]) for the first prediction direction of the current block may be induced. In MVP[0][X], [0] represents the AMVP mode and [X] represents the LX direction. That is, MVP[0][X] may represent a motion vector predictor for the LX direction induced based on the AMVP mode. MVP[0][X] can also represent the motion vector of the MVP candidate identified by the MVP index, or it can represent the sum of the motion vector of the MVP candidate identified by the MVP index and the motion vector difference induced based on the MVD information.
[0091] A merge index may be signaled for the second prediction direction. The merge index can identify one of several merge candidates belonging to the merge candidate list. Based on the motion vector for the second prediction direction of the merge candidate identified by the merge index, a motion vector predictor (MVP[1][1-X]) for the second prediction direction of the current block may be induced. In MVP[1][1-X], [1] represents the merge mode and [1-X] represents the L(1-X) direction. That is, MVP[1][1-X] may represent a motion vector predictor for the L(1-X) direction induced based on the merge mode.
[0092] Based on the motion vector predictor (MVP[0][X]) for the first prediction direction, the first predicted block of the current block may be induced. Based on the motion vector predictor (MVP[1][1-X]) for the second prediction direction, the second predicted block of the current block may be induced.
[0093] At least one of the motion vector predictors (MVP[0][X], MVP[1][1-X]) for bidirectional prediction of the current block may be corrected based on either a bidirectional matching-based correction method or a template matching-based correction method, as described later. Based on the corrected motion vector predictor, the first and second predicted blocks of the current block may be induced, respectively.
[0094] 1. Bidirectional matching-based correction method
[0095] Within a predetermined search range, a cost array can be calculated by performing a search based on the position of the current block's reference block.
[0096] The reference block of the current block may include a reference block for the first prediction direction (hereinafter referred to as the first reference block) and a reference block for the second prediction direction (hereinafter referred to as the second reference block). The first reference block may be identified based on motion information for the first prediction direction. The second reference block may be identified by motion information for the second prediction direction. The motion information for the first prediction direction may be acquired based on either AMVP mode or merge mode, and the motion information for the second prediction direction may be acquired based on the other mode of AMVP mode or merge mode.
[0097] The cost array may consist of multiple costs calculated for each search position within the search range. Each cost may be calculated as the sample difference between at least two blocks searched in both directions. For example, the cost may be calculated as the SAD (sum of absolute difference) between at least two blocks searched in both directions. In this case, the block searched in the first prediction direction will be called the LX block, and the block searched in the second prediction direction will be called the L(1-X) block. The cost may be calculated based on all samples belonging to the LX and L(1-X) blocks, or it may be calculated based on some samples within the LX and L(1-X) blocks.
[0098] Here, "some samples" means subblocks of LX and L(1-X) blocks, and at least one of the width or height of the subblock may be half the width or height of the LX and L(1-X) blocks. That is, LX and L(1-X) blocks have a size of WxH, and the "some samples" may be subblocks with a size of WxH / 2, W / 2xH, or W / 2xH / 2. In this case, if the "some samples" is a WxH / 2 subblock, the "some samples" may be the upper subblock (or lower subblock) within the LX and L(1-X) blocks. If the "some samples" is a W / 2xH subblock, the "some samples" may be the left subblock (or right subblock) within the LX and L(1-X) blocks. If the "some samples" is a W / 2xH / 2 subblock, the "some samples" may be the upper left subblock within the LX and L(1-X) blocks, but are not limited to these.
[0099] Alternatively, some samples may be defined as at least one even-numbered sample line or at least one odd-numbered sample line in the LX and L(1-X) blocks. In this case, the sample lines may be vertical or horizontal.
[0100] Alternatively, some samples may refer to samples at positions identically defined in the encoding / decoding device. For example, the samples at the predefined positions may refer to at least one of the following: the top-left sample, top-right sample, bottom-left sample, bottom-right sample, center sample, center sample of a sample column / row adjacent to the boundary of the current block, or a sample located on a diagonal line within the current block.
[0101] The search positions within the search range may be shifted by p units in the x-axis direction and q units in the y-axis direction from the position of the current block's reference block. For example, if p and q are integers in the range of -1 to 1, the maximum number of integer-per unit search positions within the search range may be 9. Or, if p and q are integers in the range of -2 to 2, the maximum number of integer-per unit search positions within the search range may be 25. However, this is not limited to these cases; p and q may also belong to a range of integers with a size (or absolute value) greater than 2, and searches may be performed in decimal-per unit increments.
[0102] The search position within the search range may be determined based on an offset that is identically predefined for both the encoding and decoding devices. That is, the offset may be defined as the displacement vector between the current block's reference block position and the search position. The offset may include at least one of a non-directional offset or a directional offset. The directional offset may include an offset for at least one of the following directions: left, right, top, bottom, upper left, upper right, lower left, or lower right. The non-directional offset means an offset with a size of 0, and the directional offset may mean an offset where the size (or absolute value) of at least one of the x or y components of the offset is greater than or equal to 1.
[0103] As an example, the offset may be defined as shown in Table 1 below.
[0104] [Table 1]
[0105] Table 1 defines offsets that specify the search position by index, where dX[i] represents the x-component of the i-th offset and dY[i] represents the y-component of the i-th offset. The offsets in Table 1 may include the non-directional offset (0,0) and eight directional offsets. However, the index in Table 1 is merely for distinguishing offsets and does not limit the position of the offset corresponding to the index or limit the priority between offsets. Also, Table 1 shows the case where the size of the x and y components of the offset is 1, but this is just an example, and offsets may be defined where the size of at least one of the x or y components is greater than or equal to 2.
[0106] The aforementioned offsets may be defined for the L0 and L1 directions, respectively. The offset for searching in the L1 direction may be determined dependently on the offset for searching in the L0 direction. For example, if the offset for searching in the L0 direction is (p,q), the offset for searching in the L1 direction may be set to (-p,-q) by mirroring. Alternatively, the offset for searching in the L1 direction may be determined independently of the offset for searching in the L0 direction.
[0107] The information regarding the size and / or direction of the offset described above may be predefined identically for both the encoding and decoding devices, or it may be encoded in the encoding device and signaled to the decoding device. The information may be variably determined taking into account the block attributes described above.
[0108] The cost corresponding to each search position (or predefined offset) within the search range may be calculated using the method described above. Among the multiple costs belonging to the cost array, the cost with the minimum value may be identified, and the delta motion vector may be determined based on the offset corresponding to the identified cost. The motion vector predictor in the L0 direction may be corrected based on the delta motion vector (deltaMV), and the motion vector predictor in the L1 direction may be corrected based on the mirrored delta motion vector (-deltaMV).
[0109] 2. Template matching-based correction method
[0110] A cost array can be calculated by performing a search within a predetermined search range, based on the position of the reference block of the current block. Here, the reference block of the current block is as described in the "Bidirectional Matching-Based Correction Method" above. The cost array may consist of multiple costs calculated for each search position within the search range. Each cost may be calculated as the sample difference between the template region of the current block and the template region of the search position. As an example, the cost may be calculated as the sum of absolute difference (SAD) between the template region of the current block and the template region of the search position. The template region of the search position can mean the template region of the block that has the sample at the search position as the upper left corner sample.
[0111] The aforementioned cost may be calculated based on all samples belonging to the template region of the current block and search position, or it may be calculated based on some of the samples within the template region.
[0112] Here, some samples can mean at least one even-numbered sample line or at least one odd-numbered sample line within the template region. In this case, the sample lines may be vertical or horizontal.
[0113] Alternatively, the template region may include at least one of the following regions adjacent to a block (i.e., the current block, the block at the search position): the top region, the left region, the top-left region, the bottom-left region, or the top-right region. In this case, some samples may be restricted to samples belonging to a region at a specific position within the template region. For example, the region at a specific position may include at least one of the top region or the left region.
[0114] Alternatively, the template region may include adjacent sample lines and / or at least one non-adjacent sample line adjacent to the block. In this case, some samples may be restricted to samples belonging to a sample line at a specific location within the template region. For example, the sample line at a specific location may be predefined identically for both the encoding and decoding devices and may include at least one of the following: an adjacent sample line, a non-adjacent sample line one sample away from the block boundary, or a non-adjacent sample line two samples away from the block boundary. Information indicating the sample line at the specific location may be signaled as a bitstream.
[0115] The search position within the search range may be a position shifted by p units in the x-axis direction and q units in the y-axis direction from the position of the current block's reference block, and may be determined based on an offset that is identically defined for both the encoding and decoding devices. This is explained in the "Bidirectional Matching-Based Correction Method" above, and a redundant explanation is omitted here.
[0116] Cost arrays may be calculated for both the first and second prediction directions using the method described above. Among the multiple costs belonging to the cost array of the first prediction direction, the cost with the minimum value may be identified, and the delta motion vector for the first prediction direction may be determined based on the offset corresponding to the identified cost. Based on the determined delta motion vector for the first prediction direction, the motion vector predictor for the first prediction direction may be corrected. Similarly, among the multiple costs belonging to the cost array of the second prediction direction, the cost with the minimum value may be identified, and the delta motion vector for the second prediction direction may be determined based on the offset corresponding to the identified cost. Based on the determined delta motion vector for the second prediction direction, the motion vector predictor for the second prediction direction may be corrected.
[0117] Depending on whether the current block satisfies a predetermined first condition, either the bidirectional matching-based correction method or the template matching-based correction method described above may be selectively used. The first condition relating to this disclosure may include at least one of the following: the current block performs bidirectional prediction; bidirectional reference pictures for the current block exist before and / or after the current picture in time order; or the POC differences between the current picture and each reference picture are the same. The time order can mean picture order count (POC) or coding order.
[0118] If the first condition described above is true, a bidirectional matching-based correction method may be used, and if the first condition is false, a template matching-based correction method may be used. Alternatively, if all of the first conditions are true, a bidirectional matching-based correction method may be used, and if any one of the first conditions is false, a template matching-based correction method may be used. Alternatively, if any one of the first conditions is true, a bidirectional matching-based correction method may be used, and if all of the first conditions are false, a template matching-based correction method may be used.
[0119] For example, if the current block performs bidirectional prediction, a bidirectional matching-based correction method may be used; otherwise, a template matching-based correction method may be used.
[0120] Alternatively, if the POC difference (diffPOC0) between the current picture including the current block and the reference picture in the first prediction direction is the same as the POC difference (diffPOC1) between the current picture and the reference picture in the second prediction direction, a bidirectional matching-based correction method may be used; otherwise, a template matching-based correction method may be used.
[0121] Alternatively, if the current block performs bidirectional prediction and diffPOC is the same as diffPOC1, a bidirectional matching-based correction method may be used; otherwise (i.e., the current block does not perform bidirectional prediction, or diffPOC is not the same as diffPOC1), a template matching-based correction method may be used.
[0122] Alternatively, if the current block performs bidirectional prediction, a bidirectional matching-based correction method may be used even if diffPOC is not identical to diffPOC1. On the other hand, if the current block does not perform bidirectional prediction, a template matching-based correction method may be used.
[0123] The AMVP-MERGE mode described herein may also be applied when the current block is a block encoded in IBC (intra block copy) mode. When the interprediction mode of the current block is IBC mode, the current picture to which the current block belongs is used as the reference picture, and otherwise the method according to Embodiment 1 described above may be applied identically.
[0124] Example 2
[0125] This disclosure relates to the case where the first and second prediction modes are AMVP-MERGE modes, and more particularly to the case where motion information for bidirectional prediction is acquired for each of the AMVP and Merge modes.
[0126] In AMVP-MERGE mode, two or more bidirectional motion information entries may be acquired for the current block. For example, bidirectional motion information may be acquired based on AMVP mode, and bidirectional motion information may be acquired based on merge mode. However, it is not limited to this; unidirectional motion information may be acquired based on AMVP mode, and bidirectional motion information may be acquired based on merge mode. Alternatively, bidirectional motion information may be acquired based on AMVP mode, and unidirectional motion information may be acquired based on merge mode. For the sake of explanation, we will assume below that bidirectional motion information is acquired for both AMVP mode and merge mode.
[0127] Currently, motion information may be acquired for each prediction direction of the block based on AMVP mode. For this purpose, at least one of the following may be signaled: information on whether or not to perform bidirectional prediction, the MVP index, the reference picture index, or MVD information. If the block currently performs bidirectional prediction, the MVP index, the reference picture index, and MVD information may be signaled for each prediction direction.
[0128] Specifically, an MVP candidate list may be constructed for each prediction direction in the current block. The MVP candidate list may contain multiple MVP candidates. A motion vector predictor (MVP[0][X], MVP[0][1-X]) for each prediction direction may be derived from the MVP candidate list for that prediction direction. In MVP[0][X], [0] represents the AMVP mode and [X] represents the LX direction. That is, MVP[0][X] may represent a motion vector predictor for the LX direction derived based on the AMVP mode. MVP[0][X] may be derived based on an MVP candidate identified by the MVP index for the LX direction. For example, MVP[0][X] can mean the motion vector of an MVP candidate identified by the MVP index for the LX direction. Alternatively, MVP[0][X] can mean a motion vector derived based on the motion vector difference between the motion vector of an MVP candidate identified by the MVP index for the LX direction and the motion vector difference. Here, the motion vector difference may be derived based on MVD information for the LX direction. Similarly, in MVP[0][1-X], [0] represents the AMVP mode and [1-X] represents the L(1-X) direction. That is, MVP[0][1-X] may represent a motion vector predictor in the L(1-X) direction induced based on the AMVP mode. MVP[0][1-X] may be induced based on MVP candidates identified by the MVP index in the L(1-X) direction. For example, MVP[0][1-X] can mean the motion vector of an MVP candidate identified by the MVP index in the L(1-X) direction. Alternatively, MVP[0][1-X] can mean a motion vector induced based on the motion vector of an MVP candidate identified by the MVP index in the L(1-X) direction and the motion vector difference. Here, the motion vector difference may be induced based on MVD information in the L(1-X) direction.
[0129] On the other hand, motion information of the current block may be obtained based on the merge mode. For this purpose, at least one of the merge index or MVD information that identifies one of several merge candidates belonging to the merge candidate list may be signaled. Motion information of the current block may be obtained based on the motion information of the merge candidate identified by the merge index. Here, motion information may include at least one of the motion vector (or motion vector predictor) or reference picture index. If the identified merge candidate performs bidirectional prediction, bidirectional motion information may be obtained for the current block. However, if the identified merge candidate performs unidirectional prediction, unidirectional motion information may be obtained for the current block.
[0130] Specifically, a merge candidate list containing multiple merge candidates may be constructed for the current block. Motion vector predictors in both directions (MVP[1][X], MVP[1][1-X]) may be derived from the merge candidate list. In MVP[1][X], [1] represents the merge mode and [X] represents the LX direction. That is, MVP[1][X] may represent the LX direction motion vector predictor derived based on the merge mode. MVP[1][X] may be derived based on the merge candidate identified by the merge index. For example, MVP[1][X] can mean the LX direction motion vector of the merge candidate identified by the merge index. Alternatively, MVP[1][X] can mean the motion vector derived based on the LX direction motion vector of the merge candidate identified by the merge index and the motion vector difference. Here, the motion vector difference may be derived based on MVD information. Similarly, in MVP[1][1-X], [1] represents the merge mode and [1-X] represents the L(1-X) direction. In other words, MVP[1][1-X] may represent a motion vector predictor in the L(1-X) direction induced based on the merge mode. MVP[1][1-X] may be induced based on merge candidates identified by the merge index. For example, MVP[1][1-X] can mean the motion vector in the L(1-X) direction of a merge candidate identified by the merge index. Alternatively, MVP[1][1-X] can mean a motion vector induced based on the motion vector difference between the L(1-X) direction motion vector of a merge candidate identified by the merge index. Here, the motion vector difference may be induced based on MVD information.
[0131] Bidirectional motion vector predictors (MVP[0][X],MVP[0][1-X]) induced based on AMVP mode and bidirectional motion vector predictors (MVP[1][X],MVP[1][1-X]) induced based on merge mode can be paired in a predefined manner. This pairing may induce one or more bidirectional motion vector predictors for the current block.
[0132] As an example, the pairing described above may induce a single bidirectional motion vector predictor for the current block. In this case, the bidirectional motion vector predictor for the current block may be {MVP[0][X],MVP[1][1-X]}. That is, the bidirectional motion vector predictor for the current block can be induced to be a combination of a motion vector predictor in the LX direction using AMVP mode and a motion vector predictor in the L(1-X) direction using merge mode. Alternatively, the bidirectional motion vector predictor for the current block may be {MVP[0][1-X],MVP[1][X]}. That is, the bidirectional motion vector predictor for the current block may be induced to be a combination of a motion vector predictor in the L(1-X) direction using AMVP mode and a motion vector predictor in the LX direction using merge mode.
[0133] Alternatively, the pairing may induce two bidirectional motion vector predictors for the current block. In this case, the bidirectional motion vector predictors for the current block may be {MVP[0][X],MVP[1][1-X]} and {MVP[0][1-X],MVP[1][X]}. That is, one of the two bidirectional motion vector predictors may be induced by a combination of a motion vector predictor in the LX direction using AMVP mode and a motion vector predictor in the L(1-X) direction using merge mode. The other of the two bidirectional motion vector predictors may be induced by a combination of a motion vector predictor in the L(1-X) direction using AMVP mode and a motion vector predictor in the LX direction using merge mode.
[0134] When a single bidirectional motion vector predictor is induced for the current block, the first predicted block of the current block may be induced based on MVP[0][X], and the second predicted block of the current block may be induced based on MVP[1][1-X]. Alternatively, when a single bidirectional motion vector predictor is induced for the current block, the first predicted block of the current block may be induced based on MVP[0][1-X], and the second predicted block of the current block may be induced based on MVP[1][X].
[0135] When two bidirectional motion vector predictors are induced for the current block, the first predicted block of the current block may be induced based on MVP[0][X] and MVP[1][1-X]. For example, the first predicted block of the current block may be induced based on a weighted sum of the reference block in the LX reference picture identified by MVP[0][X] and the reference block in the L(1-X) reference picture identified by MVP[1][1-X]. Similarly, the second predicted block of the current block may be induced based on MVP[0][1-X] and MVP[1][X]. For example, the first predicted block of the current block may be induced based on a weighted sum of the reference block in the L(1-X) reference picture identified by MVP[0][1-X] and the reference block in the LX reference picture identified by MVP[1][X].
[0136] Alternatively, if two bidirectional motion vector predictors are induced for the current block, one of the two predictors may be selectively used. For this purpose, an index may be used to identify one of the two bidirectional motion predictors. The index may be signaled in the bitstream or induced based on predefined conditions identical to those of the encoding and decoding devices. Based on the bidirectional motion vector predictor identified by the index, the first and second predicted blocks of the current block may be induced, respectively.
[0137] For the sake of explanation, we will refer to one of the two pairs induced by the aforementioned pairing, {MVP[0][X],MVP[1][1-X]} and {MVP[0][1-X],MVP[1][X]}, as the first motion pair, and the other as the second motion pair.
[0138] At least one of the first motion pair and the second motion pair may be corrected based on either the bidirectional matching-based correction method or the template matching-based correction method described above. In this case, the first predicted block and / or the second predicted block of the current block may be induced based on the corrected motion pair.
[0139] Depending on whether the block currently satisfies a predetermined first condition, the first motion pair may be corrected based on either a bidirectional matching-based correction method or a template matching-based correction method. The first condition here is as described in Example 1, and a redundant explanation is omitted here.
[0140] Depending on whether the current block satisfies a predetermined second condition, the second motion pair may be corrected based on either a bidirectional matching-based correction method or a template matching-based correction method.
[0141] The second condition for the second motion pair may be defined independently of the first condition for the first motion pair. For example, the second condition may include at least one of the following: the current block makes bidirectional predictions; the bidirectional reference pictures for the current block exist before and / or after the current picture in chronological order; or the POC differences between the current picture and each reference picture are identical. The second condition for the second motion pair may be the same as the first condition for the first motion pair. Alternatively, the second condition for the second motion pair may be defined differently from the first condition for the first motion pair.
[0142] Alternatively, the second condition for the second motion pair may include at least one of the following: the first condition for the first motion pair is false; the current block performs bidirectional prediction; the POC difference between the current picture and each reference picture is the same; the cost for the first motion pair calculated by a bidirectional matching-based correction method or a template matching-based correction method is less than the first threshold; the POC difference between the current picture and the reference picture for the second motion pair is less than or equal to the POC difference between the current picture and the reference picture for the first motion pair; or the difference between the first motion pair (or the corrected first motion pair) and the second motion pair is greater than the second threshold.
[0143] If the aforementioned second condition is true, a bidirectional matching-based correction method may be used; if the second condition is false, a template matching-based correction method may be used. Alternatively, if all of the second conditions are true, a bidirectional matching-based correction method may be used; if any one of the second conditions is false, a template matching-based correction method may be used. Alternatively, if any one of the second conditions is true, a bidirectional matching-based correction method may be used; if all of the second conditions are false, a template matching-based correction method may be used.
[0144] For example, if the first condition for the first movement pair among the second conditions is true, or if all the remaining conditions among the second conditions are not true, a template matching-based correction method may be used; otherwise, a bidirectional matching-based correction method may be used.
[0145] The interpretation method relating to this disclosure may be applied with the following modifications.
[0146] In AMVP-MERGE mode, the current block may be restricted to guiding bidirectional motion information. For example, only merge candidates with bidirectional motion information may be included in the merge candidate list. Alternatively, if a merge candidate identified by the merge index has unidirectional motion information, bidirectional motion information may be guiding based on that unidirectional motion information.
[0147] Alternatively, in AMVP-MERGE mode, either MVP[0][X] or MVP[0][1-X] may be induced based on the other of the previously induced MVP[0][X] or MVP[0][1-X]. For example, a previously induced MVP[0][X] or MVP[0][1-X] can be mirrored to a reference picture in the opposite direction relative to the current picture to induce an MVP with the opposite sign.
[0148] Alternatively, in AMVP-MERGE mode, either MVP[1][X] or MVP[1][1-X] may be induced based on the other of the pre-induced MVP[1][X] or MVP[1][1-X]. For example, a pre-induced MVP[1][X] or MVP[1][1-X] can be mirrored to a reference picture in the opposite direction relative to the current picture to induce an MVP with the opposite sign.
[0149] The AMVP-MERGE mode described herein may also be applied when the current block is a block encoded in IBC (intra block copy) mode. When the interprediction mode of the current block is IBC mode, the current picture to which the current block belongs is used as the reference picture, and otherwise the method according to Embodiment 2 described above may be applied identically.
[0150] Example 3
[0151] This disclosure relates to the case where the first and second prediction modes are CIIP (combined inter and intra prediction) modes. When the CIIP mode is applied to the current block, a first prediction block and a second prediction block may be induced for the current block, respectively.
[0152] The first prediction block relating to this disclosure may be derived based on interpretation.
[0153] As an example, the first predicted block may be guided based on unidirectional or bidirectional motion information of the current block. Here, the unidirectional or bidirectional motion information may be guided based on AMVP mode or merge mode, as described in Examples 1 and 2.
[0154] The second prediction block relating to this disclosure may be guided based on intra-prediction.
[0155] As an example, one intra-prediction mode may be determined for the current block, and a second prediction block may be induced based on that intra-prediction mode. The said intra-prediction mode may be one of the following: Planar mode, MPM (Most Probable Mode), MIP mode (matrix-based intra-prediction mode), DIMD-based intra-prediction mode (described later), or TIMD-based intra-prediction mode (described later). The MPM may be one or more MPM candidates from a list of MPM candidate candidates, which may be identified by an MPM index signaled in the bitstream.
[0156] The aforementioned intra-prediction mode may be one that is predefined for the encoding and decoding devices. Alternatively, one of the predefined intra-prediction modes for the CIIP mode may be selected, and the selected intra-prediction mode may be set as one of the intra-prediction modes for the current block. Here, the predefined intra-prediction modes for the CIIP mode may include at least one of the planar mode, MPM, MIP mode, DIMD-based intra-prediction mode, or TIMD-based intra-prediction mode. An index may be used to identify one of the predefined intra-prediction modes for selection. The index may be signaled in the bitstream or induced based on conditions that are identically predefined for the encoding and decoding devices.
[0157] 1. Decoder side intra-mode derivation method (DIMD)
[0158] The gradient can be calculated based on at least two samples belonging to the surrounding region of the current block. Here, the gradient may include at least one of either a horizontal gradient or a vertical gradient. Based on at least one of the calculated gradient or gradient amplitude, an intra-prediction mode for the current block may be induced. Here, the gradient amplitude may be determined based on the sum of the horizontal gradient and the vertical gradient. This induction method may induced one intra-prediction mode for the current block, or two or more intra-prediction modes may be induced.
[0159] As an example, the gradient may be calculated in units of a window having a predetermined size. Based on the calculated gradient, an angle indicating the direction of the sample within the window can be calculated. The calculated angle may correspond to one of the predefined intra-prediction modes described above. The magnitude of the gradient may be saved / updated for the intra-prediction mode corresponding to the calculated angle. Through this process, an intra-prediction mode corresponding to the calculated gradient is determined for each window, and the magnitude of the gradient may be saved / updated for the determined intra-prediction mode. The top T intra-prediction modes with the largest magnitudes among the saved gradients may be selected, and the selected intra-prediction modes may be set as the intra-prediction modes for the current block. Here, T may be an integer of 1, 2, 3, or more.
[0160] The peripheral region used to calculate the gradient is a region previously restored to the current block and may include at least one of the following: the left-side region, the top-end region, the upper-left-end region, the lower-left-end region, or the upper-right-end region adjacent to the current block. The peripheral region may include at least one of the following: an adjacent sample line adjacent to the current block, a first non-adjacent sample line located one sample away from the current block, or a second non-adjacent sample line located two samples away from the current block. However, it may also include non-adjacent sample lines located N samples away from the current block, where N is an integer greater than or equal to 3.
[0161] The surrounding region may be a region identically predefined for the encoding and decoding devices in order to calculate the gradient. Alternatively, the surrounding region may be variably determined based on information that identifies the location of the surrounding region. In this case, the information that identifies the location of the surrounding region may be signaled in a bitstream. Alternatively, the location of the surrounding region may be determined based on at least one of the following: whether the current block is located at the boundary of a coding tree unit, the size of the current block (e.g., width, height, width-to-height ratio, width-to-height product), the division type of the current block, the prediction mode of the surrounding region, or the availability of the surrounding region.
[0162] For example, if the current block is located at the top boundary of a coding tree unit, at least one of the top, top-left, or top-right regions of the current block does not need to be referenced for gradient calculation. If the width of the current block is greater than its height, either the top region or the left region (e.g., the top region) is referenced for gradient calculation, and the other (e.g., the left region) does not need to be referenced for gradient calculation. Conversely, if the width of the current block is less than its height, either the top region or the left region (e.g., the left region) is referenced for gradient calculation, and the other (e.g., the top region) does not need to be referenced for gradient calculation. If the current block is generated by horizontal block division, the top region does not need to be referenced for gradient calculation. Conversely, if the current block is generated by vertical block division, the left region does not need to be referenced for gradient calculation. If the surrounding region of the current block is encoded in intermode, that surrounding region does not need to be referenced for gradient calculation. However, this is not limited to the above, and the surrounding region may be referenced for gradient calculation regardless of its prediction mode.
[0163] 2. Template-based intramode derivation (TIMD)
[0164] On the decoder side, an intra-prediction mode may be induced based on the template region adjacent to the current block, which will be explained in detail below.
[0165] The cost for each of the predetermined candidate modes may be calculated.
[0166] The predetermined candidate modes can mean a plurality of intra-prediction modes that are identically predefined for the encoding device and the decoding device. Alternatively, for template-based induction, a candidate list composed of the candidate modes may be generated, and costs may be calculated for the candidate modes belonging to the candidate list. Alternatively, costs may be calculated only for the top N candidate modes in the generated candidate list. Here, N may be a value that is identically predefined for the encoding device and the decoding device. For example, N may be an integer of 2, 3, 4, 5, or more.
[0167] The candidate list for template region-based induction may be constructed in the same manner as the MPM list described above. Alternatively, the candidate list may correspond to the first or second MPM list described above. Alternatively, the candidate list may consist of a combination of the first and second groups (i.e., the MPM lists) described above, or a combination of subgroups of the first and second groups (i.e., the first or second MPM list).
[0168] The aforementioned cost may be calculated as the sum of absolute difference (SAD) between the predicted and recovered samples in the template region. Alternatively, the aforementioned cost may be calculated as the sum of absolute transformed difference (SATD) between the predicted and recovered samples in the template region. Here, SATD can mean the SAD transformed into the frequency domain. As an example of the transformation, the Hadamard transform may be used, but is not limited to this. The predicted samples in the template region may be generated based on the candidate modes described above.
[0169] The template area for cost calculation may be an already restored area adjacent to the current block. For example, the template area may include at least one of the following areas of the current block: the area around the top edge, the area around the left edge, the area around the upper left edge, the area around the lower left edge, or the area around the upper right edge.
[0170] The template region may be a region predefined identically for both the encoding and decoding devices for calculating the cost. Alternatively, the template region may be variably determined based on information that identifies the location of the template region. In this case, the information identifying the location of the template region may be signaled as a bitstream. Alternatively, the location of the template region may be determined based on at least one of the following: whether the current block is located at the boundary of a coding tree unit, the size of the current block (e.g., width, height, width-to-height ratio, width-to-height product), the division type of the current block, the prediction mode of the surrounding region, or the availability of the surrounding region.
[0171] For example, if the current block is located at the top boundary of a coding tree unit, at least one of the top, top-left, or top-right peripheral regions of the current block does not need to be referenced for cost calculation. If the width of the current block is greater than its height, either the top peripheral region or the left peripheral region (e.g., the top peripheral region) is referenced for cost calculation, while the other (e.g., the left peripheral region) does not need to be referenced for cost calculation. Conversely, if the width of the current block is less than its height, either the top peripheral region or the left peripheral region (e.g., the left peripheral region) is referenced for cost calculation, while the other (e.g., the top peripheral region) does not need to be referenced for cost calculation. If the current block is generated by horizontal block division, the top peripheral region does not need to be referenced for cost calculation. Conversely, if the current block is generated by vertical block division, the left peripheral region does not need to be referenced for cost calculation. If the peripheral regions of the current block are encoded in intermode, those peripheral regions do not need to be referenced for cost calculation. However, this is not limited to the case where, regardless of the prediction mode of the surrounding region, the surrounding region may be referenced for cost calculation.
[0172] The template region may consist of N reference sample lines, where N may be an integer of 1, 2, 3, 4, or more. The number of reference sample lines constituting the template region may be the same regardless of the location of the surrounding region, or it may differ depending on the location of the surrounding region. The cost may be calculated based on all samples belonging to the template region, or the cost may be calculated using only the reference sample lines at a predetermined location within the template region, or the cost may be calculated based on all samples belonging to the reference sample line at the predetermined location, or the cost may be calculated using only the sample at the predetermined location within the reference sample line at the predetermined location. The location of the reference sample lines and / or samples for cost calculation may be determined based on at least one of the following: whether the current block is located at the boundary of a coding tree unit, the size of the current block (e.g., width, height, width-to-height ratio, width-to-height product), the division type of the current block, the prediction mode of the surrounding region, or the availability of the surrounding region. Alternatively, information indicating the location of the reference sample lines for cost calculation may be signaled in a bitstream.
[0173] From the costs calculated for each of the candidate modes, the one candidate mode with the minimum cost may be selected. For example, the cost may be calculated for each of the five candidate modes in the candidate list. The five candidate modes in the candidate list may be rearranged in ascending order of the calculated costs. The top one candidate mode may be selected from the five rearranged candidate modes.
[0174] Alternatively, at least two candidate modes having the minimum cost among the costs calculated for each candidate mode may be selected. For example, the cost may be calculated for each of the five candidate modes in the candidate list. The five candidate modes in the candidate list may be rearranged in ascending order of the calculated costs. The top two candidate modes may be selected from the five rearranged candidate modes.
[0175] One or more candidate modes selected through the process described above may be set as the intra-prediction mode for the current block.
[0176] Alternatively, if at least two candidate modes are selected by the process described above, the intra-prediction mode of the current block may be induced based on a comparison between the selected candidate modes and / or a comparison between at least one of the selected candidate modes and a threshold. For example, the intra-prediction mode of the current block may be induced based on whether or not the selected candidate modes satisfy the following conditions.
[0177] [Condition] costMode2<(KxcostMode1)
[0178] In the above conditions, costMode1 may mean the cost calculated based on one of the selected candidate modes, and costMode2 may mean the cost calculated based on the remaining selected candidate mode. For example, costMode1 may mean the cost calculated based on the candidate mode with the smaller cost among the selected candidate modes, and costMode2 may mean the cost calculated based on the candidate mode with the larger cost among the selected candidate modes. In the above conditions, K represents a predetermined comparison factor, which may be a value that is identically predefined for the encoding device and the decoding device. For example, K may be an integer of 1, 2, or more, or it may mean a real number such as 1 / 2 or 1 / 4.
[0179] If the above conditions are met, the selected candidate mode may be set as the intra-prediction mode for the current block. On the other hand, if the above conditions are not met, the candidate mode having costMode1 is set as the intra-prediction mode for the current block, and the candidate mode having costMode2 does not need to be used as the intra-prediction mode for the current block.
[0180] Alternatively, two or more intra-prediction modes may be determined for the current block, and a second prediction block may be induced based on the two or more determined intra-prediction modes. That is, prediction blocks may be induced based on two or more intra-prediction modes, and the second prediction block may be induced based on the weighted sum of the two or more induced prediction blocks. As an example, when two intra-prediction modes are determined for the current block, the second prediction block may be induced as shown in Equation 1 below.
[0181] [Formula 1]
number
[0182] In formula 1, P intra This represents the second prediction block (or a sample of the second prediction block). intra0 This represents a prediction block (or a sample of prediction blocks) that is guided based on one of the two intra-prediction modes, and P intra1 w1 represents a prediction block (or a sample of prediction blocks) induced based on the remaining intra-prediction mode of the two prediction blocks. w1 represents a weight value for the weighted sum between the two prediction blocks. For example, when shift is 2 and offset is 2, w1 may be determined to be an integer in the range of 1 to 3. Or, when shift is 3, 4, or 5, offset may be 4, 8, or 16, respectively, and w1 may be determined to be an integer in the range of 1 to 7, 1 to 15, or 1 to 32, respectively.
[0183] The two or more intra-prediction modes may include at least two of the following: planar mode, MPM, MIP mode, DIMD-based intra-prediction mode, or TIMD-based intra-prediction mode. The number of intra-prediction modes used to induce the second prediction block may be determined by a predefined number.
[0184] Example 4
[0185] This disclosure relates to the case where the first prediction mode is CIIP mode and the second prediction mode is intra mode. In this case, the first prediction block and the second prediction block may be induced for the current block, respectively.
[0186] The first prediction block relating to this disclosure may be derived based on interpretation and intrapretation.
[0187] For example, an inter-prediction block can be induced based on an inter-prediction, and an intra-prediction block can be induced based on an intra-prediction. A first prediction block may be induced based on the aforementioned induced inter-prediction block and intra-prediction block.
[0188] The interpretation block may be guided based on unidirectional or bidirectional motion information of the current block. Here, the unidirectional or bidirectional motion information may be guided based on AMVP mode or merge mode, as described in Examples 1 and 2.
[0189] Furthermore, one intra-prediction mode may be determined for the current block, and the intra-block may be guided based on that intra-prediction mode. The said intra-prediction mode may be one of the following: planar mode, MPM, MIP mode, DIMD-based intra-prediction mode, or TIMD-based intra-prediction mode. The method for determining one intra-prediction mode for the current block is as described in Example 3, and a redundant explanation will be omitted here.
[0190] The first prediction block may be induced by a weighted sum of the induced interpretation block and intraprediction block. The first prediction block may be induced as shown in Equation 2 below.
[0191] [Formula 2]
number
[0192] In equation 2, P temp This represents the first prediction block (or a sample of the first prediction block). inter represents an interpretation block (or a sample of an interpretation block), and P intra w2 represents an intra-prediction block (or a sample of intra-prediction blocks). w2 represents the weights for the weighted sum between the intra and inter-prediction blocks. The values of shift, offset, and weights in Equation 2 are as described in Equation 1.
[0193] The second prediction block relating to this disclosure may be guided based on intra-prediction.
[0194] As an example, a further intra-prediction mode may be determined for the current block, and a second prediction block may be induced based on that intra-prediction mode. The further intra-prediction mode is a different intra-prediction mode from the intra-prediction mode for the intra-prediction block of the first prediction block, and may be one of the following: planar mode, MPM, MIP mode, DIMD-based intra-prediction mode, or TIMD-based intra-prediction mode. The further intra-prediction mode may be determined based on the method for determining an intra-prediction mode for the current block described in Example 3, which is omitted here in detail.
[0195] However, the induction of the second prediction block (or the determination of further intra-prediction modes) may be omitted by conditions that are identically predefined for both the encoding and decoding devices.
[0196] Example 5
[0197] This disclosure relates to the case where the first prediction mode is CIIP mode and the second prediction mode is intra mode. In this case, the first prediction block and the second prediction block may be induced for the current block, respectively.
[0198] The first prediction block relating to this disclosure may be derived based on interpretation and intrapretation.
[0199] For example, an inter-prediction block can be induced based on an inter-prediction, and an intra-prediction block can be induced based on an intra-prediction. A first prediction block may be induced based on a weighted sum between the induced inter-prediction block and the intra-prediction block.
[0200] The interpretation block may be guided based on unidirectional or bidirectional motion information of the current block. Here, the unidirectional or bidirectional motion information may be guided based on AMVP mode or merge mode, as described in Examples 1 and 2.
[0201] Furthermore, two or more intra-prediction modes may be determined for the current block, and an intra-block may be induced based on any one of the two or more intra-prediction modes. The two or more intra-prediction modes may include at least two of the following: planar mode, MPM, MIP mode, DIMD-based intra-prediction mode, or TIMD-based intra-prediction mode. The method for determining two or more intra-prediction modes for the current block is as described in Example 3, and a redundant explanation will be omitted here.
[0202] The first prediction block may be derived by a weighted sum of the derived interpretation block and intraprediction block. The first prediction block may be derived as shown in the following equation 3.
[0203] [Formula 3]
number
[0204] In equation 3, P temprepresents the first prediction block (or samples of the first prediction block). P inter represents the inter prediction block (or samples of the inter prediction block), P intra represents the intra prediction block (or samples of the intra prediction block). w3 represents a weighting value for the weighted sum between the inter and intra prediction blocks. The values of shift, offset, and the weighting value in Equation 3 are as described in Equation 1.
[0205] The second prediction block according to the present disclosure may be derived based on intra prediction.
[0206] As an example, the second prediction block may be derived based on the remaining one of two or more intra prediction modes determined for the current block.
[0207] Example 6
[0208] The present disclosure relates to a case where the first prediction mode is the IBC mode and the second prediction mode is the intra mode. In this case, the first prediction block and the second prediction block may be respectively derived for the current block.
[0209] The first prediction block according to the present disclosure may be derived based on the IBC mode.
[0210] An IBC candidate list may be configured for the current block. The IBC candidate list may include a plurality of IBC candidates, and the plurality of IBC candidates may be derived based on surrounding blocks (or block vectors of surrounding blocks) adjacent to the current block. One or more block vectors may be derived from the IBC candidate list. The first prediction block may be derived based on the derived block vector.
[0211] As an example, one of several IBC candidates belonging to the current block's IBC candidate list may be selected. An index may be used to identify one of the several IBC candidates for the selection. Here, the index may be signaled by a bitstream. Based on the block vector of the selected IBC candidate, the block vector of the current block may be derived. Based on the derived block vector, a reference block may be identified and set as the first prediction block. Here, the reference block may belong to the current picture to which the current block belongs.
[0212] Alternatively, at least two IBC candidates may be selected from a list of IBC candidates belonging to the current block. For this purpose, a first index may be used to identify one of the multiple IBC candidates, and a second index may be used to identify the remaining one of the multiple IBC candidates. The first and second indices may be signaled by a bitstream. Alternatively, the first index may be signaled by a bitstream, and the second index may be induced based on the signaled first index. The first block vector of the current block may be induced based on the block vector of the IBC candidate selected based on the first index. The second block vector of the current block may be induced based on the block vector of the IBC candidate selected based on the second index. The first predicted block may be induced by a weighted sum of the first reference block identified based on the first block vector and the second reference block identified based on the second block vector. Here, the first and second reference blocks may belong to the current picture to which the current block belongs.
[0213] On the other hand, two different IBC candidates may be selected from among multiple IBC candidates based on the first and second indices. For example, an index may be assigned to each of the multiple IBC candidates belonging to the IBC candidate list. An IBC candidate having the same index as the first index may be selected from the IBC candidate list. If the second index is smaller than the first index, an IBC candidate having the same index as the second index may be selected from the IBC candidate list. On the other hand, if the second index is greater than or equal to the first index, an IBC candidate having the same index as the second index plus 1 may be selected from the IBC candidate list.
[0214] The second prediction block relating to this disclosure may be guided based on intra-prediction.
[0215] As an example, one intra-prediction mode may be determined for the current block, and a second prediction block may be induced based on that intra-prediction mode. The one intra-prediction mode may be any one of the following: planar mode, MPM, MIP mode, DIMD-based intra-prediction mode, or TIMD-based intra-prediction mode. This is as described in Example 3.
[0216] Alternatively, two or more intra-prediction modes may be determined for the current block, and a second prediction block may be induced based on the two or more determined intra-prediction modes. That is, prediction blocks may be induced based on two or more intra-prediction modes, and a second prediction block may be induced based on the weighted sum of the two or more induced prediction blocks. This is as explained in Example 3, and a redundant explanation will be omitted here.
[0217] Referring to Figure 4, the final predicted block for the current block can be guided based on multiple predicted blocks for the current block (S410).
[0218] In other words, the predicted block of the current block may be derived based on the weighted sum of the first and second predicted blocks of the current block. For example, the predicted block of the current block may be derived as shown in equation 4 below.
[0219] [Equation 4]
number
[0220] In formula 4, P final P0 represents the predicted block (or sample of the predicted block) of the current block, and P1 represents the second predicted block (or sample of the second predicted block). w4 represents the weight value for the weighted sum between the first and second predicted blocks. For example, when shift is 2 and offset is 2, w4 may be determined to be an integer in the range of 1 to 3. Or, when shift is 3, 4, or 5, offset may be 4, 8, or 16, respectively, and w4 may be determined to be an integer in the range of 1 to 7, 1 to 15, or 1 to 32, respectively.
[0221] Figure 5 shows the general configuration of the inter-prediction unit 332 that performs the inter-prediction method according to this disclosure.
[0222] Referring to Figure 5, the interpretation unit 332 may include a multiple prediction block guidance unit 500 and a multiple prediction block weighting unit 510.
[0223] The multiple prediction block guidance unit 500 can guide multiple prediction blocks for the current block. The multiple prediction blocks may include a first prediction block and a second prediction block. The first prediction block may be guided based on a first prediction mode, and the second prediction block may be guided based on a second prediction mode. Here, the first prediction mode and the second prediction mode may be the same prediction mode or may be different prediction modes. The first and second prediction blocks based on the first and second prediction modes may be guided based on any one of the embodiments 1 to 6 described above, and a detailed explanation is omitted here.
[0224] The multiple prediction block weighting unit 510 can guide the final prediction block of the current block based on multiple prediction blocks for the current block. That is, the multiple prediction block weighting unit 510 can guide the prediction block of the current block based on the weighted sum of the first prediction block and the second prediction block of the current block. This is explained with reference to Figure 4.
[0225] Figure 6 is a diagram illustrating an embodiment of the present disclosure, showing an interpretation method performed by an encoding device 200.
[0226] Referring to Figure 6, multiple predicted blocks can be induced for the current block (S600).
[0227] The plurality of prediction blocks may include a first prediction block and a second prediction block. The first prediction block may be induced based on a first prediction mode, and the second prediction block may be induced based on a second prediction mode. Here, the first prediction mode and the second prediction mode may be the same prediction mode or may be different prediction modes. The first and second prediction blocks based on the first and second prediction modes may be induced based on any one of the embodiments 1 to 6 described above, and a detailed explanation is omitted here.
[0228] Referring to Figure 6, the final predicted block of the current block can be derived based on multiple predicted blocks for the current block (S610). That is, the predicted block of the current block may be derived based on a weighted sum of the first and second predicted blocks of the current block, as explained with reference to Figure 4.
[0229] Figure 7 shows a schematic configuration of the inter-prediction unit 221 that performs the inter-prediction method according to this disclosure.
[0230] Referring to Figure 7, the interpretation unit 221 may include a multiple prediction block guidance unit 700 and a multiple prediction block weighting unit 710.
[0231] The multiple prediction block guidance unit 700 can guide multiple prediction blocks for the current block. The multiple prediction blocks may include a first prediction block and a second prediction block. The first prediction block may be guided based on a first prediction mode, and the second prediction block may be guided based on a second prediction mode. Here, the first prediction mode and the second prediction mode may be the same prediction mode or may be different prediction modes. The first and second prediction blocks based on the first and second prediction modes may be guided based on any one of the embodiments 1 to 6 described above, and a detailed explanation is omitted here.
[0232] The multiple prediction block weighting unit 710 can guide the final prediction block of the current block based on multiple prediction blocks for the current block. That is, the multiple prediction block weighting unit 510 can guide the prediction block of the current block based on the weighted sum of the first prediction block and the second prediction block of the current block. This is explained with reference to Figure 4.
[0233] In the embodiments described above, the method is explained based on a sequence diagram in a series of steps or blocks, but the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the sequence diagram are not exclusive, and other steps may be included, or one or more steps in the sequence diagram may be omitted without affecting the scope of the embodiments described herein.
[0234] The methods relating to the embodiments of this document described above may be implemented in software form, and the encoding and / or decoding devices relating to this document may be included in image processing devices such as TVs, computers, smartphones, set-top boxes, and display devices.
[0235] When the embodiments described in this document are implemented as software, the methods described above may be implemented as modules (processes, functions, etc.) that perform the functions described above. The modules may be stored in memory and executed by a processor. The memory may be located inside or outside the processor and may be connected to the processor by various well-known means. The processor may include an ASIC (application-specific integrated circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices. In other words, the embodiments described in this document may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.
[0236] Furthermore, the decoding and encoding devices to which the embodiments of this specification apply may include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video conferencing equipment, real-time communication equipment such as video communication, mobile streaming equipment, storage media, camcorders, video-on-demand (VoD) service providers, over-the-top (OTT) video equipment, internet streaming service providers, 3D video equipment, virtual reality (VR) equipment, argumentative reality (AR) equipment, image-phone video equipment, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video equipment, and may be used to process video signals or data signals. For example, over-the-top (OTT) video equipment may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0237] Furthermore, the processing methods to which the embodiments of this specification apply may be produced in the form of programs executed on a computer and stored on a computer-readable recording medium. Similarly, multimedia data having the data structures according to the embodiments of this specification may also be stored on a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices on which computer-readable data is stored. The computer-readable recording medium may include, for example, Blu-ray discs (BDs), general-purpose serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission over the Internet). Furthermore, a bitstream generated by an encoding method may be stored on a computer-readable recording medium or transmitted over a wireless communication network.
[0238] Furthermore, the embodiments of this specification may be embodied as computer program products in the form of program code, and the program code may be executed on a computer according to the embodiments of this specification. The program code may be stored on a computer-readable carrier.
[0239] Figure 8 shows an example of a content streaming system to which the embodiments of this disclosure can be applied.
[0240] Referring to Figure 8, the content streaming system to which the embodiments described herein apply may broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.
[0241] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and transmitting this bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server may be omitted.
[0242] The bitstream may be generated by an encoding method or bitstream generation method to which any embodiment of this specification is applied, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0243] The streaming server transmits multimedia data to the user's device based on a user request via a web server, and the web server acts as an intermediary to inform the user about available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server is responsible for controlling the commands and responses between the devices in the content streaming system.
[0244] The streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0245] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, and HMDs), digital TVs, desktop computers, and digital signage.
[0246] Each server in the aforementioned content streaming system may be operated as a distributed server, in which case the data received by each server may be processed in a distributed manner.
[0247] The claims described herein may be combined in various ways. For example, the technical features of the method claims herein may be combined to be embodied as an apparatus, and the technical features of the apparatus claims herein may be combined to be embodied as a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims herein may be combined to be embodied as an apparatus, and the technical features of the method claims and the technical features of the apparatus claims herein may be combined to be embodied as a method.
Claims
1. The current step involves inducing multiple prediction blocks from a given block, wherein the multiple prediction blocks include a first prediction block and a second prediction block. A step of inducing the final prediction block of the current block based on the first prediction block and the second prediction block, An image decoding method comprising the step of restoring the current block based on the final predicted block of the current block.
2. The first prediction block is guided based on at least one of the first L0 motion vector predictor or the second L1 motion vector predictor. The second prediction block is guided based on at least one of the first L1 motion vector predictor or the second L0 motion vector predictor. The image decoding method according to claim 1, wherein at least one of the first L0 motion vector predictor or the first L1 motion vector predictor is guided based on a first prediction mode, and at least one of the second L0 motion vector predictor or the second L1 motion vector predictor is guided based on a second prediction mode.
3. The first prediction mode is the AMVP mode, The image decoding method according to claim 2, wherein the second prediction mode is a merge mode.
4. The first L0 motion vector predictor is guided based on the first prediction mode, The image decoding method according to claim 1, wherein the first L1 motion vector predictor is guided based on the first L0 motion vector predictor which has been guided in advance.
5. At least one of the first L0 motion vector predictor or the second L1 motion vector predictor is corrected based on either a bidirectional matching-based correction method or a template matching-based correction method. The image decoding method according to claim 2, wherein one of the two-way matching-based correction method or the template matching-based correction method is selected based on a predefined first condition.
6. At least one of the first L1 motion vector predictor or the second L0 motion vector predictor is corrected based on either a bidirectional matching-based correction method or a template matching-based correction method. The image decoding method according to claim 5, wherein one of the two-way matching-based correction method or the template matching-based correction method is selected based on a predefined second condition.
7. The image decoding method according to claim 6, wherein whether or not the second condition is met is determined based on whether or not the first condition is met.
8. The aforementioned first prediction block is induced based on intermode, The image decoding method according to claim 1, wherein the second prediction block is induced based on one or more intra-prediction modes for the current block.
9. The image decoding method according to claim 8, wherein one or more intra-prediction modes for the current block include at least one of planar mode, MPM, MIP mode, DIMD-based intra-prediction mode, or TIMD-based intra-prediction mode.
10. The first prediction block is guided based on the inter-prediction block and intra-prediction block of the current block, The image decoding method according to claim 1, wherein the second prediction block is induced based on a further intra-prediction mode for the current block.
11. The first predicted block is derived based on one or more block vectors derived from the IBC candidate list of the current block. The image decoding method according to claim 1, wherein the second prediction block is induced based on one or more intra-prediction modes for the current block.
12. The current step involves inducing multiple prediction blocks from a given block, wherein the multiple prediction blocks include a first prediction block and a second prediction block. A step of inducing the final prediction block of the current block based on the first prediction block and the second prediction block, The steps include: inducing the current block's residual block based on the final predicted block of the current block; An image encoding method comprising the step of encoding the current block's residual block.
13. A computer-readable recording medium for storing a bitstream generated by an image encoding method, wherein the image encoding method is The current step involves inducing multiple prediction blocks from a given block, wherein the multiple prediction blocks include a first prediction block and a second prediction block. A step of inducing the final prediction block of the current block based on the first prediction block and the second prediction block, The steps include: inducing the current block's residual block based on the final predicted block of the current block; A computer-readable storage medium comprising the steps of encoding the current block and the current block.
14. A step of acquiring a bitstream for image information, wherein the bitstream is generated based on the steps of: inducing a plurality of prediction blocks including a first prediction block and a second prediction block for the current block; inducing a final prediction block for the current block based on the first prediction block and the second prediction block; inducing a residual block for the current block based on the final prediction block for the current block; and encoding the residual block for the current block. A data transmission method comprising the step of transmitting data including the bitstream.