A video decoding method and apparatus for deriving weighted index information for generating predictive samples.

The video decoding method enhances video coding efficiency by deriving weighted index information for affine merge candidates, addressing the need for efficient compression of high-resolution and immersive media.

JP7869388B2Active Publication Date: 2026-06-02LG ELECTRONICS INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2025-08-14
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images/videos, as well as immersive media, necessitates a highly efficient video compression technology to effectively compress, transmit, and reproduce information while reducing transmission and storage costs.

Method used

A video decoding method that includes deriving weighted index information for generating prediction samples, particularly for affine merge candidates in inter prediction, using Control Point Motion Vectors (CPMVs) to enhance video coding efficiency.

Benefits of technology

Improves overall image/video compression efficiency and enables efficient construction of motion vector candidates and weighted-value-based dual prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007869388000056
    Figure 0007869388000056
  • Figure 0007869388000057
    Figure 0007869388000057
  • Figure 0007869388000058
    Figure 0007869388000058
Patent Text Reader

Abstract

To provide a video decoding method to be executed by a decoding device.SOLUTION: A method includes: generating, based on inter-prediction mode information, a merge candidate list for a current block; deriving motion information on the current block based on a candidate selected from the merge candidate list; generating L0 prediction samples and L1 prediction samples of the current block based on the motion information; and generating prediction samples of the current block based on the L0 prediction samples, the L1 prediction samples, and a weight index for the current block. The candidates include a constructed affine merge candidate. A weight index for the constructed affine merge candidate is set to be equal to a weight index for CP 1, based on a configuration where the constructed affine merge candidate is constructed based on CPMV1 for CP1, CPMV2 for CP2 and CPMV3 for CP3.SELECTED DRAWING: Figure 14
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to a video decoding method for deriving weighted index information for generating prediction samples, and an apparatus therefor.

Background Art

[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or UHD (Ultra High Definition) images / videos of 8K or higher has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to existing image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line, or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003] Also, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of images / videos having image characteristics different from real images, such as game images, has been increasing.

[0004] Thus, there is a need for a highly efficient image / video compression technology to effectively compress, transmit, store, and reproduce information of high-resolution and high-quality images / videos having various characteristics as described above.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The technical problem of this document is to provide a method and an apparatus for enhancing the efficiency of video coding.

[0006] Another technical problem of this document is to provide a method and an apparatus for deriving weighted index information for generating prediction samples in inter prediction.

[0007] Another technical challenge of this paper is to provide a method and apparatus for deriving weighted index information for candidates in an affine merge candidate list during biprediction. [Means for solving the problem]

[0008] According to one embodiment of this document, a video decoding method performed by a decoding device is provided. The method includes the steps of: receiving video information including interprediction mode information via a bitstream; generating a list of merge candidates for the current block based on the interprediction mode information; deriving motion information for the current block based on a selected candidate from the merge candidate list; generating L0 prediction samples and L1 prediction samples for the current block based on the motion information; and generating prediction samples for the current block based on the L0 prediction samples, the L1 prediction samples and weighted value information, wherein the weighted value information is derived based on weighted value index information for the selected candidate, and the candidates include affine merge candidates, the affine merge candidates include CPMV (Control Point Motion Vector), and the affine merge candidate is located on the upper left side of the current block. If the affine merge candidate includes CPMV0 for 0), the weighted index information for the affine merge candidate is derived based on the 0th weighted index information for CP0. If the affine merge candidate does not include CPMV0 for CP0 located on the upper left side of the current block, the weighted index information for the affine merge candidate is derived based on the 1st weighted index information for CP1 (Control Point 1) located on the upper right side of the current block.

[0009] According to another embodiment of this document, a video encoding method is provided that is performed by an encoding device. The method includes the steps of: determining an interprediction mode of a current block and generating interprediction mode information indicating the interprediction mode; generating a list of merge candidates for the current block based on the interprediction mode; generating selection information indicating one of the candidates included in the merge candidate list; and encoding video information including the interprediction mode information and the selection information, wherein the candidates include affine merge candidates, and the affine merge candidates include CPMV (Control Point Motion Vector), and if the affine merge candidate includes CPMV0 for CP0 (Control Point 0) located on the upper left side of the current block, the weighted index information for the affine merge candidate is indicated based on the 0th weighted index information for CP0, and if the affine merge candidate does not include CPMV0 for CP0 located on the upper left side of the current block, the weighted index information for the affine merge candidate is indicated based on the 1st weighted index information for CP1 (Control Point 1) located on the upper right side of the current block.

[0010] In yet another embodiment of this document, a computer-readable digital storage medium is provided which stores a bitstream containing video information that causes a decoding device to perform a video decoding method. The video decoding method includes the steps of: acquiring video information including interprediction mode information via a bitstream; generating a list of merge candidates for the current block based on the interprediction mode information; selecting one candidate from among the candidates included in the merge candidate list; deriving motion information for the current block based on the selected candidate; generating L0 and L1 predicted samples for the current block based on the motion information; and generating predicted samples for the current block based on the L0 predicted samples, the L1 predicted samples and weighted value information, wherein the weighted value information is derived based on weighted value index information for the selected candidate, the candidate includes an affine merge candidate, the affine merge candidate includes a CPMV (Control Point Motion Vector), and the affine merge candidate is located on the upper left side of the current block. If the affine merge candidate includes CPMV0 for 0), the weighted index information for the affine merge candidate is derived based on the 0th weighted index information for CP0. If the affine merge candidate does not include CPMV0 for CP0 located on the upper left side of the current block, the weighted index information for the affine merge candidate is derived based on the 1st weighted index information for CP1 (Control Point 1) located on the upper right side of the current block. [Effects of the Invention]

[0011] According to this document, it is possible to improve the overall image / video compression efficiency.

[0012] According to this document, motion vector candidates can be efficiently constructed during interpretation.

[0013] According to this document, weighted-value-based dual prediction can be efficiently performed.

Brief Description of Drawings

[0014] [Figure 1] An example of a video / image coding system to which embodiments of this document can be applied is schematically shown. [Figure 2] It is a diagram schematically explaining the configuration of a video / image encoding apparatus to which embodiments of this document can be applied. [Figure 3] It is a diagram schematically explaining the configuration of a video / image decoding apparatus to which embodiments of this document can be applied. [Figure 4] An example of the inter prediction procedure is illustratively shown. [Figure 5] It is a drawing for explaining the merge mode in inter prediction. [Figure 6] An example of the motion expressed through an affine motion model is illustratively shown. [Figure 7a] An example of the CPMV for affine motion prediction is illustratively shown. [Figure 7b] An example of the CPMV for affine motion prediction is illustratively shown. [Figure 8] An example of the case where the affine MVF is determined in units of sub-blocks is illustratively shown. [Figure 9] It is a drawing for explaining the affine merge mode in inter prediction. [Figure 10] It is a drawing for explaining the candidate positions in the affine merge mode. [Figure 11] It is a drawing for explaining SbTMVP in inter prediction. [Figure 12] An example of a video / video encoding method according to embodiments of this document and related components is schematically shown. [Figure 13] An example of a video / video encoding method according to embodiments of this document and related components is schematically shown. [Figure 14]An example of a video / video decoding method and related components according to an embodiment of this document is schematically shown. [Figure 15] An example of a video / video decoding method and related components according to an embodiment of this document is schematically shown. [Figure 16] An example of a content streaming system to which the embodiments disclosed in this document can be applied is shown.

Embodiments for Carrying Out the Invention

[0015] This disclosure can be modified in various ways and can have various embodiments. Here, specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this disclosure to specific embodiments. The terms commonly used in this document are merely used to describe specific embodiments and are not intended to limit the technical idea of this disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as "including" or "having" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should be understood that the presence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof is not precluded in advance.

[0016] On the other hand, each configuration in the drawings described in this disclosure is independently illustrated for the convenience of explaining different characteristic functions, and it does not mean that each configuration is realized by separate hardware or separate software. For example, among each configuration, two or more configurations may be combined to form one configuration, and one configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of rights of this disclosure as long as they do not depart from the essence of this document.

[0017] In this specification, "A or B" can mean "A only," "B only," or "both A and B." Alternatively, in this specification, "A or B" can be interpreted as "A and / or B." For example, in this specification, "A, B, or C" can mean "A only," "B only," "C only," or "any combination of A, B, and C."

[0018] In this specification, slashes ( / ) and commas can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B, or C".

[0019] In this specification, "at least one of A and B" may mean "A only," "B only," or "both A and B." Furthermore, in this specification, the expressions "at least one of A or B" and "at least one of A and / or B" may be interpreted similarly to "at least one of A and B."

[0020] Furthermore, in this specification, "at least one of A, B and C" may mean "A only," "B only," "C only," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."

[0021] Furthermore, parentheses used in this specification can mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" is proposed as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra prediction," but rather "intra prediction" is proposed as an example of "prediction." Similarly, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" is proposed as an example of "prediction."

[0022] Technical features described individually in each drawing in this specification can be implemented individually or simultaneously.

[0023] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the attached drawings. Hereafter, the same reference numerals will be used for the same components in the drawings, and redundant descriptions of the same components may be omitted.

[0024] Figure 1 schematically illustrates an example of a video / image coding system to which this disclosure may apply.

[0025] As shown in Figure 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or network.

[0026] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may consist of a separate device or external component.

[0027] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, or a video / image archive containing previously captured video / images. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process can be replaced by the process of generating the relevant data.

[0028] An encoding device can encode input video / images. For compression and coding efficiency, the encoding device can perform a series of steps, including prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in bitstream format.

[0029] The transmitting unit can transmit encoded video / image information or data output in bitstream format to the receiving unit of a receiving device via a digital storage medium or network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit may include elements for generating media files via a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0030] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of the encoding device.

[0031] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0032] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to methods disclosed in the VVC (versatile video coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (e.g., H.267 or H.268).

[0033] This document presents various embodiments relating to video / image coding, and unless otherwise noted, these embodiments can be implemented in combination with each other.

[0034] In this document, "video" can mean a collection of images over time. "Picture" generally refers to a single image representing a specific time period, while "slice" or "tile" is a unit that constitutes part of a picture in coding. A slice or tile can contain one or more CTUs (coding tree units). A single picture can consist of one or more slices or tiles.

[0035] A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture.A tile scan can represent a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in a CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice can contain multiple complete tiles or multiple consecutive rows of CTUs within a single tile of a picture that can be contained within a single NAL unit. In this document, tile groups and slices may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.

[0036] On the other hand, a single picture can be divided into two or more subpictures. A subpicture can be a rectangular region of one or more slices within a picture.

[0037] A pixel or pel can refer to the smallest unit that makes up a picture (or image). Alternatively, the term "sample" may be used as a counterpart to pixel. A sample can generally represent a pixel or a pixel value, or it may represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. Alternatively, a sample can refer to a pixel value in the spatial domain, and if such a pixel value is converted to the frequency domain, it can also refer to the conversion coefficient in the frequency domain.

[0038] A unit can represent a basic unit of image processing. A unit can contain at least one of a specific region of a picture and information associated with that region. A unit can contain one luma block and two chroma (e.g., cb, cr) blocks. The term unit may sometimes be used interchangeably with terms such as block or area. In general, an M×N block can contain a sample (or sample array) consisting of M columns and N rows, or a set (or array) of transform coefficients.

[0039] Figure 2 is a schematic diagram illustrating the configuration of a video / image encoding device to which this disclosure may apply. Hereinafter, the term "video encoding device" may include an image encoding device.

[0040] As shown in Figure 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be called a rebuilder or a reconstructed block generator. The image segmentation unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 described above can be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. The memory 270 may also include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.

[0041] The image splitting unit 210 can split an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units can be recursively split from a coding tree unit (CTU) or the largest coding unit (LCU) using a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary-tree structure and / or the ternary structure. Alternatively, the binary-tree structure may be applied first. The coding procedure according to this disclosure may be performed based on the final coding unit that is not further split. In this case, based on coding efficiency due to image characteristics, the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further comprise a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can each be separated or partitioned from the final coding unit described above.The prediction unit may be a unit of sample prediction, and the conversion unit may be a unit for deriving conversion coefficients and / or a unit for deriving a residual signal from conversion coefficients.

[0042] The term "unit" can sometimes be used interchangeably with terms such as "block" or "area." Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the luminance (luma) component pixel / pixel value, or only the chroma component pixel / pixel value. A sample can be used as the term corresponding to a single picture (or image) pixel or pel.

[0043] The encoding device 200 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (predicted block, predicted sample array) output from the inter-prediction unit 221 or intra-prediction unit 222 from the input image signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoding device 200 can be called the subtraction unit 231. The prediction unit can make predictions for the block to be processed (hereinafter referred to as the current block) and generate a predicted block that includes the predicted sample for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied on a current block or CU basis. The prediction unit can generate various prediction-related information, such as prediction mode information, and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each prediction mode. The prediction information can be encoded by the entropy encoding unit 240 and output in bitstream format.

[0044] The intra-prediction unit 222 can predict the current block by referring to a sample in the current picture. The referenced sample can be located in the vicinity (neighbor) of the current block or at a distance, depending on the prediction mode. The prediction mode in intra-prediction can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is illustrative, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 222 can also determine the prediction mode to apply to the current block using the prediction modes applied to the surrounding blocks.

[0045] The interprediction unit 221 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the surrounding block and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, col CU, etc., and the reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the interpretation unit 221 can construct a motion information candidate list based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 221 can use the motion information of surrounding blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0046] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for a prediction of a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also be based on an intra-block copy (IBC) prediction mode or a palette mode for predictions of blocks. The IBC prediction mode or palette mode can be used for content image / video coding such as in games, for example, as in SCC (screen content coding). IBC basically performs predictions within the current picture, but can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, sample values ​​within the picture can be signaled based on information about the palette table and palette index.

[0047] The prediction signal generated via the prediction unit (comprising the inter-prediction unit 221 and / or the intra-prediction unit 222) can be used to generate a reconstructed signal or a residual signal. The transformation unit 232 can generate transformation coefficients by applying a transformation technique to the residual signal. For example, the transformation technique may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented by this graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels and obtaining a transformation based on it. The transformation process can be applied to pixel blocks of the same size that are square, or to blocks of variable size that are not square.

[0048] The quantization unit 233 quantizes the conversion coefficients and transmits them to the entropy encoding unit 240, which can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients can be called residual information. The quantization unit 233 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients. The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy encoding unit 240 can also encode information necessary for video / image restoration (e.g., the values ​​of syntax elements) together with, or separately from, the quantized conversion coefficients. Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. Information and / or syntax elements transmitted / signaled from the encoding device to the decoding device in this document may be included in the video / image information. The video / image information may be encoded via the encoding procedure described above and included in the bitstream.The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be transmitted by a transmitting unit (not shown) and / or stored by a storage unit (not shown) which are configured as internal / external elements of the encoding device 200, or the transmitting unit may be provided in the entropy encoding unit 240.

[0049] The quantized conversion coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 234 and the inverse transformation unit 235. The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 250 can be called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.

[0050] On the other hand, LMCS (luma mapping with chromium ascaling) can also be applied during the picture encoding and / or restoration process.

[0051] The filtering unit 260 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 260 can generate various filtering-related information and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each filtering method. The filtering-related information can be encoded by the entropy encoding unit 240 and output in bitstream form.

[0052] The corrected restored picture sent to memory 270 can be used as a reference picture in the interpretation unit 221. When interpretation is applied via this, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and can also improve encoding efficiency.

[0053] The DPB in memory 270 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 221. Memory 270 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 221 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 270 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 222.

[0054] On the other hand, in this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. If quantization / inverse quantization is omitted, the quantized transformation coefficient may be called a transformation coefficient. If transformation / inverse transformation is omitted, the transformation coefficient may be called a coefficient or residual coefficient, or for consistency of expression, may still be called a transformation coefficient.

[0055] Furthermore, in this document, quantized transformation coefficients and transformation coefficients may be referred to as transformation coefficients and scaled transformation coefficients, respectively. In this case, residual information may include information about the transformation coefficients (etc.), and such information may be signaled via residual coding syntax. Transformation coefficients may be derived based on the residual information (or information about the transformation coefficients (etc.)), and scaled transformation coefficients may be derived via an inverse transformation (scaling) of the transformation coefficients. Residual samples may be derived based on an inverse transformation (transformation) of the scaled transformation coefficients. This can be similarly applied / expressed in other parts of this document.

[0056] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which this disclosure may apply.

[0057] As shown in Figure 3, the decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra-predictor 331 and an inter-predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. The entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering device 350 described above can be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. The memory 360 may include a decoded picture buffer (DPB) and may also be configured by a digital storage medium. The aforementioned hardware component may also further include memory 360 as an internal / external component.

[0058] When a bitstream containing video / image information is input, the decoding device 300 can reconstruct the image corresponding to the process by which the video / image information was processed in the encoding device shown in Figure 2. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the decoding processing unit can be, for example, a coding unit, which can be divided from a coding tree unit or a maximum coding unit according to a quad-tree structure, a binary tree structure, and / or a terminally tree structure. One or more conversion units can be derived from the coding unit. The reconstructed image signal decoded and output via the decoding device 300 can then be reproduced via a playback device.

[0059] The decoding device 300 can receive the signal output from the encoding device shown in Figure 3 in bitstream form, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can further decode the picture based on the parameter set information and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements necessary for image reconstruction and the quantized values ​​of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the information of the syntax element to be decoded, the decoded information of the surrounding and decoded blocks, or the symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.Of the information decoded by the entropy decoding unit 310, information related to prediction is provided to the prediction unit (inter-prediction unit 332 and intra-prediction unit 331), and the residual values ​​that have been entropy decoded by the entropy decoding unit 310, i.e., quantized conversion coefficients and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, of the information decoded by the entropy decoding unit 310, information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives signals output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoding unit 310. On the other hand, the decoding device relating to this document may be called a video / image / picture decoding device, and the decoding device may also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, inverse transformation unit 322, addition unit 340, filtering unit 350, memory 360, inter-prediction unit 332, and intra-prediction unit 331.

[0060] The inverse quantization unit 321 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 321 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients.

[0061] In the inverse conversion unit 322, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0062] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 310, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block, and can determine a specific intra / inter-prediction mode.

[0063] The prediction unit 330 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for predictions on a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also base its predictions on an intra-block copy (IBC) prediction mode or a palette mode for predictions on a block. The IBC prediction mode or palette mode can be used for content image / video coding such as in games, for example, as in SCC (screen content coding). IBC basically performs predictions within the current picture, but can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.

[0064] The intra-prediction unit 331 can predict the current block by referring to a sample in the current picture. The referenced sample can be located in the vicinity (neighbor) of the current block or at a distance from it, depending on the prediction mode. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-prediction unit 331 can also use the prediction modes applied to the surrounding blocks to determine the prediction mode to be applied to the current block.

[0065] The interprediction unit 332 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the interprediction unit 332 can construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes, and the prediction information may include information indicating the mode of interprediction for the current block.

[0066] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (which comprises an inter-prediction unit 332 and / or an intra-prediction unit 331). If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block may be used as the restored block.

[0067] The addition unit 340 may be called the restoration unit or restoration block generation unit. The generated restoration signal can be used for intra-prediction of the next block to be processed in the current picture, and can be output after filtering as described later, or it can be used for intra-prediction of the next picture.

[0068] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.

[0069] The filtering unit 350 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.

[0070] The (modified) restored picture stored in the DPB of memory 360 can be used as a reference picture by the inter-prediction unit 332. Memory 360 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 332 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 360 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 331.

[0071] In this specification, the embodiments described for the filtering unit 260, inter-prediction unit 221, and intra-prediction unit 222 of the encoding device 200 can be applied identically or in a corresponding manner to the filtering unit 350, inter-prediction unit 332, and intra-prediction unit 331 of the decoding device 300, respectively.

[0072] On the other hand, as mentioned above, prediction is performed to improve compression efficiency when performing video coding. This makes it possible to generate a predicted block that includes predicted samples for the current block, which is the block to be coded. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is similarly derived by the encoding device and the decoding device, and the encoding device can improve image coding efficiency by signaling the decoding device with information about the residual between the original block and the predicted block (residual information), which is not the original sample value of the original block itself. The decoding device can derive a residual block that includes residual samples based on the residual information, and can generate a restored block that includes restored samples by adding the residual block and the predicted block, and can generate a restored picture that includes the restored block.

[0073] The residual information can be generated through transformation and quantization procedures. For example, an encoding device can signal the relevant residual information (via a bitstream) to a decoding device by deriving a residual block between the original block and the predicted block, performing a transformation procedure on the residual samples (residual sample array) contained in the residual block to derive transformation coefficients, and performing a quantization procedure on the transformation coefficients to derive quantized transformation coefficients. Here, the residual information may include information such as the value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transformation coefficients. The decoding device can derive a residual sample (or residual block) by performing an inverse quantization / inverse transformation procedure based on the residual information. The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also derive a residual block by inverse quantization / inverse transformation of the quantized transformation coefficients for reference for subsequent interpretation of the picture, and generate a reconstructed picture based on this.

[0074] Figure 4 illustrates the procedure for interpretation.

[0075] Referring to Figure 4, the interpretation procedure may include an interpretation mode determination step, a motion information derivation step based on the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information. The interpretation procedure can be performed in an encoding device and a decoding device, as described above. In this document, a coding device may include an encoding device and / or a decoding device.

[0076] Referring to Figure 4, the coding device determines the interpretation mode for the current block (S400). A variety of interpretation modes can be used for predicting the current block in the picture. For example, various modes such as merge mode, skip mode, MVP (motion vector prediction) mode, affine mode, subblock merge mode, and MMVD (merge with MVD) mode can be used. DMVR (Decoder side motion vector refinement) mode, AMVR (adaptive motion vector resolution) mode, Bi-prediction with CU-level weight (BCW), and Bi-directional optical flow (BDOF) can be used as additional or alternative modes. The affine mode may also be called the affine motion prediction mode. The MVP mode may also be called the AMVP (advanced motion vector prediction) mode. In this document, candidate motion information derived by some modes and / or some modes may be included among the candidate motion information for other modes. For example, an HMVP candidate may be added to the merge candidates in the merge / skip mode, or to the mvp candidates in the MVP mode. When the HMVP candidate is used as a candidate for motion information in the merge mode or skip mode, the HMVP candidate may also be called an HMVP merge candidate.

[0077] Prediction mode information indicating the inter-prediction mode of the current block can be signaled from the encoding device to the decoding device. The prediction mode information can be received by the decoding device as part of a bitstream. The prediction mode information may include index information indicating one of a number of candidate modes. Alternatively, the inter-prediction mode can be indicated via hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether the skip mode is applicable, and if the skip mode is not applicable, a merge flag may be signaled to indicate whether the merge mode is applicable, and if the merge mode is not applicable, it may indicate that the MVP mode is applicable, or additional flags for further distinctions may be signaled. Affine modes may be signaled as independent modes or as modes dependent on the merge mode or MVP mode, etc. For example, affine modes may include affine merge mode and affine MVP mode.

[0078] The coding device derives motion information for the current block (S410). The derivation of the motion information can be performed based on the inter-prediction mode.

[0079] The coding device can perform interpretation using the motion information of the current block. The encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding device can use the original block in the original picture for the current block to search for highly correlated similar reference blocks in fractional pixel units within a defined search range in the reference picture, thereby deriving motion information. Block similarity can be derived based on the difference in phase-based sample values. For example, block similarity can be calculated based on the SAD between the current block (or template of the current block) and the reference block (or template of the reference block). In this case, motion information can be derived based on the reference block with the smallest SAD within the search area. The derived motion information can be signaled to the decoding device in various ways based on the interpretation mode.

[0080] The coding device performs interpretation based on motion information for the current block (S420). The coding device can derive predicted samples for the current block based on the motion information. The current block containing the predicted samples may be called a predicted block.

[0081] Figure 5 is a diagram illustrating the merge mode in interpretation.

[0082] When merge mode is applied, the movement information of the current predicted block is not transmitted directly, but rather the movement information of the surrounding predicted blocks is used to derive the movement information of the current predicted block. Therefore, the movement information of the current predicted block can be indicated by transmitting flag information that indicates the use of merge mode and a merge index that indicates which surrounding predicted block was used. This merge mode can be called regular merge mode. For example, this merge mode can be applied when the value of the regular_merge_flag syntax element is 1.

[0083] The encoding device must search for merge candidate blocks to be used to derive motion information for the currently predicted block in order to perform merge mode. For example, up to five merge candidate blocks may be used, but the embodiments of this document (etc.) are not limited thereto. The maximum number of merge candidate blocks may be transmitted in the slice header or tile group header, but the embodiments of this document (etc.) are not limited thereto. After searching for the merge candidate blocks, the encoding device can generate a merge candidate list and select the merge candidate block with the lowest cost from among them as the final merge candidate block.

[0084] This document can provide various embodiments for merge candidate blocks that constitute the merge candidate list.

[0085] For example, the merge candidate list can use five merge candidate blocks. For instance, it can use four spatial merge candidates and one temporal merge candidate. As a specific example, in the case of a spatial merge candidate, the block shown in Figure 4 can be used as a spatial merge candidate. Hereinafter, the spatial merge candidate or the spatial MVP candidate described later may be called an SMVP, and the temporal merge candidate or the temporal MVP candidate described later may be called a TMVP.

[0086] The merge candidate list for the current block can be constructed, for example, based on the following procedure:

[0087] A coding device (encoding device / decoding device) can search for spatially surrounding blocks of the current block and insert the derived spatial merge candidates into a merge candidate list. For example, the spatially surrounding blocks may include the surrounding block at the lower left corner of the current block, the surrounding block to the left, the surrounding block at the upper right corner, the surrounding block above, and the surrounding block at the upper left corner. However, this is an example, and additional surrounding blocks such as the surrounding block to the right, the surrounding block below, and the surrounding block at the lower right can also be used as spatially surrounding blocks. The coding device can search for the spatially surrounding blocks based on priority to detect available blocks and derive the movement information of the detected blocks as spatial merge candidates. For example, an encoding device or a decoding device can search the five blocks shown in Figure 5 in the order A1->B1->B0->A0->B2, sequentially index the available candidates, and construct a merge candidate list.

[0088] The coding device can search for temporally surrounding blocks of the current block and insert the derived temporal merge candidates into the merge candidate list. The temporally surrounding blocks can be located on a reference picture that is a different picture from the current picture on which the current block is located. The reference picture on which the temporally surrounding blocks are located can be called a collocated picture or col picture. The temporally surrounding blocks can be searched for in the order of the lower right corner surrounding block and the lower right center block of the co-located blocks with respect to the current block on the col picture. On the other hand, when motion data compression is applied, specific motion information can be stored as representative motion information for each fixed storage unit in the col picture. In this case, it is not necessary to store motion information for all blocks within the fixed storage unit, and the motion data compression effect can be obtained through this. In this case, the fixed storage unit can be predetermined, for example, a 16x16 sample unit or an 8x8 sample unit, or size information for the fixed storage unit can be signaled from the encoding device to the decoding device. When motion data compression is applied, the motion information of the temporal peripheral block can be replaced with representative motion information of the fixed storage unit in which the temporal peripheral block is located. That is, in this case, from an implementation standpoint, the temporal merge candidate can be derived based on the motion information of a predicted block that covers the position after being arithmetically shifted to the right by a certain amount based on the coordinates (upper left sample position) of the temporal peripheral block, and then arithmetically shifted to the left, rather than being a predicted block located at the coordinates of the temporal peripheral block.For example, when the fixed storage unit is a 2n×2n sample unit, if the coordinates of the temporal neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>n)<<n), (yTnb>>n)<<n)) can be used for the temporal merge candidate. Specifically, for example, when the fixed storage unit is a 16×16 sample unit, if the coordinates of the temporal neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>4)<<4), (yTnb>>4)<<4)) can be used for the temporal merge candidate. Or, for example, when the fixed storage unit is an 8×8 sample unit, if the coordinates of the temporal neighboring block are (xTnb, yTnb), the motion information of the prediction block located at the corrected position ((xTnb>>3)<<3), (yTnb>>3)<<3)) can be used for the temporal merge candidate.

[0089] The coding device can check whether the number of current merge candidates is smaller than the number of maximum merge candidates. The number of the maximum merge candidates can be predefined or signaled from the encoding device to the decoding device. For example, the encoding device can generate information regarding the number of the maximum merge candidates, encode it, and transmit it to the decoder in the form of a bitstream. If all the numbers of the maximum merge candidates are filled, the subsequent candidate addition process cannot proceed.

[0090] If the verification results show that the number of current merge candidates is less than the number of maximum merge candidates, the coding device may insert additional merge candidates into the merge candidate list. For example, the additional merge candidates may include at least one of the following: history-based merge candidates(s), pair-wise average merge candidates(s), ATMVP, combined bi-predictive merge candidates (if the slice / tile group type of the current slice / tile group is type B), and / or zero-vector merge candidates.

[0091] If, as a result of the above verification, the number of current merge candidates is not less than the number of maximum merge candidates, the coding device can terminate the configuration of the merge candidate list. In this case, the encoding device can select the optimal merge candidate from among the merge candidates that make up the merge candidate list on a rate-distortion (RD) cost basis, and can signal selection information (e.g., merge index) pointing to the selected merge candidate to the decoding device. The decoding device can select the optimal merge candidate based on the merge candidate list and the selection information.

[0092] As described above, the motion information of the selected merge candidate can be used as the motion information of the current block, and predicted samples of the current block can be derived based on the motion information of the current block. The encoding device can derive the residual samples of the current block based on the predicted samples and signal the decoding device to the residual information of the residual samples. As described above, the decoding device can generate restored samples based on the residual samples derived based on the residual information and the predicted samples, and generate a restored picture based on these.

[0093] When skip mode is applied, the motion information of the current block can be derived in the same way as when merge mode is applied as described above. However, when skip mode is applied, the residual signal for the block is omitted, and therefore, the predicted sample can be immediately used as the restored sample. The skip mode can be applied, for example, when the value of the cu_skip_flag syntax element is 1.

[0094] On the other hand, the pair-wise average merge candidate can also be called a pair-wise average candidate or pair-wise candidate. A pair-wise average candidate (etc.) can be generated by averaging pairs of predefined candidates in an existing list of merge candidates. A predefined pair can be defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers can represent merge indices relative to the list of merge candidates. An averaged motion vector can be calculated separately for each reference list. For example, if two motion vectors are available in a list, the two motion vectors can be averaged even if they point to different reference pictures. For example, if only one motion vector is available, it can be used directly. For example, if no motion vectors are available, the list can be kept in an invalid state.

[0095] For example, if the merge candidate list is not filled even after pairwise average merge candidates have been added, i.e., if the current number of merge candidates in the merge candidate list is less than the maximum number of merge candidates, a zero vector (zero MVP) may be inserted at the end until the maximum merge candidate number appears. In other words, a zero vector can be inserted until the current number of merge candidates in the merge candidate list reaches the maximum number of merge candidates.

[0096] On the other hand, existing methods could use only one motion vector to represent the movement of a coding block; that is, a translation motion model could be used. However, while this method may represent the optimal movement at the block level, coding efficiency can be improved if the optimal motion vector can be determined on a sample-by-sample basis, rather than the optimal movement for each sample. For this reason, an affine motion model can be used. The affine motion prediction method, which uses an affine motion model for coding, can efficiently represent four types of motion, as will be explained later.

[0097] Figure 6 illustrates motion as represented through an affine motion model.

[0098] Referring to Figure 6, the motion that can be represented through the affine motion model can include translational motion, scaling motion, rotational motion, and shear motion. That is, not only translational motion in which (part of) the image moves in a plane according to the passage of time as shown in Figure 6, but also scaling motion in which (part of) the image is scaled according to the passage of time, rotational motion in which (part of) the image rotates according to the passage of time, and shear motion in which (part of) the image is deformed into a parallelogram according to the passage of time can be efficiently represented through the affine motion prediction.

[0099] The encoding / decoding device can predict the distortion pattern of the image based on the motion vector at the control point (CP) of the current block via the affine motion prediction, thereby improving the accuracy of the prediction and thus improving the compression performance of the image. Furthermore, the motion vector for at least one control point of the current block can be derived using the motion vectors of the surrounding blocks of the current block, reducing the data burden for the additional information and significantly improving the efficiency of interpretation.

[0100] Affine motion models that represent three of the motions that an affine motion model can represent (translation, scale, and rotation) can be called similar (or simplified) affine motion models. However, affine motion models are not limited to the aforementioned motion models.

[0101] A method for predicting affine motion can use two, three, or four motion vectors to represent the motion vector for each sample in a block.

[0102] Figures 7a and 7b illustrate CPMV for affine motion prediction.

[0103] Affine motion prediction can determine the motion vector of a sample position contained within a block using two or more control point motion vectors (CPMVs). In this case, the set of motion vectors can be referred to as an affine motion vector field (MVF).

[0104] For example, Figure 7a can show the case where two CPMVs are used, which can be called a four-parameter affine model. In this case, the motion vector at the sample position (x,y) can be determined, for example, as shown in Equation 1.

[0105]

number

[0106] For example, Figure 7b can show the case where three CPMVs are used, which can be called a 6-parameter affine model. In this case, the motion vector at the (x,y) sample position can be determined, for example, as shown in Equation 2.

[0107]

number

[0108] In equations 1 and 2, {v x ,v y} can represent the motion vector at the (x,y) position. Also, {v 0x ,v 0y} can indicate the CPMV of the control point (CP) at the upper left corner of the coding block, {v 1x ,v 1y} can indicate the CPMV of the CP at the upper right corner, {v 2x ,v 2y} can indicate the CPMV of the CP at the lower left corner. Also, W can indicate the current block width, and H can indicate the current block height.

[0109] Figure 8 illustrates an example where the affine MVF is determined at the subblock level.

[0110] During the encoding / decoding process, the affine MVF can be determined on a sample-by-sample basis or on a predefined subblock basis. For example, when determined on a sample-by-sample basis, the motion vector is obtained based on each sample value. Alternatively, when determined on a subblock basis, the motion vector for that block is obtained based on the sample value of the center of the subblock (the lower right side of the center, i.e., the lower right sample of the four central samples). In other words, in affine motion prediction, the motion vector of the current block can be derived on a sample-by-sample basis or on a subblock basis.

[0111] In Figure 8, the affine MVF is determined in units of 4x4 subblocks, but the size of the subblocks can be varied in many ways.

[0112] In other words, if affine prediction is available, there are currently three motion models that can be applied to a block: a translational motion model, a four-parameter affine motion model, and a six-parameter affine motion model. Here, the translational motion model can represent a model in which existing block unit motion vectors are used, the four-parameter affine motion model can represent a model in which two CPMVs are used, and the six-parameter affine motion model can represent a model in which three CPMVs are used.

[0113] On the other hand, affine motion prediction can include affine MVP (or affine inter) mode or affine merge mode.

[0114] Figure 9 is a diagram illustrating the affine merge mode in interpretation.

[0115] For example, in affine merge mode, CPMV can be determined by the affine motion model of the surrounding blocks coded with affine motion prediction. For instance, surrounding blocks coded with affine motion prediction on the search order can be used for affine merge mode. That is, if at least one of the surrounding blocks is coded with affine motion prediction, then the current block can be coded in affine merge mode. Here, affine merge mode can be called AF_MERGE.

[0116] When affine merge mode is applied, the CPMV of the current block may be derived using the CPMV of the surrounding blocks. In this case, the CPMV of the surrounding blocks can be used as is as the CPMV of the current block, or the CPMV of the surrounding blocks can be modified based on the size of the surrounding blocks and the size of the current block, etc., and then used as the CPMV of the current block.

[0117] On the other hand, in the case of affine merge mode, where motion vectors (MV) are derived on a subblock basis, this can be called subblock merge mode, which can be indicated based on the subblock merge flag (or the merge_subblock_flag syntax element). Alternatively, if the value of the merge_subblock_flag syntax element is 1, it may be indicated that subblock merge mode is applied. In this case, the affine merge candidate list described later can also be called the subblock merge candidate list. In this case, the subblock merge candidate list may further include candidates derived by SbTMVP described later. In this case, the candidates derived by SbTMVP can be used as the candidate at index 0 of the subblock merge candidate list. In other words, the candidates derived by SbTMVP may be located before the inherited affine candidate or constructed affine candidate described later in the subblock merge candidate list.

[0118] When affine merge mode is applied, an affine merge candidate list may be constructed for CPMV derivation for the current block. For example, the affine merge candidate list may include at least one of the following candidates: 1) inherited affine merge candidates; 2) constructed affine merge candidates; 3) zero motion vector candidates (or zero vectors). Here, the inherited affine merge candidates are those derived based on the CPMVs of the surrounding block if the surrounding block is coded in affine mode; the constructed affine merge candidates are those derived by constructing CPMVs based on the MVs of the surrounding block for each CPMV; and the zero motion vector candidates may represent candidates constructed with CPMVs whose value is 0.

[0119] The aforementioned list of affine merge candidates can be structured, for example, as follows:

[0120] There can be up to two inherited affine candidates, and these inherited affine candidates can be derived from the affine motion model of the peripheral block. A peripheral block can include one left peripheral block and an above peripheral block. Candidate blocks can be positioned as shown in Figure 4. The scan order for the left predictor can be A1→A0, and the scan order for the above predictor can be B1→B0→B2. Only one inherited candidate can be selected from each of the left and above blocks. A pruning check may not be performed between the two inherited candidates.

[0121] If a peripheral affine block is identified, the control point motion vectors of the identified block can be used to derive CPMVP candidates in the current block's affine merge list. Here, a peripheral affine block can refer to a block among the peripheral blocks of the current block that is coded in affine prediction mode. For example, referring to Figure 8, if peripheral block A on the bottom-left is coded in affine prediction mode, the motion vectors v2, v3, and v4 of the top-left, top-right, and bottom-left corners of peripheral block A can be obtained. If peripheral block A is coded with a 4-parameter affine motion model, the two CPMVs of the current block can be calculated using v2 and v3. If peripheral block A is coded with a 6-parameter affine motion model, the three CPMVs of the current block can be calculated using v2, v3, and v4.

[0122] Figure 10 is a diagram illustrating the candidate positions in affine merge mode.

[0123] A constructed affine candidate can mean a candidate formed by combining translational motion information around each control point. The motion information of a control point can be derived from its identified spatial and temporal periphery. CPMVk(k=0,1,2,3) can represent the k-th control point.

[0124] Referring to Figure 10, for CPMV0, blocks can be checked in the order B2->B3->A2, and the motion vector of the first available block can be used. For CPMV1, blocks can be checked in the order B1->B0, and for CPMV2, blocks can be checked in the order A1->A0. TMVP (temporal motion vector predictor) can be used for CPMV3 if available.

[0125] After the motion vectors of the four control points are obtained, affine merge candidates can be generated based on the acquired motion information. The combination of control point motion vectors can be any one of the following: {CPMV0, CPMV1, CPMV2}, {CPMV0, CPMV1, CPMV3}, {CPMV0, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV3}, {CPMV0, CPMV1}, and {CPMV0, CPMV2}.

[0126] A combination of three CPMVs can constitute a candidate for a 6-parameter affine merge, and a combination of two CPMVs can constitute a candidate for a 4-parameter affine merge. To avoid the motion scaling process, related combinations of control point motion vectors can be discarded if the reference indices of the control points are different from each other.

[0127] Figure 11 is a diagram illustrating SbTMVP in interpretation.

[0128] Alternatively, the SbTMVP (subblock-based temporal motion vector prediction) method can be used. For example, SbTMVP can be called ATMVP (advanced temporal motion vector prediction). SbTMVP can utilize motion fields in collocated pictures to improve motion vector prediction and the merge mode for CUs in the current picture. Here, collocated pictures can also be called col pictures.

[0129] For example, SbTMVP can predict motion at the subblock (or subCU) level. Furthermore, SbTMVP can apply a motion shift before fetching temporal motion information from the colpicture. Here, the motion shift can be obtained from the motion vector of one of the spatially surrounding blocks of the current block.

[0130] SbTMVP can predict the motion vectors of subblocks (or subCUs) within a current block (or CU) in two steps.

[0131] In the first step, spatial peripheral blocks can be tested in the order A1, B1, B0, and A0 in Figure 4. The first spatial peripheral block may be identified that has a motion vector using the col picture as its reference picture, and the motion vector can be selected with the applied motion shift. If no such motion is identified from the spatial peripheral block, the motion shift can be set to (0,0).

[0132] In the second step, the motion shift confirmed in the first step can be applied to obtain subblock-level motion information (motion vectors and reference indices) from the col picture. For example, the motion shift can be added to the coordinates of the current block. For example, the motion shift can be set as the motion of A1 in Figure 4. In this case, for each subblock, the motion information of the corresponding block in the col picture can be used to derive the subblock motion information. Temporal motion scaling can be applied to align the reference picture of the temporal motion vector with the reference picture of the current block.

[0133] A combined subblock-based merge list containing all SbTVMP candidates and affine merge candidates can be used for signaling the affine merge mode. Here, the affine merge mode can be called the subblock-based merge mode. The SbTVMP mode can be available or unavailable by a flag included in the SPS (sequence parameter set). If the SbTMVP mode is available, the SbTMVP predictor can be added as the first entry in the list of subblock-based merge candidates, followed by the affine merge candidates. The maximum allowed size of the affine merge candidate list can be 5.

[0134] The size of the subCU (or subblock) used in SbTMVP can be fixed at 8x8, and, similar to affine merge mode, SbTMVP mode can only be applied to blocks where both the width and height are 8 or greater. The encoding logic for additional SbTMVP merge candidates can be the same as for other merge candidates. That is, an RD check utilizing an additional RD (rate-distortion) cost can be performed for each CU in a P or B slice to determine whether to use an SbTMVP candidate.

[0135] On the other hand, a predicted block can be derived for the current block based on motion information derived by the prediction mode. The predicted block may include predicted samples (predicted sample arrays) of the current block. If the motion vector of the current block points to fractional sample units, an interpolation procedure may be performed, through which predicted samples of the current block can be derived based on reference samples in fractional sample units within the reference picture. When affine interpretation (affine prediction mode) is applied to the current block, predicted samples can be generated based on sample / subblock unit MV. When biprediction is applied, predicted samples derived via a (phase-based) weighted sum or weighted average of predicted samples derived based on L0 prediction (i.e., prediction using reference pictures in reference picture list L0 and MVL0) and predicted samples derived based on L1 prediction (i.e., prediction using reference pictures in reference picture list L1 and MVL1) can be used as predicted samples of the current block. Here, the motion vector in the L0 direction can be called the L0 motion vector or MVL0, and the motion vector in the L1 direction can be called the L1 motion vector or MVL1. When bi-prediction is applied, if the reference picture used for L0 prediction and the reference picture used for L1 prediction are located in different temporal directions relative to the current picture (i.e., it is bi-prediction but corresponds to bidirectional prediction), this can be called true bi-prediction.

[0136] As mentioned earlier, reconstructed samples and pictures can be generated based on the derived predicted samples, and then procedures such as in-loop filtering can be performed.

[0137] On the other hand, when bi-prediction is applied to a block, the predicted samples can be derived based on a weighted average. For example, bi-prediction using a weighted average can be called BCW (Bi-prediction with CU-level Weight), BWA (Bi-prediction with Weighted Average), or weighted averaging bi-prediction.

[0138] Previously, the dual prediction signal (i.e., dual prediction sample) could be derived by a simple average of the L0 prediction signal (L0 prediction sample) and the L1 prediction signal (L1 prediction sample). That is, the dual prediction sample was derived as the average of the L0 prediction sample based on the L0 reference picture and MVL0 and the L1 prediction sample based on the L1 reference picture and MVL1. However, when dual prediction is applied, the dual prediction signal (dual prediction sample) can also be derived by a weighted average of the L0 prediction signal and the L1 prediction signal, as follows. For example, the dual prediction signal (dual prediction sample) can be derived as shown in Equation 3.

[0139]

number

[0140] In equation 3, Pbi-pred can represent the value of the biprediction signal, i.e., the predicted sample value derived by applying biprediction, and w can represent the weighted value. Furthermore, P0 can represent the value of the L0prediction signal, i.e., the predicted sample value derived by applying L0prediction, and P1 can represent the value of the L1prediction signal, i.e., the predicted sample value derived by applying L1prediction.

[0141] For example, in weighted average biprediction, five weight values ​​may be allowed. For example, the five weight values ​​(w) may include -2, 3, 4, 5, or 10. That is, the weight value (w) can be determined to be one of the weight value candidates including -2, 3, 4, 5, or 10. For each CU to which biprediction is applied, the weight value w can be determined by one of two methods. The first method is that for unmerged CUs, the weight value index can be signaled after the motion vector difference. The second method is that for merged CUs, the weight value index can be inferred from surrounding blocks based on the merge candidate index.

[0142] For example, weighted average biprediction can be applied to CUs with 256 or more luma samples. That is, weighted average biprediction can be applied when the product of the width and height of the CU is greater than or equal to 256. For low-delay pictures, five weight values ​​may be used, and for non-low-delay pictures, three weight values ​​may be used. For example, the three weight values ​​may include 3, 4, or 5.

[0143] For example, in an encoding device, a fast search algorithm can be applied to find weighted indexes without significantly increasing the complexity of the encoding device. Such algorithms can be summarized as follows: For example, when combined with AMVR (adaptive motion vector resolution) (when AMVR is used as the interprediction mode), if the current picture is a low-latency picture, non-identical weighted values ​​can be conditionally checked for 1-pel and 4-pel motion vector precision. For example, when combined with affine (when affine prediction mode is used as the interprediction mode), if the affine prediction mode is selected as the best mode, affine ME (Motion Estimation) can be performed on non-identical weighted values. For example, if the two reference pictures in a biprediction are identical, non-identical weighted values ​​can be conditionally checked. For example, depending on the POC distance between the current picture and the reference picture, the coding QP (quantization parameter), and the temporal level, certain conditions may be met that result in non-identical weighted values ​​not being found.

[0144] For example, a BCW weighted index (or weighted index) can be coded using one context-coded bin and subsequent bypass-coded bins. The first context-coded bin can indicate whether the same weights are used or not. If non-identical weights are used based on the first context-coded bin, an additional bin may be signaled using bypass coding to indicate the non-identical weights to be used.

[0145] On the other hand, when dual prediction is applied, the weighted value information used to generate prediction samples can be derived based on the weighted value index information for the selected candidate from among the candidates included in the merge candidate list.

[0146] According to one embodiment of this document, when constructing motion vector candidates for merge mode, weighted index information for temporal motion vector candidates can be derived as follows. For example, when a temporal motion vector candidate uses bi-prediction, weighted index information for weighted average can be derived. That is, when the inter-prediction type is bi-prediction, weighted index information for temporal merge candidates (or temporal motion vector candidates) in the merge candidate list can be derived.

[0147] For example, the weighted index information for the weighted average of a candidate time motion vector can always be derived to 0. Here, a weighted index information of 0 can mean that the weights for each reference direction (i.e., the L0 and L1 prediction directions in biprediction) are the same. For example, the procedure for deriving the motion vector of the luma component for merge mode may be as shown in Table 1 below.

[0148] [Table 1-1]

[0149] [Table 1-2]

[0150] [Table 1-3]

[0151] [Table 1-4]

[0152] [Table 1-5]

[0153] Referring to Table 1 above, gbiIdx can represent a bipredictive weighted index, and gbiIdxCol can represent a bipredictive weighted index for a temporal merge candidate (e.g., a temporal motion vector candidate in the merge candidate list). In the procedure for deriving the motion vector of the luma component for the merge mode (Table of Contents 3 in 8.4.2.2), gbiIdxCol can be derived to 0. That is, the weighted index of the temporal motion vector candidate can be derived to 0.

[0154] Alternatively, the weighted index for the weighted average of the temporal motion vector candidates can be derived based on the weighted index information of the collocated block. Here, a collocated block can be called a col block, co-located block, or co-located reference block, and a col block can represent a block at the same position as the current block on the reference picture. For example, the procedure for deriving the motion vector of the luma component for merge mode may be as shown in Table 2 below.

[0155] [Table 2-1]

[0156] [Table 2-2]

[0157] [Table 2-3]

[0158] [Table 2-4]

[0159] [Table 2-5]

[0160] Referring to Table 2 above, gbiIdx can represent a bipredictive weighted index, and gbiIdxCol can represent a bipredictive weighted index for a temporal merge candidate (e.g., a temporal motion vector candidate in the merge candidate list). When the slice type or tile group type is B in the procedure for deriving the motion vector of the luma component for the merge mode (Table of Contents 4 in 8.4.2.2), then gbiIdxCol can be derived using gbiIdxCol. That is, the weighted index of the temporal motion vector candidate can be derived using the weighted index of the col block.

[0161] On the other hand, according to other embodiments of this document, when constructing motion vector candidate configurations for subblock-level merge modes, a weighted index for a weighted average of temporal motion vector candidates can be derived. Here, the subblock-level merge mode may be called an affine merge mode (subblock-level). Temporal motion vector candidates may represent subblock-based temporal motion vector candidates and may also be called SbTMVP (or ATMVP) candidates. That is, when the inter-prediction type is bi-prediction, weighted index information for SbTMVP candidates (or subblock-based temporal motion vector candidates) in the affine merge candidate list or subblock merge candidate list can be derived.

[0162] For example, the weighted index information for the weighted average of the temporal motion vector candidates for the subblock base can always be derived to 0. Here, a weighted index information of 0 can mean that the weights for each reference direction (i.e., the L0 and L1 prediction directions in biprediction) are the same. For example, the procedure for deriving the motion vectors and reference indices in the subblock merge mode and the procedure for deriving the temporal merge candidates for the subblock base can be as shown in Tables 3 and 4 below, respectively.

[0163] [Table 3-1]

[0164] [Table 3-2]

[0165] [Table 3-3]

[0166] [Table 3-4]

[0167] [Table 3-5]

[0168] [Table 3-6]

[0169] [Table 3-7]

[0170] [Table 3-8]

[0171] [Table 4]

[0172] Referring to Tables 3 and 4 above, gbiIdx can represent a biprediction weighted index, gbiIdxSbCol can represent a biprediction weighted index for a subblock base temporal merge candidate (e.g., a temporal motion vector candidate in the subblock base merge candidate list), and in the procedure for deriving the subblock base temporal merge candidate (8.4.4.3), gbiIdxSbCol can be derived to 0. That is, the weighted index for a subblock base temporal motion vector candidate can be derived to 0.

[0173] Alternatively, weighted index information for a weighted average of candidate time motion vectors for a subblock base can be derived based on weighted index information for a time center block. For example, the time center block can represent a col block or a subblock or sample located in the center of a col block, specifically, a subblock or sample located in the lower right of the four central subblocks or samples in a col block. For example, in this case, the procedure for deriving motion vectors and reference indices within the subblock merge mode, the procedure for deriving candidate time merges for the subblock base, and the procedure for deriving base motion information for time merging of the subblock base may be as shown in Tables 5, 6, and 7 below, respectively.

[0174] [Table 5-1]

[0175] [Table 5-2]

[0176] Table 5-3

[0177] Table 5-4

[0178] Table 5-5

[0179] Table 5-6

[0180] Table 5-7

[0181] Table 5-8

[0182] Table 6-1

[0183] Table 6-2

[0184] Table 6-3

[0185] Table 6-4

[0186] [Table 7-1]

[0187] [Table 7-2]

[0188] [Table 7-3]

[0189] [Table 7-4]

[0190] Referring to Tables 5, 6, and 7, gbiIdx can represent a bipredictive weighted index, and gbiIdxSbCol can represent a bipredictive weighted index for a subblock-based temporal merge candidate (e.g., a temporal motion vector candidate in the subblock-based merge candidate list). In the procedure for deriving base motion information for subblock-based temporal merging (8.4.4.4), gbiIdxSbCol can be derived from gbiIdxcolCb. That is, the weighted index for a subblock-based temporal motion vector candidate can be derived from the weighted index of the temporal center block. For example, the temporal center block can represent a col block or a subblock or sample located in the center of a col block, specifically, a subblock or sample located in the lower right of the four central subblocks or samples in a col block.

[0191] Alternatively, the weighted index information for the weighted average of the temporal motion vector candidates for the subblock base can be derived based on the weighted index information for each subblock unit, or, if the subblock is not available, based on the weighted index information of the temporal center block. For example, the temporal center block can represent a col block or a subblock or sample located in the center of a col block, specifically, a subblock or sample located in the lower right of the four central subblocks or samples in a col block. For example, in this case, the procedure for deriving the motion vectors and reference indices in the subblock merge mode, the procedure for deriving the temporal merge candidates for the subblock base, and the procedure for deriving the base motion information for the temporal merge of the subblock base may be as shown in Tables 8, 9, and 10 below.

[0192] [Table 8-1]

[0193] [Table 8-2]

[0194] [Table 8-3]

[0195] [Table 8-4]

[0196] [Table 8-5]

[0197] [Table 8-6]

[0198] Table 8-7

[0199] Table 8-8

[0200] Table 9-1

[0201] Table 9-2

[0202] Table 9-3

[0203] Table 9-4

[0204] Table 9-5

[0205] Table 10-1

[0206] Table 10-2

[0207] Table 10-3

[0208]

Table 10-4

[0209] Referring to Table 8, Table 9, and Table 10 above, gbiIdx can represent the dual-prediction weighted value index, and gbiIdxSbCol can represent the dual-prediction weighted value index for the time merge candidates of the sub-block base (e.g., the time motion vector candidates in the merge candidate list of the sub-block base). In the procedure (8.4.4.3) for deriving the base motion information for the time merge of the sub-block base, the gbiIdxSbCol can be derived from gbiIdxcolCb. Or, depending on conditions (e.g., when both availableFlagL0SbCol and availableFlagL1SbCol are 0), in the procedure (8.4.4.3) for deriving the base motion information for the time merge of the sub-block base, the gbiIdxSbCol can be derived from ctrgbiIdx, and in the procedure (8.4.4.4) for deriving the base motion information for the time merge of the sub-block base, the ctrgbiIdx can be derived from gbiIdxSbCol. That is, the weighted value index of the time motion vector candidates of the sub-block base can be derived from the weighted value index of each sub-block unit, and when the sub-block is not available, it can be derived from the weighted value index of the time center block. For example, the time center block can represent a col block or a sub-block or sample located at the center of the col block. Specifically, it can represent the sub-block or sample located in the lower right of the four sub-blocks or samples at the center of the col block.

[0210] On the other hand, according to yet another embodiment of this document, when constructing motion vector candidate configurations for merge mode, weighted index information for pair-wise candidates can be derived. For example, pair-wise candidates may be included in the merge candidate list, and weighted index information for the weighted average of the pair-wise candidates can be derived. The pair-wise candidates can be derived based on other merge candidates in the merge candidate list, and if the pair-wise candidates use bi-prediction, weighted index information for the weighted average can be derived. That is, if the inter-prediction type is bi-prediction, weighted index information for pair-wise candidates in the merge candidate list can be derived.

[0211] The pairwise candidates can be derived based on two other merge candidates (e.g., cand0 and cand1) from among the candidates included in the merge candidate list.

[0212] For example, the weighted index information for the pairwise candidate can be derived based on the weighted index information of one of the two merge candidates (e.g., merge candidate cand0 or merge candidate cand1). For example, the weighted index information for the pairwise candidate can be derived based on the weighted index information of the candidate that uses dual predictions among the two merge candidates.

[0213] Alternatively, if the weighted index information for each of the other two merge candidates is the same as the first weighted index information, the weighted index information for the pairwise candidate can be derived based on the first weighted index information. On the other hand, if the weighted index information for each of the other two merge candidates is not the same as each other, the weighted index information for the pairwise candidate can be derived based on the default weighted index information. The default weighted index information may correspond to weighted index information that assigns the same weight to each of the L0 and L1 predicted samples.

[0214] Alternatively, if the weighted index information of the other two merge candidates is the same as the first weighted index information, the weighted index information for the pairwise candidate can be derived based on the first weighted index information. On the other hand, if the weighted index information of the other two merge candidates is not the same as each other, the weighted index information for the pairwise candidate can be derived based on the weighted index information of the other two candidates that is not the default weighted index information. The default weighted index information may correspond to weighted index information that assigns the same weight to each of the L0 and L1 prediction samples.

[0215] On the other hand, according to yet another embodiment of this document, when constructing motion vector candidate configurations for a subblock-based merge mode, weighted index information for a weighted average of the temporal motion vector candidates can be derived. Here, the subblock-based merge mode can be called an affine merge mode (subblock-based). The temporal motion vector candidate can represent a subblock-based temporal motion vector candidate and can also be called an SbTMVP (or ATMVP) candidate. The weighted index information for the SbTMVP candidate can be derived based on the weighted index information of the surrounding blocks to the left of the current block. That is, if the candidate derived by SbTMVP uses biprediction, the weighted index of the surrounding blocks to the left of the current block can be derived as a weighted index for the subblock-based merge mode.

[0216] For example, since a col block can be derived from a spatially adjacent block to the left (or surrounding block to the left) of the current block, the weighted index of the surrounding block to the left can be considered reliable. Thus, the weighted index for the SbTMVP candidate can be derived using the weighted index of the surrounding block to the left.

[0217] On the other hand, according to yet another embodiment of this document, when constructing motion vector candidates for affine merge mode, weighted index information for a weighted average can be derived when the affine merge candidates use bi-prediction. That is, when the inter-prediction type is bi-prediction, weighted index information can be derived for the candidates in the affine merge candidate list or the subblock merge candidate list.

[0218] For example, among the affine merge candidates, a constructed affine merge candidate can represent a candidate from which CP0, CP1, CP2, or CP3 candidates are derived in the affine model, based on the motion information of spatially adjacent blocks (or spatially surrounding blocks) or temporally adjacent blocks (or temporally surrounding blocks) of the current block, thereby deriving the MVF. For example, CP0 can represent a control point located at the upper left sample position of the current block, CP1 can represent a control point located at the upper right sample position of the current block, and CP2 can represent a control point located at the lower left sample position of the current block. Additionally, CP3 can represent a control point located at the lower right sample position of the current block.

[0219] For example, among the affine merge candidates, constructed affine merge candidates can be generated based on the combination of each control point in the current block, such as {CP0,CP1,CP2}, {CP0,CP1,CP3}, {CP0,CP2,CP3}, {CP1,CP2,CP3}, {CP0,CP1}, and {CP0,CP2}. For example, an affine merge candidate may include at least one of {CPMV0,CPMV1,CPMV2}, {CPMV0,CPMV1,CPMV3}, {CPMV0,CPMV2,CPMV3}, {CPMV1,CPMV2,CPMV3}, {CPMV0,CPMV1}, and {CPMV0,CPMV2}. CPMV0, CPMV1, CPMV2, and CPMV3 may correspond to the motion vectors for CP0, CP1, CP2, and CP3, respectively.

[0220] In one embodiment, if the affine merge candidate includes a CPMV0 for a CP0 (Control Point 0) located on the upper left side of the current block, the weighted index information for the affine merge candidate can be derived based on the zero-weighted index information for the CP0. The zero-weighted index information may correspond to the weighted index information of a block among the surrounding blocks of the CP0 that is used for deriving the CPMV0. In this case, the surrounding blocks of the CP0 may include the surrounding block at the upper left corner of the current block, the surrounding block to the left adjacent to the lower side of the surrounding block at the upper left corner, and the surrounding block to the upper side adjacent to the right side of the surrounding block at the upper left corner.

[0221] On the other hand, if the affine merge candidate does not include CPMV0 for CP0 located on the upper left side of the current block, the weighted index information for the affine merge candidate can be derived based on the first weighted index information for CP1 (Control Point 1) located on the upper right side of the current block. The first weighted index information may correspond to the weighted index information of the block used for deriving CPMV1 among the surrounding blocks of CP1. In that case, the surrounding blocks of CP1 may include the surrounding block at the upper right corner of the current block and the upper surrounding block adjacent to the left of the surrounding block at the upper right corner.

[0222] According to the method described above, the weighted index information for the affine merge candidates can be derived for each of {CPMV0,CPMV1,CPMV2}, {CPMV0,CPMV1,CPMV3}, {CPMV0,CPMV2,CPMV3}, {CPMV1,CPMV2,CPMV3}, {CPMV0,CPMV1}, and {CPMV0,CPMV2} based on the weighted index information of the block used for the derivation of the first CPMV.

[0223] In another embodiment for deriving weighted index information for the affine merge candidate, if the weighted index information for CP0 located on the upper left side of the current block and the weighted index information for CP1 located on the upper right side of the current block are the same, the weighted index information for the affine merge candidate can be derived based on the zero weighted index information for CP0. The zero weighted index information may correspond to the weighted index information of the block used for deriving CPMV0 among the surrounding blocks of CP0. On the other hand, if the weighted index information for CP0 located on the upper left side of the current block and the weighted index information for CP1 located on the upper right side of the current block are not the same, the weighted index information for the affine merge candidate can be derived based on default weighted index information. The default weighted index information may correspond to weighted index information that assigns the same weight to each of the L0 prediction sample and the L1 prediction sample.

[0224] According to yet another embodiment for deriving weighted index information for the affine merge candidate, the weighted index information for the affine merge candidate can be derived using the weighted index of the candidate with the highest occurrence frequency among the weighted indexes of each candidate. For example, among the CP0 candidate block, the weighted index of the candidate block determined to be the motion vector in CP0; among the CP1 candidate block, the weighted index of the candidate block determined to be the motion vector in CP1; among the CP2 candidate block, the weighted index of the candidate block determined to be the motion vector in CP2; and / or among the CP3 candidate block, the weighted index of the candidate block determined to be the motion vector in CP3 that has the most overlap can be derived as the weighted index of the affine merge candidate.

[0225] For example, CP0 and CP1 may be used as the control points, or CP0, CP1, and CP2 may be used, and CP3 may not be used. However, for example, when trying to utilize a candidate CP3 for an affine block (a block coded in affine prediction mode), the method for deriving the weighted index in the temporal candidate block described in the above-mentioned embodiment may be used.

[0226] Figures 12 and 13 schematically show an example of a video / image encoding method and related components according to the embodiment of this document.

[0227] The method disclosed in Figure 12 can be performed by the encoding device disclosed in Figure 2 or Figure 13. Specifically, for example, steps S1200 to S1220 in Figure 12 can be performed by the prediction unit 220 of the encoding device 200 in Figure 13, and step S1230 in Figure 12 can be performed by the entropy encoding unit 240 of the encoding device 200 in Figure 13. Although not shown in Figure 12, the prediction unit 220 of the encoding device 200 can derive predicted samples or prediction-related information in Figure 12, the residual processing unit 230 of the encoding device 200 can derive residual information from original samples or predicted samples, and the entropy encoding unit 240 of the encoding device 200 can generate a bitstream from the residual information or prediction-related information. The method disclosed in Figure 12 may include the embodiments described above in this document.

[0228] Referring to Figure 12, the encoding device can determine the inter-prediction mode of the current block and generate inter-prediction mode information indicating the inter-prediction mode (S1200). For example, the encoding device can determine a merge mode, an affine (merge) mode, or a sub-block merge mode as the inter-prediction mode to be applied to the current block, and generate inter-prediction mode information indicating this.

[0229] The encoding device can generate a merge candidate list for the current block based on the inter prediction mode (S1210). For example, the encoding device can generate the merge candidate list according to the determined inter prediction mode. Here, when the determined inter prediction mode is the affine merge mode or the sub-block merge mode, the merge candidate list may be called, for example, an affine merge candidate list or a sub-block merge candidate list, etc., but may also be simply called a merge candidate list.

[0230] For example, candidates can be inserted into the merge candidate list until the number of candidates in the merge candidate list reaches the maximum number of candidates. Here, a candidate can represent a candidate or a candidate block for deriving the motion information (or motion vector) of the current block. For example, the candidate block can be derived through searching for peripheral blocks of the current block. For example, the peripheral blocks can include spatial peripheral blocks and / or temporal peripheral blocks of the current block. The spatial peripheral blocks can be searched first (spatial merge) to derive candidates, and then the temporal peripheral blocks can be searched (temporal merge) to derive candidates, and the derived candidates can be inserted into the merge candidate list. For example, even after inserting the candidates into the merge candidate list, additional candidates can be inserted if the number of candidates in the merge candidate list is less than the maximum number of candidates. For example, the additional candidates can include at least one of history based merge candidate(s), pair-wise average merge candidate(s), ATMVP, combined bi-predictive merge candidate (when the slice / tile group type of the current slice / tile group is of type B), and / or zero vector merge candidate.

[0231] Alternatively, for example, candidates may be inserted into the affine merge candidate list until the number of candidates in the affine merge candidate list reaches the maximum number of candidates. Here, a candidate may include the CPMV (Control Point Motion Vector) of the current block. Alternatively, the candidate may represent a candidate or candidate block for deriving the CPMV. The CPMV may represent the motion vector at the CP (Control Point) of the current block. For example, there may be two, three, or four CPs, and they may be located on at least part of the upper left side (or upper left corner), upper right side (or upper right corner), lower left side (or lower left corner), or lower right side (or lower right corner) of the current block, with only one CP at each position.

[0232] For example, candidates can be derived through a search of the surrounding blocks of the current block (or the surrounding blocks of the current block's CP). For example, an affine merge candidate list can contain at least one of inherited affine merge candidates, constructed affine merge candidates, or zero motion vector candidates. For example, an affine merge candidate list can first insert the inherited affine merge candidates, and then insert the constructed affine merge candidates. Also, if the number of candidates in the affine merge candidate list is less than the maximum number of candidates after inserting the constructed affine merge candidates, the remainder can be filled with zero motion vector candidates. Here, zero motion vector candidates can also be called zero vectors. For example, an affine merge candidate list can be a list of affine merge modes in which motion vectors are derived on a sample-by-sample basis, or it can be a list of affine merge modes in which motion vectors are derived on a subblock-by-subblock basis. In this case, the affine merge candidate list can be called the subblock merge candidate list, and the subblock merge candidate list can also include candidates derived by SbTMVP (or SbTMVP candidates). For example, if an SbTMVP candidate is included in the subblock merge candidate list, it can be positioned before inherited affine merge candidates and constructed affine merge candidates within the subblock merge candidate list.

[0233] The encoding device can generate selection information indicating one of the candidates included in the merge candidate list (S1220). For example, the merge candidate list may include at least some of spatial merge candidates, temporal merge candidates, pairwise candidates, or zero vector candidates, and one of these candidates can be selected for inter prediction of the current block. Alternatively, for example, the subblock merge candidate list may include at least some of inherited affine merge candidates, constructed affine merge candidates, SbTMVP candidates, or zero vector candidates, and one of these candidates can be selected for inter prediction of the current block.

[0234] For example, the selection information may include index information indicating the selected candidate within the merge candidate list. For example, the selection information may also be called merge index information or subblock merge index information.

[0235] Furthermore, the encoding device can generate inter-prediction type information that indicates the current block's inter-prediction type as a bi-prediction. For example, the current block's inter-prediction type can be determined to be a bi-prediction from among L0 prediction, L1 prediction, or bi-prediction, and the encoding device can generate inter-prediction type information that indicates this. Here, L0 prediction can indicate a prediction based on reference picture list 0, L1 prediction can indicate a prediction based on reference picture list 1, and bi-prediction can indicate a prediction based on both reference picture list 0 and reference picture list 1. For example, the encoding device can generate inter-prediction type information based on the inter-prediction type. For example, the inter-prediction type information may include the inter_pred_idc syntax element.

[0236] The encoding device can encode video information including interprediction mode information and selection information (S1230). For example, the video information may also be called video information. The video information may include various types of information according to the embodiments described above in this document. For example, the video information may include at least a portion of prediction-related information or residual-related information. For example, the prediction-related information may include at least a portion of the interprediction mode information, selection information, and interprediction type information. For example, the encoding device can encode video information including all or part of the aforementioned information (or syntax elements) and generate a bitstream or encoded information. Alternatively, it can output in the form of a bitstream. The bitstream or encoded information can also be transmitted to a decoding device via a network or storage medium.

[0237] Although not shown in Figure 12, for example, the encoding device can generate predicted samples of the current block. Alternatively, for example, the encoding device can generate predicted samples of the current block based on selected candidates. Or, for example, the encoding device can derive motion information based on selected candidates and generate predicted samples of the current block based on the motion information. For example, the encoding device can generate L0 predicted samples and L1 predicted samples by dual prediction, and generate predicted samples of the current block based on the L0 predicted samples and the L1 predicted samples. In this case, weighted index information (or weighted information) for dual prediction can be used to generate predicted samples of the current block from the L0 predicted samples and the L1 predicted samples. Here, the weighted information can be shown based on the weighted index information.

[0238] In other words, for example, the encoding device can generate L0 and L1 prediction samples for the current block based on the selected candidates. For example, if the interpretation type of the current block is determined to be biprediction, then reference picture list 0 and reference picture list 1 may be used for the prediction of the current block. For example, the L0 prediction sample may represent a prediction sample of the current block derived based on reference picture list 0, and the L1 prediction sample may represent a prediction sample of the current block derived based on reference picture list 1.

[0239] For example, the candidates may include spatial merge candidates. For example, if the selected candidate is a spatial merge candidate, L0 motion information and L1 motion information may be derived based on the spatial merge candidate, and the L0 prediction sample and the L1 prediction sample may be generated based on this.

[0240] For example, the candidates may include temporal merge candidates. For example, if the selected candidate is a temporal merge candidate, L0 motion information and L1 motion information may be derived based on the temporal merge candidate, and the L0 prediction sample and the L1 prediction sample may be generated based on this.

[0241] For example, the candidates may include pairwise candidates. For example, if the selected candidate is a pairwise candidate, L0 motion information and L1 motion information may be derived based on the pairwise candidate, and the L0 prediction sample and L1 prediction sample may be generated based on this. For example, the pairwise candidate may be derived based on two other candidates from among the candidates included in the merge candidate list.

[0242] Alternatively, for example, the merge candidate list may be a subblock merge candidate list, and affine merge candidates, subblock merge candidates, or SbTMVP candidates may be selected. Here, affine merge candidates on a subblock basis may also be called subblock merge candidates.

[0243] For example, the candidates may include subblock merge candidates. For example, if the selected candidate is a subblock merge candidate, L0 motion information and L1 motion information may be derived based on the subblock merge candidate, and L0 prediction samples and L1 prediction samples may be generated based on this. For example, the subblock merge candidate may include a CPMV (Control Point Motion Vector), and the L0 prediction samples and L1 prediction samples may be generated by performing predictions on a subblock basis based on the CPMV.

[0244] Here, CPMV can be represented based on one of the surrounding blocks of the current block's CP (Control Point). For example, there can be two, three, or four CPs, and they can be located in at least part of the upper left side (or upper left corner), upper right side (or upper right corner), lower left side (or lower left corner), or lower right side (or lower right corner) of the current block, with only one CP at each location.

[0245] For example, CP could be CP0, located to the upper left of the current block. In this case, the surrounding blocks could include the surrounding block at the upper left corner of the current block, the surrounding block to the left adjacent to the lower side of the surrounding block at the upper left corner, and the surrounding block to the upper side adjacent to the right side of the surrounding block at the upper left corner. Alternatively, the surrounding blocks could include blocks A2, B2, or B3 in Figure 10.

[0246] Alternatively, for example, the CP could be CP1 located to the upper right of the current block. In this case, the surrounding block could include the surrounding block at the upper right corner of the current block and the upper surrounding block adjacent to the left of the surrounding block at the upper right corner. Alternatively, the surrounding block could include block B0 or ​​block B1 in Figure 10.

[0247] Alternatively, for example, CP could be CP2 located to the lower left of the current block. In this case, the surrounding block could include the surrounding block at the lower left corner of the current block and the surrounding block to the left adjacent to the upper side of the surrounding block at the lower left corner. Alternatively, the surrounding block could include block A0 or block A1 in Figure 10.

[0248] Alternatively, for example, the CP could be CP3 located to the lower right of the current block. Here, CP3 may also be called RB. In this case, the surrounding block may include the col block of the current block or the surrounding block at the lower right corner of the col block. Here, the col block may include a block at the same position as the current block in a reference picture different from the current picture in which the current block is located. Alternatively, the surrounding block may include a T block in Figure 10.

[0249] Alternatively, for example, the candidates may include SbTMVP candidates. For example, if the selected candidate is an SbTMVP candidate, L0 motion information and L1 motion information may be derived based on the surrounding blocks to the left of the current block, and the L0 prediction sample and the L1 prediction sample may be generated based on this. For example, the L0 prediction sample and the L1 prediction sample may be generated by performing predictions on a subblock basis.

[0250] For example, L0 motion information may include an L0 reference picture index and an L0 motion vector, and L1 motion information may include an L1 reference picture index and an L1 motion vector. An L0 reference picture index may include information representing a reference picture in reference picture list 0, and an L1 reference picture index may include information representing a reference picture in reference picture list 1.

[0251] For example, an encoding device can generate prediction samples for the current block based on L0 prediction samples, L1 prediction samples, and weighted value information. For example, the weighted value information can be represented based on weighted value index information. The weighted value index information can represent weighted value index information for biprediction. For example, the weighted value information can include information for a weighted average of L0 prediction samples or L1 prediction samples. That is, the weighted value index information can represent index information for the weights used in the weighted average, and the weighted value index information can also be generated in a procedure for generating prediction samples based on the weighted average. For example, the weighted value index information can include information representing one of three or five weights. For example, the weighted average can represent a weighted average in BCW (Bi-prediction with CU-level Weight) or BWA (Bi-prediction with Weighted Average).

[0252] For example, the candidate may include a temporal merge candidate, and the weighted index information for the temporal merge candidate may be represented as 0. That is, the weighted index information for the temporal merge candidate may be represented as 0. Here, weighted index information of 0 can indicate that the weights for each reference direction (i.e., the L0 prediction direction and the L1 prediction direction in biprediction) are the same. Alternatively, for example, the candidate may include a temporal merge candidate, and the weighted index information may be represented based on the weighted index information of a col block. That is, the weighted index information for the temporal merge candidate may be represented based on the weighted index information of a col block. Here, the col block may include a block at the same position as the current block in a reference picture different from the current picture in which the current block is located.

[0253] For example, the candidates may include pairwise candidates, and the weighted index information may be represented based on the weighted index information of one of the other two candidates in the merge candidate list used to derive the pairwise candidates. That is, the weighted index information for the pairwise candidates may be represented based on the weighted index information of one of the other two candidates in the merge candidate list used to derive the pairwise candidates.

[0254] For example, the candidates include pairwise candidates, and the pairwise candidates can be represented based on two other candidates. If the weighted index information of the two other candidates is the same as the first weighted index information, the weighted index information for the pairwise candidates can be represented based on the first weighted index information. If the weighted index information of the two other candidates is not the same as the first weighted index information, the weighted index information for the pairwise candidates can be represented based on default weighted index information. In this case, the default weighted index information can correspond to weighted index information that assigns the same weight to each of the L0 prediction samples and the L1 prediction samples.

[0255] For example, the candidates include pairwise candidates, and the pairwise candidates can be represented based on two other candidates. If the weighted index information of the two other candidates is the same as the first weighted index information, the weighted index information for the pairwise candidates can be represented based on the first weighted index information. If the weighted index information of the two other candidates is not the same as the first weighted index information, the weighted index information for the pairwise candidates can be represented based on the weighted index information of the two other candidates that is not the default weighted index information. The default weighted index information can correspond to weighted index information that assigns the same weight to each of the L0 prediction samples and the L1 prediction samples.

[0256] For example, the merge candidate list may be a subblock merge candidate list, and affine merge candidates, subblock merge candidates, or SbTMVP candidates may be selected. Here, affine merge candidates on a subblock basis may also be called subblock merge candidates.

[0257] For example, the candidates may include affine merge candidates, and the affine merge candidates may include CPMV (Control Point Motion Vector).

[0258] For example, if the affine merge candidate includes CPMV0 for CP0 (Control Point 0) located on the upper left side of the current block, the weighted index information for the affine merge candidate can be shown based on the zeroth weighted index information for CP0. If the affine merge candidate does not include CPMV0 for CP0 located on the upper left side of the current block, the weighted index information for the affine merge candidate can be shown based on the first weighted index information for CP1 (Control Point 1) located on the upper right side of the current block.

[0259] The aforementioned zero-weighted index information corresponds to the weighted index information of the block used for deriving the CPMV0 among the surrounding blocks of CP0, and the surrounding blocks of CP0 may include the surrounding block of the upper left corner of the current block, the surrounding block to the left adjacent to the lower side of the surrounding block of the upper left corner, and the surrounding block to the upper side adjacent to the right side of the surrounding block of the upper left corner.

[0260] The first weighted index information corresponds to the weighted index information of the block used for deriving the CPMV1 among the surrounding blocks of CP1, and the surrounding blocks of CP1 may include the surrounding block of the upper right corner of the current block and the upper surrounding block adjacent to the left of the surrounding block of the upper right corner.

[0261] Alternatively, for example, the candidates may include SbTMVP candidates, and the weighted index information for the SbTMVP candidates may be represented based on the weighted index information of the surrounding blocks to the left of the current block. That is, the weighted index information for the SbTMVP candidates may be represented based on the weighted index information of the surrounding blocks to the left.

[0262] Alternatively, for example, the candidates may include SbTMVP candidates, and the weighted index information for the SbTMVP candidates may be represented as 0. That is, the weighted index information for SbTMVP candidates may be represented as 0. Here, a weighted index information of 0 indicates that the weights for each reference direction (i.e., the L0 prediction direction and the L1 prediction direction in biprediction) are the same.

[0263] Alternatively, for example, the candidates may include SbTMVP candidates, and the weighted index information may be represented based on the weighted index information of the center block within the col block. That is, the weighted index information for SbTMVP candidates may be represented based on the weighted index information of the center block within the col block. Here, the col block may include a block in the same position as the current block in a reference picture different from the current picture in which the current block is located, and the center block may include the lower right subblock of the four subblocks located in the center of the col block.

[0264] Alternatively, for example, the candidates may include SbTMVP candidates, and the weighted index information may be represented based on the weighted index information of each subblock of the col block. That is, the weighted index information for SbTMVP candidates may be represented based on the weighted index information of each subblock of the col block.

[0265] Alternatively, although not shown in Figure 12, for example, the encoding device can derive a residual sample based on the predicted sample and the original sample. In this case, residual-related information can be derived based on the residual sample. A residual sample can be derived based on the residual-related information. A restored sample can be generated based on the residual sample and the predicted sample. A restored block and a restored picture can be derived based on the restored sample. Or, for example, the encoding device can encode video information including residual-related information or predicted-related information.

[0266] For example, an encoding device can encode video information containing all or part of the aforementioned information (or syntax elements) and generate a bitstream or encoded information. Alternatively, it can output it in the form of a bitstream. The bitstream or encoded information can be transmitted to a decoding device via a network or storage medium. Alternatively, the bitstream or encoded information can be stored on a computer-readable storage medium, and the bitstream or encoded information can be generated by the video encoding method described above.

[0267] Figures 14 and 15 schematically show an example of a video / image decoding method and related components according to the embodiment of this document.

[0268] The method disclosed in Figure 14 can be performed by the decoding device disclosed in Figure 3 or Figure 15. Specifically, for example, S1400 in Figure 14 can be performed by the entropy decoding unit 310 of the decoding device 300 in Figure 15, and S1410 to S1440 in Figure 14 can be performed by the prediction unit 330 of the decoding device 300 in Figure 15. Although not shown in Figure 14, in Figure 15, the entropy decoding unit 310 of the decoding device 300 can derive prediction-related information or residual information from the bitstream, the residual processing unit 320 of the decoding device 300 can derive residual samples from the residual information, the prediction unit 330 of the decoding device 300 can derive prediction samples from the prediction-related information, and the addition unit 340 of the decoding device 300 can derive a restored block or restored picture from the residual sample or predicted sample. The method disclosed in Figure 14 may include embodiments described above in this document.

[0269] Referring to Figure 14, the decoding device can receive video information including interpredictive mode information via a bitstream (S1400). For example, the video information may also be called video information. The video information may include various types of information according to the embodiments described above in this document. For example, the video information may include at least a portion of prediction-related information or residual-related information.

[0270] For example, the prediction-related information may include interpretation mode information or interpretation type information. For example, the interpretation mode information may include information indicating at least some of the various interpretation modes. For example, various modes such as merge mode, skip mode, MVP (motion vector prediction) mode, affine mode, subblock merge mode, or MMVD (merge with MVD) mode can be used. In addition, DMVR (Decoder side motion vector refinement) mode, AMVR (adaptive motion vector resolution) mode, BCW (Bi-prediction with CU-level weight), or BDOF (Bi-directional optical flow) may be used as additional or alternative modes. For example, the interpretation type information may include the inter_pred_idc syntax element. Alternatively, the interpretation type information may include information indicating either L0 prediction, L1 prediction, or bi-prediction.

[0271] The decoding device can generate a merge candidate list for the current block based on the inter-prediction mode information (S1410). For example, based on the inter-prediction mode information, the decoding device can determine the inter-prediction mode of the current block to be a merge mode, an affine (merge) mode, or a sub-block merge mode, and generate a merge candidate list based on the determined inter-prediction mode. Here, if the inter-prediction mode is determined to be an affine merge mode or a sub-block merge mode, the merge candidate list may be called an affine merge candidate list or a sub-block merge candidate list, etc., but may also be simply called a merge candidate list.

[0272] For example, candidates may be inserted into the merge candidate list until the number of candidates in the merge candidate list reaches the maximum number of candidates. Here, a candidate can represent a candidate or candidate block for deriving motion information (or motion vector) of the current block. For example, a candidate block can be derived by searching for surrounding blocks of the current block. For example, surrounding blocks may include spatially surrounding blocks and / or temporally surrounding blocks of the current block, and spatially surrounding blocks may be searched first to derive (spatial merge) candidates, and then temporally surrounding blocks may be searched to derive (temporal merge) candidates, and the derived candidates can be inserted into the merge candidate list. For example, even after the candidates have been inserted, if the number of candidates in the merge candidate list is less than the maximum number of candidates, additional candidates may be inserted. For example, additional candidates may include at least one of the following: history-based merge candidate(s), pair-wise average merge candidate(s), ATMVP, combined bi-predictive merge candidate(s) (if the slice / tile group type of the current slice / tile group is type B), and / or zero-vector merge candidate(s).

[0273] Alternatively, for example, candidates may be inserted into the affine merge candidate list until the number of candidates in the affine merge candidate list reaches the maximum number of candidates. Here, a candidate may include the CPMV (Control Point Motion Vector) of the current block. Alternatively, the candidate may represent a candidate or candidate block for deriving the CPMV. The CPMV may represent the motion vector at the CP (Control Point) of the current block. For example, there may be two, three, or four CPs, and they may be located on at least part of the upper left side (or upper left corner), upper right side (or upper right corner), lower left side (or lower left corner), or lower right side (or lower right corner) of the current block, with only one CP at each position.

[0274] For example, candidate blocks can be derived by searching the surrounding blocks of the current block (or the surrounding blocks of the current block's CP). For example, an affine merge candidate list can contain at least one of inherited affine merge candidates, constructed affine merge candidates, or zero motion vector candidates. For example, an affine merge candidate list can first insert the inherited affine merge candidates, and then insert the constructed affine merge candidates. Also, if the number of candidates in the affine merge candidate list is less than the maximum number of candidates after inserting the constructed affine merge candidates, the remainder can be filled with zero motion vector candidates. Here, zero motion vector candidates can also be called zero vectors. For example, an affine merge candidate list can be a list in an affine merge mode in which motion vectors are derived on a sample-by-sample basis, or it can be a list in an affine merge mode in which motion vectors are derived on a subblock basis. In this case, the affine merge candidate list can also be called the subblock merge candidate list, and the subblock merge candidate list can also include candidates derived by SbTMVP (or SbTMVP candidates). For example, if an SbTMVP candidate is included in the subblock merge candidate list, it can be positioned before inherited affine merge candidates and constructed affine merge candidates within the subblock merge candidate list.

[0275] The decoding device can derive motion information of the current block based on a selected candidate from the merge candidate list (S1420). For example, the merge candidate list may include at least some of spatial merge candidates, temporal merge candidates, pairwise candidates, or zero vector candidates, and one of these candidates can be selected for inter prediction of the current block. Alternatively, for example, the subblock merge candidate list may include at least some of inherited affine merge candidates, configured affine merge candidates, SbTMVP candidates, or zero vector candidates, and one of these candidates can be selected for inter prediction of the current block. For example, the selected candidate can be selected from the merge candidate list based on selection information. For example, the selection information may include index information indicating the selected candidate in the merge candidate list. For example, the selection information may also be called merge index information or subblock merge index information. For example, the selection information may be included in the video information. Alternatively, the selection information may be included in the inter prediction mode information.

[0276] The decoding device can generate L0 and L1 prediction samples for the current block based on motion information (S1430). For example, if the inter-prediction type is derived as bi-prediction, the decoding device can derive L0 motion information and L1 motion information based on the selected candidate. The decoding device can derive the inter-prediction type of the current block as bi-prediction based on the inter-prediction type information. For example, the inter-prediction type of the current block can be derived as bi-prediction from among L0 prediction, L1 prediction, or bi-prediction based on the inter-prediction type information. Here, L0 prediction can represent a prediction based on reference picture list 0, L1 prediction can represent a prediction based on reference picture list 1, and bi-prediction can represent a prediction based on both reference picture list 0 and reference picture list 1. For example, the inter-prediction type information may include the inter_pred_idc syntax element.

[0277] For example, L0 motion information may include an L0 reference picture index and an L0 motion vector, and L1 motion information may include an L1 reference picture index and an L1 motion vector. The L0 reference picture index may include information indicating a reference picture in reference picture list 0, and the L1 reference picture index may include information indicating a reference picture in reference picture list 1.

[0278] For example, the candidates may include spatial merge candidates. For example, if the selected candidate is a spatial merge candidate, L0 motion information and L1 motion information may be derived based on the spatial merge candidate, and the L0 prediction sample and the L1 prediction sample may be generated based on this.

[0279] For example, the candidates may include temporal merge candidates. For example, if the selected candidate is a temporal merge candidate, L0 motion information and L1 motion information may be derived based on the temporal merge candidate, and the L0 prediction sample and the L1 prediction sample may be generated based on this.

[0280] For example, the candidates may include pairwise candidates. For example, if the selected candidate is a pairwise candidate, L0 motion information and L1 motion information may be derived based on the pairwise candidate, and the L0 prediction sample and L1 prediction sample may be generated based on this. For example, the pairwise candidate may be derived based on two other candidates from among the candidates included in the merge candidate list.

[0281] Alternatively, for example, the merge candidate list may be a subblock merge candidate list, and affine merge candidates, subblock merge candidates, or SbTMVP candidates may be selected. Here, affine merge candidates on a subblock basis may also be called subblock merge candidates.

[0282] For example, the candidates may include affine merge candidates. For example, if the selected candidate is an affine merge candidate, L0 motion information and L1 motion information may be derived based on the affine merge candidate, and L0 prediction samples and L1 prediction samples may be generated based on this. For example, the affine merge candidate may include a CPMV (Control Point Motion Vector), and the L0 prediction samples and L1 prediction samples may be generated by performing predictions on a subblock basis based on the CPMV.

[0283] Here, CPMV can be derived based on one of the surrounding blocks of the current block's CP (Control Point). For example, there can be two, three, or four CPs, and they can be located in at least part of the upper left side (or upper left corner), upper right side (or upper right corner), lower left side (or lower left corner), or lower right side (or lower right corner) of the current block, with only one CP at each location.

[0284] For example, CP could be CP0, located to the upper left of the current block. In this case, the surrounding blocks could include the surrounding block at the upper left corner of the current block, the surrounding block to the left adjacent to the lower side of the surrounding block at the upper left corner, and the surrounding block to the upper side adjacent to the right side of the surrounding block at the upper left corner. Alternatively, the surrounding blocks could include blocks A2, B2, or B3 in Figure 10.

[0285] Alternatively, for example, the CP could be CP1 located to the upper right of the current block. In this case, the surrounding block could include the surrounding block at the upper right corner of the current block and the upper surrounding block adjacent to the left of the surrounding block at the upper right corner. Alternatively, the surrounding block could include block B0 or ​​block B1 in Figure 10.

[0286] Alternatively, for example, CP could be CP2 located to the lower left of the current block. In this case, the surrounding block could include the surrounding block at the lower left corner of the current block and the surrounding block to the left adjacent to the upper side of the surrounding block at the lower left corner. Alternatively, the surrounding block could include block A0 or block A1 in Figure 10.

[0287] Alternatively, for example, the CP could be CP3 located to the lower right of the current block. Here, CP3 may also be called RB. In this case, the surrounding block may include the col block of the current block or the surrounding block at the lower right corner of the col block. Here, the col block may include a block at the same position as the current block in a reference picture different from the current picture in which the current block is located. Alternatively, the surrounding block may include a T block in Figure 10.

[0288] Alternatively, for example, the candidates may include SbTMVP candidates. For example, if the selected candidate is an SbTMVP candidate, L0 motion information and L1 motion information can be derived based on the surrounding blocks to the left of the current block, and based on this, the L0 prediction sample and the L1 prediction sample can be generated. For example, the L0 prediction sample and the L1 prediction sample can be generated by performing predictions on a subblock basis.

[0289] The decoding device can generate prediction samples for the current block based on L0 prediction samples, L1 prediction samples, and weighted value information (S1440). For example, the weighted value information can be derived based on weighted value index information for a selected candidate from among the candidates included in the merge candidate list. For example, the weighted value information can include information for a weighted average of L0 prediction samples or L1 prediction samples. That is, the weighted value index information can indicate index information for the weights used in the weighted average, and the weighted average can be performed based on the weighted value index information. For example, the weighted value index information can include information indicating the weights of any three or five weights. For example, the weighted average can indicate a weighted average using BCW (Bi-prediction with CU-level Weight) or BWA (Bi-prediction with Weighted Average).

[0290] For example, the candidates may include temporal merge candidates, and the weighted index information for the temporal merge candidates may be derived to 0. That is, the weighted index information for temporal merge candidates may be derived to 0. Here, a weighted index information of 0 can indicate that the weights for each reference direction (i.e., the L0 prediction direction and the L1 prediction direction in biprediction) are the same.

[0291] For example, the candidates may include temporal merge candidates, and the weighted index information for the temporal merge candidates can be derived based on the weighted index information of the col block. That is, the weighted index information for the temporal merge candidates can be derived based on the weighted index information of the col block. Here, the col block may include a block at the same position as the current block in a reference picture different from the current picture in which the current block is located.

[0292] For example, the candidates may include pairwise candidates, and the weighted index information may be derived from the weighted index information of one of the other two candidates in the merge candidate list used to derive the pairwise candidates. That is, the weighted index information for the pairwise candidates may be derived from the weighted index information of one of the other two candidates in the merge candidate list used to derive the pairwise candidates.

[0293] For example, the candidates include pairwise candidates, and the pairwise candidates can be derived based on two other candidates. If the weighted index information of the two other candidates is the same as the first weighted index information, the weighted index information for the pairwise candidates can be derived based on the first weighted index information. If the weighted index information of the two other candidates is not the same, the weighted index information for the pairwise candidates can be derived based on default weighted index information. In this case, the default weighted index information can correspond to weighted index information that assigns the same weight to each of the L0 prediction samples and the L1 prediction samples.

[0294] For example, the candidates include pairwise candidates, and the pairwise candidates can be derived based on two other candidates. If the weighted index information of the two other candidates is the same as the first weighted index information, the weighted index information for the pairwise candidates can be derived based on the first weighted index information. If the weighted index information of the two other candidates is not the same as the first weighted index information, the weighted index information for the pairwise candidates can be derived based on the weighted index information of the two other candidates that is not the default weighted index information. The default weighted index information may correspond to weighted index information that assigns the same weight to each of the L0 prediction samples and the L1 prediction samples.

[0295] For example, the merge candidate list may be a subblock merge candidate list, and affine merge candidates, subblock merge candidates, or SbTMVP candidates may be selected. Here, affine merge candidates on a subblock basis may also be called subblock merge candidates.

[0296] For example, the candidates may include affine merge candidates, and the affine merge candidates may include CPMV (Control Point Motion Vector).

[0297] For example, if the affine merge candidate includes CPMV0 for CP0 (Control Point 0) located on the upper left side of the current block, the weighted index information for the affine merge candidate can be derived based on the zeroth weighted index information for CP0. If the affine merge candidate does not include CPMV0 for CP0 located on the upper left side of the current block, the weighted index information for the affine merge candidate can be derived based on the first weighted index information for CP1 (Control Point 1) located on the upper right side of the current block.

[0298] The zero weighted index information corresponds to the weighted index information of the block used for deriving CPMV0 among the surrounding blocks of CP0, and the surrounding blocks of CP0 may include the surrounding block at the upper left corner of the current block, the surrounding block to the left adjacent to the lower side of the surrounding block at the upper left corner, and the surrounding block above adjacent to the right side of the surrounding block at the upper left corner.

[0299] The first weighted index information corresponds to the weighted index information of the block used for deriving the CPMV1 among the surrounding blocks of CP1, and the surrounding blocks of CP1 may include the surrounding block of the upper right corner of the current block and the upper surrounding block adjacent to the left of the surrounding block of the upper right corner.

[0300] Alternatively, for example, the candidates may include SbTMVP candidates, and the weighted index information for the SbTMVP candidates can be derived based on the weighted index information of the surrounding blocks to the left of the current block. That is, the weighted index information for the SbTMVP candidates can be derived based on the weighted index information of the surrounding blocks to the left.

[0301] Alternatively, for example, the candidates may include SbTMVP candidates, and the weighted index information for the SbTMVP candidates may be derived as 0. That is, the weighted index information for SbTMVP candidates may be derived as 0. Here, a weighted index information of 0 indicates that the weights for each reference direction (i.e., the L0 prediction direction and the L1 prediction direction in biprediction) are the same.

[0302] Alternatively, for example, the candidates may include SbTMVP candidates, and the weighted index information may be derived based on the weighted index information of the center block within the col block. That is, the weighted index information for SbTMVP candidates may be derived based on the weighted index information of the center block within the col block. Here, the col block may include a block in the same position as the current block in a reference picture different from the current picture in which the current block is located, and the center block may include the lower right subblock of the four subblocks located in the center of the col block.

[0303] Alternatively, for example, the candidates may include SbTMVP candidates, and the weighted index information can be derived based on the weighted index information of each subblock of the col block. That is, the weighted index information for SbTMVP candidates can be derived based on the weighted index information of each subblock of the col block.

[0304] Although not shown in Figure 14, for example, the decoding device can derive a residual sample based on residual-related information contained in the video information. Furthermore, the decoding device can generate a restored sample based on the predicted sample and the residual sample. Based on the restored sample, a restored block and a restored picture can be derived.

[0305] For example, a decoding device can decode a bitstream or encoded information to obtain video information containing all or part of the aforementioned information (or syntax elements). Furthermore, the bitstream or encoded information can be stored on a computer-readable storage medium, and the aforementioned decoding method can be executed.

[0306] In the embodiments described above, the method is explained based on a flowchart in a series of steps or blocks, but the embodiments are not limited to the order of the steps, and some steps may occur with other steps, in a different order, or simultaneously. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments described herein.

[0307] The methods according to the embodiments of this document described above can be implemented in software form, and the encoding and / or decoding devices according to this document can be included in devices that perform video processing, such as TVs, computers, smartphones, set-top boxes, and display devices.

[0308] In this document, when embodiments are implemented in software, the methods described above can be implemented by modules (processes, functions, etc.) that perform the functions described above. These modules are stored in memory and can be executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by a variety of well-known means. The processor may include an ASIC (application-specific integrated circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.

[0309] Furthermore, the decoding and encoding devices to which the embodiments of this document apply can include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providers, over-the-top (OTT) video equipment, internet streaming service providers, 3D video equipment, virtual reality (VR) equipment, argumentative reality (AR) equipment, image-phone video equipment, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video equipment, and can be used to process video signals or data signals. For example, over-the-top (OTT) video equipment can include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).

[0310] Furthermore, the processing methods to which the embodiments of this document apply can be produced in the form of programs executed on a computer and stored on a computer-readable recording medium. Multimedia data having the data structure according to the embodiments of this document can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store data that can be read by a computer. The computer-readable recording medium can include, for example, Blu-ray discs (BDs), general-purpose serial buses (USB), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media implemented in the form of carrier waves (e.g., transmission over the Internet). Furthermore, a bitstream generated by an encoding method can be stored on a computer-readable recording medium or transmitted over a wireless network.

[0311] Furthermore, the embodiments described in this document can be implemented as a computer program product using program code, and the program code can be executed on a computer according to the embodiments described in this document. The program code can be stored on a computer-readable carrier.

[0312] Figure 16 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.

[0313] Referring to Figure 16, the content streaming system to which the embodiments described in this document apply can broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.

[0314] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then transmitting this bitstream to the streaming server. In other cases, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server can be omitted.

[0315] The bitstream can be generated by an encoding method or bitstream generation method applicable to the embodiments of this document, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0316] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server acts as an intermediary to inform users about available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.

[0317] The streaming server can receive content from a media storage and / or encoding server. For example, if it starts receiving content from the encoding server, it can receive the content in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0318] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, and digital signage.

[0319] Each server within the aforementioned content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.

[0320] The claims described herein can be combined in various ways. For example, the technical features of the method claims herein can be combined and implemented in an apparatus, and the technical features of the apparatus claims herein can be combined and implemented in a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims herein can be combined and implemented in an apparatus, and the technical features of the method claims and the technical features of the apparatus claims herein can be combined and implemented in a method.

Claims

1. In a video decoding method performed by a decoding device, A step of acquiring video information including interprediction mode information via a bitstream, The steps include generating a list of merge candidates for the current block based on the aforementioned inter-prediction mode information, The steps include: deriving movement information for the current block based on a candidate selected from among the candidates in the merge candidate list; The steps include generating L0 and L1 predicted samples for the current block based on the aforementioned motion information, A step of generating a prediction sample for the current block based on the L0 prediction sample, the L1 prediction sample, and the weighted index for the current block, wherein the weighted index for the current block is derived based on the weighted index for the selected candidate, The aforementioned candidates include the constructed affine merge candidates, The configured affine merge candidate is constructed based on three CPMVs (Control Point Motion Vectors) for the current block and a weighted index for the configured affine merge candidate. The three CPMVs of the configured affine merge candidate are assigned from among four CPMVs, including CPMV(0), CPMV(1), CPMV(2), and CPMV(3). The CPMV(0) relates to the surrounding block associated with the upper left corner of the current block, The CPMV(1) relates to the surrounding block associated with the upper right corner of the current block, The CPMV(2) relates to the surrounding block associated with the lower left corner of the current block, The CPMV(3) relates to the surrounding block associated with the lower right corner of the current block, The weighted index for the configured affine merge candidate represents one of the five weighted values. The weighted index for the configured affine merge candidates is: Based on the three CPMVs of the configured affine merge candidate, including the CPMV(0), the weighted index for the configured affine merge candidate is determined as a weighted index associated with the CPMV(0). A method in which, based on the three CPMVs of the configured affine merge candidate, including CPMV(1), CPMV(2), and CPMV(3), the weighted index for the configured affine merge candidate is determined as a weighted index associated with CPMV(1), such that the weighted index for the CPMV(i) among the three CPMVs of the configured affine merge candidate has the smallest i value.

2. In a video encoding method performed by an encoding device, The steps include determining the inter-prediction mode of the current block and generating inter-prediction mode information indicating the inter-prediction mode, The steps include generating a list of merge candidates for the current block based on the inter prediction mode, The steps include generating selection information that indicates one of the candidates included in the merge candidate list, The step includes encoding video information including the interprediction mode information and the selection information, The aforementioned candidates include the constructed affine merge candidates, The configured affine merge candidate is constructed based on three CPMVs (Control Point Motion Vectors) for the current block and a weighted index for the configured affine merge candidate. The three CPMVs of the configured affine merge candidate are assigned from among four CPMVs, including CPMV(0), CPMV(1), CPMV(2), and CPMV(3). The CPMV(0) relates to the surrounding block associated with the upper left corner of the current block, The CPMV(1) relates to the surrounding block associated with the upper right corner of the current block, The CPMV(2) relates to the surrounding block associated with the lower left corner of the current block, The CPMV(3) relates to the surrounding block associated with the lower right corner of the current block, The weighted index for the configured affine merge candidate represents one of the five weighted values. The weighted index for the configured affine merge candidates is: Based on the three CPMVs of the configured affine merge candidate, including the CPMV(0), the weighted index for the configured affine merge candidate is determined as a weighted index associated with the CPMV(0). A method in which, based on the three CPMVs of the configured affine merge candidate, including CPMV(1), CPMV(2), and CPMV(3), the weighted index for the configured affine merge candidate is determined as a weighted index associated with CPMV(1), such that the weighted index for the CPMV(i) among the three CPMVs of the configured affine merge candidate has the smallest i value.

3. A method for transmitting video data, A step of obtaining a bitstream relating to the video, wherein the bitstream is The steps include determining the inter-prediction mode of the current block and generating inter-prediction mode information indicating the inter-prediction mode, The steps include generating a list of merge candidates for the current block based on the inter prediction mode, The steps include generating selection information that indicates one of the candidates included in the merge candidate list, A step of encoding video information including the interprediction mode information and the selection information, and a step of generating based on the above, The step of transmitting the data, which includes the bitstream, The aforementioned candidates include the constructed affine merge candidates, The configured affine merge candidate is constructed based on three CPMVs (Control Point Motion Vectors) for the current block and a weighted index for the configured affine merge candidate. The three CPMVs of the configured affine merge candidate are assigned from among four CPMVs, including CPMV(0), CPMV(1), CPMV(2), and CPMV(3). The CPMV(0) relates to the surrounding block associated with the upper left corner of the current block, The CPMV(1) relates to the surrounding block associated with the upper right corner of the current block, The CPMV(2) relates to the surrounding block associated with the lower left corner of the current block, The CPMV(3) relates to the surrounding block associated with the lower right corner of the current block, The weighted index for the configured affine merge candidate represents one of the five weighted values. The weighted index for the configured affine merge candidates is: Based on the three CPMVs of the configured affine merge candidate, including the CPMV(0), the weighted index for the configured affine merge candidate is determined as a weighted index associated with the CPMV(0). A method in which, based on the three CPMVs of the configured affine merge candidate, including CPMV(1), CPMV(2), and CPMV(3), the weighted index for the configured affine merge candidate is determined as a weighted index associated with CPMV(1), such that the weighted index for the CPMV(i) among the three CPMVs of the configured affine merge candidate has the smallest i value.