Video or image coding to induce weighted index information for biprediction.

By using weighted index information for biprediction in video coding, the method addresses inefficiencies in high-resolution video compression, enhancing efficiency and reducing costs for high-quality video transmission and storage.

JP7835793B2Active Publication Date: 2026-03-25LG ELECTRONICS INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-05-02
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality video, particularly in immersive media formats like VR and AR, has led to a surge in video data volume, resulting in higher transmission and storage costs due to inefficient compression methods.

Method used

Implementing weighted index information for biprediction during video coding, specifically through dual prediction and deriving weighted index information for merge candidate lists, to enhance video compression efficiency.

Benefits of technology

This approach improves the overall efficiency of video compression by efficiently configuring candidate motion vectors and performing weighted-value-based dual prediction, reducing transmission and storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007835793000033
    Figure 0007835793000033
  • Figure 0007835793000034
    Figure 0007835793000034
  • Figure 0007835793000035
    Figure 0007835793000035
Patent Text Reader

Abstract

To provide video or image coding that derives weighted index information for bi-prediction.SOLUTION: According to the disclosure of the present document, when the inter prediction type of a current block indicates bi-prediction, weight index information for candidates in a merge candidate list or a sub-block merge candidate list can be induced or derived, and coding efficiency can be increased.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This technology relates to video or image coding that derives weighted index information for biprediction. [Background technology]

[0002] In recent years, demand for high-resolution, high-quality video, such as 4K or 8K or higher UHD (Ultra High Definition) video, has been increasing in various fields. As video data becomes higher resolution and higher quality, the amount of information or bits transmitted increases relative to existing video data. Therefore, when transmitting video data using existing wired or wireless broadband lines, or storing video data using existing storage media, transmission and storage costs increase.

[0003] Furthermore, in recent years, interest in and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality), and holograms have increased, and there has been a rise in broadcasting of video content with different visual characteristics from real-world footage, such as game footage.

[0004] Therefore, highly efficient video compression technology is required to effectively compress, transmit, store, and play back high-resolution, high-quality video information that possesses the various characteristics described above. [Overview of the project] [Means for solving the problem]

[0005] According to one embodiment of this document, a method and apparatus for improving the efficiency of video coding are provided.

[0006] According to one embodiment of this document, a method and apparatus for using weighted value-based dual prediction during video coding are provided.

[0007] According to one embodiment of this document, a method and apparatus for deriving weighted index information for biprediction in interprediction is provided.

[0008] According to one embodiment of this document, a method and apparatus are provided for deriving weighted index information for candidates in a merge candidate list or an affine merge candidate list during biprediction.

[0009] According to one embodiment of this document, a video / image decoding method performed by a decoding device is provided.

[0010] According to one embodiment of this document, a decoding device for performing video / image decoding is provided.

[0011] According to one embodiment of this document, a video / image encoding method performed by an encoding device is provided.

[0012] According to one embodiment of this document, an encoding device for performing video / image encoding is provided.

[0013] According to one embodiment of this document, a computer-readable digital storage medium is provided which stores encoded video / image information generated by a video / image encoding method disclosed in at least one embodiment of this document.

[0014] According to one embodiment of this document, a computer-readable digital storage medium is provided which stores encoded information or encoded video / image information, which is triggered by a decoding device to perform the video / image decoding method disclosed in at least one embodiment of this document. [Effects of the Invention]

[0015] According to one embodiment of this document, the overall efficiency of video compression can be improved.

[0016] According to one embodiment of this document, candidate motion vectors can be efficiently configured during inter prediction.

[0017] According to one embodiment of this document, weighted-value-based dual prediction can be efficiently performed.

Brief Description of Drawings

[0018] [Figure 1] An example of a video / video coding system to which the embodiments of this document can be applied is schematically shown. [Figure 2] It is a diagram schematically explaining the configuration of a video / video encoding apparatus to which the embodiments of this document can be applied. [Figure 3] It is a diagram schematically explaining the configuration of a video / video decoding apparatus to which the embodiments of this document can be applied. [Figure 4] It is a diagram for explaining the merge mode in inter prediction. [Figure 5A] CPMV for affine motion prediction is exemplarily shown. [Figure 5B] CPMV for affine motion prediction is exemplarily shown. [Figure 6] A case where the affine MVF is determined in units of sub-blocks is exemplarily shown. [Figure 7] It is a diagram for explaining the affine merge mode in inter prediction. [Figure 8] It is a diagram for explaining the position of candidates in the affine merge mode. [Figure 9] It is a diagram for explaining SbTMVP in inter prediction. [Figure 10] An example of a video / video encoding method according to the embodiments of this document and related components is schematically shown. [Figure 11] An example of a video / video encoding method according to the embodiments of this document and related components is schematically shown. [Figure 12]This document outlines an example of a video decoding method and related components relating to the embodiments described herein. [Figure 13] This document outlines an example of a video decoding method and related components relating to the embodiments described herein. [Figure 14] Examples of content streaming systems to which the embodiments disclosed in this document can be applied are shown below. [Modes for carrying out the invention]

[0019] The disclosures in this document can be modified in various ways and may have various embodiments, but specific embodiments are illustrated in the drawings and described in detail. However, this does not mean that the disclosure is limited to any particular embodiment. The terms used in this document are used solely to describe specific embodiments and are not intended to limit the technical ideas of the embodiments described herein. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this document, terms such as “includes” or “has” are intended to indicate the existence of features, figures, stages, operations, components, parts, or combinations thereof described in the document, and should be understood not to preemptively exclude the possibility of the existence or addition of one or more different features, figures, stages, operations, components, parts, or combinations thereof.

[0020] On the other hand, each configuration shown in the drawings described in this document is shown independently for the convenience of explaining its distinct characteristic functions, and does not mean that each configuration is embodied in separate hardware or separate software. For example, two or more configurations may be combined to form a single configuration, and one configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included within the scope of disclosure in this document.

[0021] The embodiments described in this document will be explained below with reference to the attached diagrams. The same reference numerals may be used for the same components in the diagrams, and redundant explanations for the same components may be omitted.

[0022] Figure 1 schematically shows an example of a video / image coding system to which the embodiments described in this document can be applied.

[0023] As shown in Figure 1, a video / image coding system may comprise a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or network.

[0024] The source device may comprise a video source, an encoding device, and a transmitter. The receiving device may comprise a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be provided in the encoding device. The receiver may be provided in the decoding device. The renderer may comprise a display unit, which may consist of a separate device or external component.

[0025] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, or video / image archives containing previously captured video / images. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process may be replaced by the process of generating the associated data.

[0026] An encoding device can encode input video / image data. For compression and coding efficiency, the encoding device can perform a series of steps, including prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in bitstream format.

[0027] The transmitting unit can transmit encoded video / image information or data output in bitstream format to the receiving unit of a receiving device via a digital storage medium or network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit may include elements for generating media files via a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0028] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of an encoding device.

[0029] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0030] This document relates to video / image coding. For example, the methods / examples disclosed in this document can be applied to methods disclosed in the VVC (versatile video coding) standard. Furthermore, the methods / examples disclosed in this document can be applied to methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (e.g., 267 or H.268).

[0031] This document presents various examples of video / image coding, and unless otherwise noted, these examples may be combined with each other.

[0032] In this document, "video" can mean a collection of images over time. "Picture" generally refers to a unit representing a single image at a specific time point in time, and "slice" or "tile" is a unit that constitutes part of a picture in coding. A slice or tile can contain one or more coding tree units (CTUs). A single picture can consist of one or more slices or tiles. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture.A tile scan may represent a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that may be exclusively contained in a single NAL unit.

[0033] On the other hand, a single picture can be divided into two or more subpictures. A subpicture can be a rectangular region of one or more slices within a picture.

[0034] A pixel or pel can refer to the smallest unit that makes up a picture (or image). Alternatively, the term "sample" may be used as a counterpart to pixel. A sample can generally represent a pixel or a pixel value, and can represent only the luma component pixel / pixel value, or only the chroma component pixel / pixel value.

[0035] A unit can represent a basic unit of image processing. A unit can contain at least one of a specific region of a picture and information associated with that region. A unit can contain one luma block and two chroma (e.g., cb, cr) blocks. The term unit may sometimes be used interchangeably with terms such as block or area. In general, an M×N block can contain a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0036] In this document, "A or B" may mean "just A," "just B," or "both A and B." In other words, in this document, "A or B" may be interpreted as "A and / or B." For example, in this document, "A, B or C" may mean "just A," "just B," "just C," or "any combination of A, B and C."

[0037] The slashes ( / ) and commas used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "just A", "just B", or "both A and B". For example, "A, B, C" can mean "A, B or C".

[0038] In this document, "at least one of A and B" may mean "just A," "just B," or "both A and B." Furthermore, in this document, the expressions "at least one of A or B" and "at least one of A and / or B" may be interpreted similarly to "at least one of A and B."

[0039] Furthermore, in this document, "at least one of A, B and C" may mean "just A," "just B," "just C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."

[0040] Furthermore, parentheses used in this document may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction." In other words, "prediction" in this document is not limited to "intra prediction," and "intra prediction" may be proposed as an example of "prediction." Also, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction."

[0041] Technical features described individually within a single drawing in this document may be embodied individually or simultaneously.

[0042] Figure 2 is a schematic diagram illustrating the configuration of a video / image encoding device to which the embodiments described in this document can be applied. Hereinafter, the term "encoding device" may include an image encoding device and / or a video encoding device.

[0043] As shown in Figure 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be called a reconstructor or a reconstructed block generator. The aforementioned video splitting unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 can be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. The memory 270 may also include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.

[0044] The video splitting unit 210 can split the input video (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units can be recursively split from a coding tree unit (CTU) or the largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary-tree structure and / or the ternary structure. Alternatively, the binary-tree structure may be applied first. The coding procedure according to this disclosure may be performed based on the final coding unit that is not further split. In this case, based on coding efficiency due to video characteristics, the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further comprise a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can each be separated or partitioned from the final coding unit described above.The prediction unit may be a unit of sample prediction, and the conversion unit may be a unit for deriving conversion coefficients and / or a unit for deriving a residual signal from conversion coefficients.

[0045] The term "unit" can sometimes be used interchangeably with terms such as "block" or "area." Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the luminance (luma) component pixel / pixel value, or only the chroma component pixel / pixel value. A sample can be used as the term corresponding to a single picture (or image) pixel or pel.

[0046] The encoding device 200 can generate a residual signal (residual block, residual sample array) by subtracting the predicted signal (predicted block, predicted sample array) output from the inter-prediction unit 221 or intra-prediction unit 222 from the input video signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the predicted signal (predicted block, predicted sample array) from the input video signal (original block, original sample array) within the encoder 200 can be called the subtraction unit 231. The prediction unit can make predictions for the block to be processed (hereinafter referred to as the current block) and generate a predicted block that includes predicted samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied on a current block or CU basis. The prediction unit can generate various information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each prediction mode. Prediction information can be encoded by the entropy encoding unit 240 and output in bitstream format.

[0047] The intra-prediction unit 222 can predict the current block by referring to a sample in the current picture. The referenced sample can be located in the vicinity (neighbor) of the current block or at a distance, depending on the prediction mode. The prediction mode in intra-prediction can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is illustrative, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 222 can also determine the prediction mode to apply to the current block using the prediction modes applied to the surrounding blocks.

[0048] The interprediction unit 221 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the surrounding block and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, col CU, etc., and the reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the interpretation unit 221 can construct a motion information candidate list based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 221 can use the motion information of surrounding blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0049] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for predictions on a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also base its predictions on an intra-block copy (IBC) prediction mode or a palette mode for predictions on a block. The IBC prediction mode or palette mode can be used for content video / video coding such as in games, for example, in SCC (screen content coding). IBC basically performs predictions within the current picture, but can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can use at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, sample values ​​within the picture can be signaled based on information about the palette table and palette index.

[0050] The prediction signal generated via the prediction unit (including the inter-prediction unit 221 and / or the intra-prediction unit 222) can be used to generate a reconstructed signal or a residual signal. The transformation unit 232 can apply a transformation technique to the residual signal to generate transformation coefficients. For example, the transformation technique may include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph when relational information between pixels is represented by this graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels and based on that. The transformation process may also be applied to pixel blocks of the same size that are square, or to non-square blocks of variable size.

[0051] The quantization unit 233 quantizes the conversion coefficients and transmits them to the entropy encoding unit 240, which can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients can be called residual information. The quantization unit 233 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients. The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy encoding unit 240 can also encode information necessary for video / image restoration (e.g., the values ​​of syntax elements) together with or separately from the quantized conversion coefficients. Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. In this document, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information.The video / image information can be encoded via the encoding procedure described above and included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network can include broadcast networks and / or communication networks, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be transmitted by a transmitting unit (not shown) and / or stored by a storage unit (not shown) which are configured as internal / external elements of the encoding device 200, or the transmitting unit may be included in the entropy encoding unit 240.

[0052] The quantized conversion coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 234 and the inverse transformation unit 235. The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 250 can be called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.

[0053] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture encoding and / or restoration process.

[0054] The filtering unit 260 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 260 can generate various filtering-related information and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each filtering method. The filtering-related information can be encoded by the entropy encoding unit 240 and output in bitstream format.

[0055] The corrected restored picture sent to memory 270 can be used as a reference picture in the interpretation unit 221. When interpretation is applied via this, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve encoding efficiency.

[0056] The DPB in memory 270 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 221. Memory 270 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 221 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 270 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 222.

[0057] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which the embodiments described in this document can be applied. Hereinafter, the term "decoding device" may include an image decoding device and / or a video decoding device.

[0058] As shown in Figure 3, the decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. The entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering unit 350 described above can be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. The memory 360 may include a decoded picture buffer (DPB) and may also be configured by a digital storage medium. The aforementioned hardware component may also further include memory 360 as an internal / external component.

[0059] When a bitstream containing video / image information is input, the decoding device 300 can reconstruct the image in accordance with the process by which the video / image information was processed in the encoding device shown in Figure 3. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the decoding processing unit can be, for example, a coding unit, which can be divided from a coding tree unit or a maximum coding unit according to a quad-tree structure, a binary tree structure, and / or a terminally tree structure. One or more conversion units can be derived from the coding unit. The reconstructed video signal decoded and output via the decoding device 300 can then be played back via a playback device.

[0060] The decoding device 300 can receive the signal output from the encoding device shown in Figure 3 in bitstream form, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information necessary for video restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can further decode the picture based on the parameter set information and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements necessary for image restoration and the quantized values ​​of conversion coefficients related to the resistivity. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded and the decoding information of the surrounding and decoded blocks or the symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, performs arithmetic decoding of the bins, and generates symbols corresponding to the values ​​of each syntax element.In this case, the CABAC entropy decoding method can update the context model after determining the context model by utilizing the decoded symbol / bin information for the context model of the next symbol / bin. Of the information decoded by the entropy decoding unit 310, information related to prediction is provided to the prediction unit (inter-prediction unit 332 and intra-prediction unit 331), and the residual values ​​that have been entropy decoded by the entropy decoding unit 310, i.e., quantized conversion coefficients and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, of the information decoded by the entropy decoding unit 310, information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives signals output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoding unit 310. On the other hand, the decoding device relating to this document may be called a video / image / picture decoding device, and the decoding device may also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, inverse transformation unit 322, addition unit 340, filtering unit 350, memory 360, interpretation unit 332, and intraprediction unit 331.

[0061] The inverse quantization unit 321 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 321 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients.

[0062] In the inverse conversion unit 322, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0063] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 310, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block, and can determine a specific intra / inter-prediction mode.

[0064] The prediction unit 320 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for prediction of a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also base its predictions on an intra-block copy (IBC) prediction mode or a palette mode for predictions on a block. The IBC prediction mode or palette mode can be used for content video / movie coding such as games, for example, as in SCC (screen content coding). IBC basically performs predictions within the current picture, but can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, information about the palette table and palette index can be included in the video / movie information and signaled.

[0065] The intra-prediction unit 331 can predict the current block by referring to a sample in the current picture. The referenced sample can be located in the vicinity (neighbor) of the current block or at a distance from it, depending on the prediction mode. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-prediction unit 331 can also determine the prediction mode to be applied to the current block using the prediction modes applied to the surrounding blocks.

[0066] The interprediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, neighboring blocks may include spatial neighboring blocks that exist in the current picture and temporal neighboring blocks that exist in the reference picture. For example, the interprediction unit 332 can construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes, and the prediction information may include information indicating the mode of interprediction for the current block.

[0067] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (which comprises an inter-prediction unit 332 and / or an intra-prediction unit 331). If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restored block.

[0068] The summing unit 340 may be called the restoration unit or restoration block generation unit. The generated restoration signal can be used for intra-prediction of the next block to be processed in the current picture, and can be output after filtering as described later, or it can be used for intra-prediction of the next picture.

[0069] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.

[0070] The filtering unit 350 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.

[0071] The (modified) restored picture stored in the DPB of memory 360 can be used as a reference picture by the inter-prediction unit 332. Memory 360 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 260 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 360 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 331.

[0072] In this specification, embodiments described for the filtering unit 260, inter-prediction unit 221, and intra-prediction unit 222 of the encoding device 200 can also be applied identically or in a corresponding manner to the filtering unit 350, inter-prediction unit 332, and intra-prediction unit 331 of the decoding device 300, respectively.

[0073] When inter-prediction is applied, the prediction unit of the encoding / decoding device can perform inter-prediction on a block-by-block basis to derive predicted samples. Inter-prediction can be a prediction derived in a manner that is dependent on data elements (e.g., sample values ​​or motion information) of picture(s) other than the current picture. When inter-prediction is applied to the current block, a predicted block (predicted sample array) for the current block can be derived based on the reference block (reference sample array) identified by the motion vector on the reference picture pointed to by the index of the reference picture. In this case, in order to reduce the amount of motion information transmitted in inter-prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between surrounding blocks and the current block. The motion information may include the motion vector and the index of the reference picture. The motion information may further include information on the inter-prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). When interpretation is applied, a neighboring block can include a spatial neighboring block currently present in the picture and a temporal neighboring block present in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to by names such as a collocated reference block or colCU, and the reference picture containing the temporal neighboring block may be referred to as a collocated picture (colPic).For example, a candidate list of motion information can be constructed based on the surrounding blocks of the current block, and a flag or index information indicating which candidate is selected (used) can be signaled to derive the motion vector and / or the index of the reference picture of the current block. Interpretation is performed based on various prediction modes; for example, in skip mode and merge mode, the motion information of the current block may be identical to the motion information of the selected surrounding block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected surrounding block can be used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0074] The motion information may include L0 motion information and / or L1 motion information depending on the interpretation type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. A prediction based on an L0 motion vector may be called an L0 prediction, a prediction based on an L1 motion vector may be called an L1 prediction, and a prediction based on both the L0 motion vector and the L1 motion vector may be called a bi (Bi) prediction. Here, an L0 motion vector may represent a motion vector associated with a reference picture list L0 (L0), and an L1 motion vector may represent a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include pictures earlier in the output order than the current picture, and the reference picture list L1 may include pictures later in the output order than the current picture. The aforementioned earlier picture may be called a forward (reference) picture, and the aforementioned later picture may be called a reverse (reference) picture. The reference picture list L0 may include further reference pictures that are later in the output order than the current picture. In this case, the earlier picture may be indexed first in the reference picture list L0, and the later picture may be indexed afterward. The reference picture list L1 may include further reference pictures that are earlier in the output order than the current picture. In this case, the later picture may be indexed first in the reference picture list L1, and the earlier picture may be indexed afterward. Here, the output order may correspond to the POC (picture order count) order.

[0075] A variety of interpretation modes can be used to predict the current block within a picture. For example, various modes such as merge mode, skip mode, MVP (motion vector prediction) mode, affine mode, subblock merge mode, and MMVD (merge with MVD) mode can be used. DMVR (Decoder side motion vector refinement) mode, AMVR (adaptive motion vector resolution) mode, Bi-prediction with CU-level weight (BCW), and Bi-directional optical flow (BDOF) can be used as additional or alternative modes. The affine mode is sometimes called the affine motion prediction mode. The MVP mode is sometimes called the AMVP (advanced motion vector prediction) mode. In this document, candidate motion information derived from some modes and / or some modes may be included as one of the candidate motion information for other modes. For example, an HMVP candidate may be added as a merge candidate in the merge / skip modes, or as an mvp candidate in the MVP mode. When the HMVP candidate is used as a candidate for motion information in the merge mode or skip mode, the HMVP candidate may be called an HMVP merge candidate.

[0076] Prediction mode information indicating the inter-prediction mode of the current block can be signaled from the encoding device to the decoding device. The prediction mode information can be included in the bitstream and received by the decoding device. The prediction mode information may include index information indicating one of a number of candidate modes. Alternatively, the inter-prediction mode may be indicated through hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether the skip mode is applicable, and if the skip mode is not applicable, a merge flag may be signaled to indicate whether the merge mode is applicable, and if the merge mode is not applicable, it may indicate that the MVP mode is applicable, or further flags for additional distinctions may be signaled. Affine modes may be signaled as independent modes, or as modes dependent on the merge mode or MVP mode, etc. For example, affine modes may include affine merge mode and affine MVP mode.

[0077] On the other hand, the current block can be signaled with information indicating whether the aforementioned list0 (L0) prediction, list1 (L1) prediction, or bi (BI) prediction is used in the current block (current coding unit). This information may be called motion prediction direction information, inter-prediction direction information, or inter-prediction instruction information, and can be composed / encoded / signaled, for example, in the form of an inter_pred_idc syntax element. That is, the inter_pred_idc syntax element can indicate whether the aforementioned L0 prediction, L1 prediction, or bi-prediction is used in the current block (current coding unit). For the sake of explanation, in this document, the inter-prediction type (L0 prediction, L1 prediction, or BI prediction) pointed to by the inter_pred_idc syntax element can be expressed as motion prediction direction. For example, L0 prediction may be represented as pred_L0, L1 prediction as pred_L1, and bi-prediction as pred_BI.

[0078] As mentioned above, a single picture can contain one or more slices. A slice can have one of the following slice types: I-slice (intra slice), P-slice (predictive slice), and B-slice (bi-predictive slice). The slice type can be indicated based on the slice type information. For blocks in an I-slice, only intra-prediction can be used for prediction, and inter-prediction is not used. Of course, even in this case, original sample values ​​may be coded and signaled without prediction. For blocks in a P-slice, intra-prediction or inter-prediction can be used, and if inter-prediction is used, only uni-prediction can be used. On the other hand, for blocks in a B-slice, intra-prediction or inter-prediction can be used, and if inter-prediction is used, up to the maximum bi-prediction can be used. That is, for blocks in a B-slice, if inter-prediction is used, either uni-prediction or bi-prediction can be used.

[0079] L0 and L1 can contain reference pictures that were encoded / decoded before the current picture. Here, L0 can refer to reference picture list 0, and L1 can refer to reference picture list 1. For example, L0 can contain reference pictures that are earlier and / or later than the current picture in POC (picture order count) order, and L1 can contain reference pictures that are later and / or earlier than the current picture in POC order. In this case, L0 may be assigned an index of a reference picture that is even lower relative to a reference picture that is earlier than the current picture in POC order, and L1 may be assigned an index of a reference picture that is even lower relative to a reference picture that is later than the current picture in POC order. In the case of a B slice, biprediction can be applied, and in this case as well, unidirectional biprediction or bidirectional biprediction can be applied. Bidirectional biprediction may be called true biprediction.

[0080] On the other hand, interpretation can be performed using motion information of the current block. The encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding device can use the original block in the original picture for the current block to search for a highly correlated similar reference block in fractional pixel units within a defined search range in the reference picture, thereby deriving motion information. Block similarity can be derived based on the difference in phase-based sample values. For example, block similarity can be calculated based on the sum of absolute differences (SAD) between the current block (or template of the current block) and the reference block (or template of the reference block). In this case, motion information can be derived based on the reference block with the smallest SAD within the search area. The derived motion information can be signaled to the decoding device in various ways based on the interpretation mode.

[0081] Figure 4 is a diagram illustrating the merge mode in interpretation.

[0082] When merge mode is applied, the movement information of the currently predicted block is not transmitted directly, but rather the movement information of the surrounding predicted block is used to guide the movement information of the currently predicted block. Therefore, the movement information of the currently predicted block can be instructed by transmitting flag information indicating that merge mode has been used and a merge index indicating which surrounding predicted block was used. This merge mode may be called regular merge mode. For example, this merge mode can be applied when the value of the syntax element of regular_merge_flag is 1.

[0083] The encoding device must search for merge candidate blocks to be used to guide the motion information of the currently predicted block in order to perform merge mode. For example, up to five merge candidate blocks may be used, but the embodiments in this document are not limited to this. The maximum number of merge candidate blocks may also be transmitted in the slice header or tile group header, and the embodiments in this document are not limited to this. After finding the merge candidate blocks, the encoding device can generate a merge candidate list and select the merge candidate block with the lowest cost from among them as the final merge candidate block.

[0084] This document can provide various embodiments for the merge candidate blocks that constitute the merge candidate list.

[0085] For example, the merge candidate list can use five merge candidate blocks. For example, it can use four spatial merge candidates and one temporal merge candidate. As a specific example, in the case of a spatial merge candidate, the block shown in Figure 4 can be used as a spatial merge candidate. Hereinafter, the spatial merge candidate or the spatial MVP candidate described later may be called an SMVP, and the temporal merge candidate or the temporal MVP candidate described later may be called a TMVP.

[0086] The merge candidate list for the current block can be constructed, for example, based on the following procedure:

[0087] A coding device (encoding device / decoding device) can search for spatially surrounding blocks of the current block and insert the derived spatial merge candidates into a merge candidate list. For example, the spatially surrounding blocks may include the surrounding block at the lower left corner of the current block, the surrounding block to the left, the surrounding block at the upper right corner, the surrounding block above, and the surrounding block at the upper left corner. However, this is an example, and additional surrounding blocks such as the surrounding block to the right, the surrounding block below, and the surrounding block at the lower right can also be used as spatially surrounding blocks. The coding device can search for the spatially surrounding blocks based on priority, detect available blocks, and derive the movement information of the detected blocks as spatial merge candidates. For example, an encoding device or a decoding device can search the five blocks shown in Figure 4 in the order A1->B1->B0->A0->B2, and sequentially index the available candidates to form a merge candidate list.

[0088] The coding device can search for the temporal neighboring blocks of the current block and insert the derived temporal merge candidates into the merge candidate list. The temporal neighboring blocks may be located on a reference picture that is a picture different from the current picture in which the current block is located. The reference picture on which the temporal neighboring blocks are located may be referred to as a collocated picture or a col picture. The temporal neighboring blocks can be searched in the order of the neighboring blocks around the lower right corner and the block at the center of the lower right side of the co-located block for the current block on the col picture. On the other hand, when motion data compression is applied, specific motion information can be stored as representative motion information for each fixed storage unit in the col picture. In this case, it is not necessary to store the motion information for all blocks within the fixed storage unit, and through this, the effect of motion data compression can be obtained. In this case, the fixed storage unit may be predetermined, for example, in a 16x16 sample unit, an 8x8 sample unit, etc., or the size information for the fixed storage unit may be signaled from the encoding device to the decoding device. When motion data compression is applied, the motion information of the temporal neighboring blocks can be replaced with the representative motion information of the fixed storage unit in which the temporal neighboring blocks are located. That is, in this case, from the aspect of implementation, instead of the prediction block located at the coordinates of the temporal neighboring blocks, based on the coordinates (upper left sample position) of the temporal neighboring blocks, after arithmetic right shift by a certain value and then arithmetic left shift, the motion information of the prediction block covering the shifted position can be used to derive the temporal merge candidate. For example, when the fixed storage unit is a 2nx2n sample unit, if the coordinates of the temporal neighboring blocks are (xTnb, yTnb), the motion information of the prediction block located at the modified positions ((xTnb>>n)<<n), (yTnb>>n)<<n)) can be used for the temporal merge candidate.Specifically, for example, if the constant storage unit is a 16x16 sample unit, and the coordinates of the temporally surrounding block are (xTnb, yTnb), then the motion information of the predicted block located at the corrected position ((xTnb>>4)<<4), (yTnb>>4)<<4)) can be used for the temporal merge candidate. Alternatively, for example, if the constant storage unit is an 8x8 sample unit, and the coordinates of the temporally surrounding block are (xTnb, yTnb), then the motion information of the predicted block located at the corrected position ((xTnb>>3)<<3), (yTnb>>3)<<3)) can be used for the temporal merge candidate.

[0089] The coding device can check whether the current number of merge candidates is less than the maximum number of merge candidates. The maximum number of merge candidates can be predefined or signaled from the encoding device to the decoding device. For example, the encoding device can generate and encode information regarding the maximum number of merge candidates and transmit it to the decoder in the form of a bitstream. Once the maximum number of merge candidates is met, the process of adding candidates further does not need to be performed.

[0090] If the result of the above check is that the current number of merge candidates is less than the maximum number of merge candidates, the coding device may insert additional merge candidates into the merge candidate list. For example, the additional merge candidates may include at least one of the following: history-based merge candidates(s), pair-wise average merge candidates(s), ATMVP, combined bi-predictive merge candidates (if the slice / tile group type of the current slice / tile group is type B), and / or zero-vector merge candidates.

[0091] If, as a result of the above check, the current number of merge candidates is not less than the maximum number of merge candidates, the coding device may terminate the construction of the merge candidate list. In this case, the encoding device can select the optimal merge candidate from among the merge candidates constituting the merge candidate list based on the RD (rate-distortion) cost, and can signal selection information (e.g., merge index) pointing to the selected merge candidate to the decoding device. The decoding device can select the optimal merge candidate based on the merge candidate list and the selection information.

[0092] As previously stated, the motion information of the selected merge candidate can be used for the motion information of the current block, and based on the motion information of the current block, predicted samples of the current block can be derived. The encoding device can derive residual samples of the current block based on the predicted samples and signal residual information regarding the residual samples to the decoding device. As previously stated, the decoding device can generate restored samples based on the residual samples derived based on the residual information and the predicted samples, and based on these, can generate a restored picture.

[0093] When skip mode is applied, the motion information of the current block can be derived in the same way as when merge mode is applied. However, when skip mode is applied, the residual signal for the block is omitted, and therefore, the predicted sample can be immediately used as the restored sample. Skip mode can be applied, for example, when the value of the syntax element of cu_skip_flag is 1.

[0094] On the other hand, the pair-wise average merge candidate may be called a pair-wise average candidate or pair-wise candidate. A pair-wise average candidate can be generated by averaging pairs of predefined candidates in an existing list of merge candidates. A predefined pair can be defined as {(0,1), (0,2), (1,2), (0,3), (1,3), (2,3)}, where the numbers can represent the merge index relative to the list of merge candidates. An averaged motion vector can be calculated separately for each reference list. For example, if two motion vectors are available in a list, they can be averaged even if they point to different reference pictures. For example, if only one motion vector is available, that one can be used directly. For example, if no motion vectors are available, the list can be kept invalid.

[0095] For example, if the merge candidate list is not filled even after the pairwise mean merge candidates have been added, i.e., if the number of current merge candidates in the merge candidate list is less than the number of the largest merge candidate, a zero vector (zero MVP) may be inserted at the end until the number of the largest merge candidate is indicated. In other words, a zero vector may be inserted until the number of current merge candidates in the merge candidate list is equal to the number of the largest merge candidate.

[0096] On the other hand, existing methods allowed for the use of a single motion vector to represent the movement of a coding block; that is, a translation motion model could be used. However, while this method may represent the optimal movement at the block level, coding efficiency can be improved if the optimal motion vector can be determined on a sample-by-sample basis, rather than the optimal movement of each actual sample. For this reason, an affine motion model can be used. The affine motion prediction method for coding using an affine motion model is as follows.

[0097] Affine motion prediction methods can represent motion vectors for each sample in a block using two, three, or four motion vectors. For example, an affine motion model can represent four types of motion. Of the motions that an affine motion model can represent, an affine motion model that represents three types of motion (translation, scale, and rotation) may be called a similarity (or simplified) affine motion model, and this will be used as a basis for explanation, but it is not limited to the motion models mentioned above.

[0098] Figures 5a and 5b illustrate CPMV for affine motion prediction.

[0099] Affine motion prediction can determine the motion vector of a sample position contained within a block using two or more control point motion vectors (CPMVs). In this case, the set of motion vectors can be referred to as an affine motion vector field (MVF).

[0100] For example, Figure 5a can show the case where two CPMVs are used, which can be called a four-parameter affine model. In this case, the motion vector at the sample position (x,y) can be determined, for example, as shown in Equation 1.

[0101]

number

[0102] For example, Figure 5b can show the case where three CPMVs are used, which can be called a 6-parameter affine model. In this case, the motion vector at the sample position (x,y) can be determined, for example, as shown in Equation 2.

[0103]

number

[0104] In equations 1 and 2, {v x ,v y} can represent the motion vector at the (x,y) position. Also, {v 0x ,v 0y} can indicate the CPMV of the control point (CP) at the upper left corner of the coding block, {v 1x ,v 1y} can indicate the CPMV of the CP at the upper right corner, {v 2x ,v 2y} can indicate the CPMV of the CP at the lower left corner. Also, W can indicate the current block width, and H can indicate the current block height.

[0105] Figure 6 illustrates an example where the affine MVF is determined at the subblock level.

[0106] During the encoding / decoding process, the affine MVF can be determined on a sample-by-sample or predefined subblock basis. For example, if determined on a sample basis, the motion vector is obtained based on each sample value. Alternatively, if determined on a subblock basis, the motion vector for that block is obtained based on the sample value of the center of the subblock (the lower right side of the center, i.e., the lower right sample of the four central samples). In other words, in affine motion prediction, the motion vector of the current block can be derived on a sample-by-sample or subblock basis.

[0107] In the embodiment, we may assume that the affine MVF is determined in units of 4x4 subblocks, but this is for the sake of explanation, and the size of the subblocks can be varied in various ways.

[0108] In other words, if affine prediction is available, there are currently three motion models applicable to a block: a translational motion model, a four-parameter affine motion model, and a six-parameter affine motion model. Here, the translational motion model can represent a model in which existing block-unit motion vectors are used, the four-parameter affine motion model can represent a model in which two CPMVs are used, and the six-parameter affine motion model can represent a model in which three CPMVs are used.

[0109] On the other hand, affine motion prediction may include affine MVP (or affine inter) mode or affine merge mode.

[0110] Figure 7 is a diagram illustrating the affine merge mode in interpretation.

[0111] For example, in affine merge mode, CPMV can be determined by the affine motion model of the surrounding blocks coded with affine motion prediction. For example, surrounding blocks coded with affine motion prediction on the search order can be used for affine merge mode. That is, if at least one of the surrounding blocks is coded with affine motion prediction, then the current block can be coded in affine merge mode. Here, affine merge mode may be called AF_MERGE.

[0112] When affine merge mode is applied, the CPMV of the current block can be derived using the CPMV of the surrounding blocks. In this case, the CPMV of the surrounding blocks may be used as is for the current block, or the CPMV of the surrounding blocks may be modified based on the size of the surrounding blocks and the size of the current block, etc., and then used as the CPMV of the current block.

[0113] On the other hand, in the case of affine merge mode, where motion vectors (MV) are derived on a subblock basis, this may be called subblock merge mode, which can be indicated based on the subblock merge flag (or the syntax element of merge_subblock_flag). Alternatively, if the value of the syntax element of merge_subblock_flag is 1, it may be indicated that subblock merge mode is applied. In this case, the affine merge candidate list described later may also be called the subblock merge candidate list. In this case, the subblock merge candidate list may further include candidates derived as SbTMVP described later. In this case, the candidates derived as SbTMVP may be used as the candidate for index 0 in the subblock merge candidate list. In other words, the candidates derived as SbTMVP may be located before the inherited affine candidate or constructed affine candidate described later in the subblock merge candidate list.

[0114] When affine merge mode is applied, an affine merge candidate list can be constructed for the derivation of CPMVs for the current block. For example, the affine merge candidate list may include at least one of the following candidates: 1) inherited affine merge candidate; 2) constructed affine merge candidate; 3) zero motion vector candidate (or zero vector). Here, the inherited affine merge candidate is a candidate derived based on the CPMVs of the surrounding block if the surrounding block is coded in affine mode; the constructed affine merge candidate is a candidate derived by constructing CPMVs based on the MVs of the surrounding block for each CPMV; and the zero motion vector candidate may indicate a candidate constructed with CPMVs whose value is 0.

[0115] The aforementioned affine merge candidate list can be structured, for example, as follows:

[0116] There can be up to two inherited affine candidates, which can be derived from the affine motion model of the surrounding block. A surrounding block can include one left-side surrounding block and one above-side surrounding block. The candidate blocks can be positioned as shown in Figure 4. The scan order for the left predictor can be A1->A0, and the scan order for the above-side predictor can be B1->B0->B2. Only one inherited candidate can be selected from each of the left and above sides. A pruning check may not be performed between the two inherited candidates.

[0117] If a peripheral affine block is identified, the control point motion vectors of the identified block can be used to derive CPMVP candidates in the current block's affine merge list. Here, a peripheral affine block can refer to a block among the peripheral blocks of the current block that is coded in affine prediction mode. For example, referring to Figure 7, if the bottom-left peripheral block A is coded in affine prediction mode, motion vectors v2, v3, and v4 can be obtained for the top-left, top-right, and bottom-left corners of peripheral block A. If peripheral block A is coded with a 4-parameter affine motion model, two CPMVs of the current block can be calculated using v2 and v3. If peripheral block A is coded with a 6-parameter affine motion model, three CPMVs of the current block can be calculated using v2, v3, and v4.

[0118] Figure 8 is a diagram illustrating the candidate positions in affine merge mode.

[0119] The constructed affine candidates may mean candidates constructed by combining translational motion information around each control point. The motion information of the control point can be derived from the specified spatial neighborhood and temporal neighborhood. CPMV k(k=1、2、3、4) can indicate the k-th control point.

[0120] Referring to FIG. 8, for CPMV1, the blocks can be checked in the order of B2->B3->A2, and the motion vector of the first available block can be used. For CPMV2, the blocks can be checked in the order of B1->B0, and for CPMV3, the blocks can be checked in the order of A1->A0. The TMVP (temporal motion vector predictor) can be used as CPMV4 if available.

[0121] After the motion vectors of the four control points are obtained, the affine merge candidates can be constructed based on the obtained motion information. The combinations of the control point motion vectors can be configured as {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2} and {CPMV1, CPMV3}, and can be configured in the listed order.

[0122] The combination of three CPMVs can constitute a 6-parameter affine merge candidate, and the combination of two CPMVs can constitute a 4-parameter affine merge candidate. To avoid the motion scaling process, if the reference indices of the control points are different from each other, the related combinations of the control point motion vectors can be discarded.

[0123] FIG. 9 is a diagram for explaining SbTMVP in inter prediction.

[0124] On the other hand, the SbTMVP (subblock-based temporal motion vector prediction) method is sometimes used. For example, SbTMVP is sometimes called ATMVP (advanced temporal motion vector prediction). SbTMVP can use motion fields in collocated pictures to improve motion vector prediction and the merge mode for CUs in the current picture. Here, collocated pictures are sometimes called col pictures.

[0125] For example, SbTMVP can predict motion at the subblock (or subCU) level. Furthermore, SbTMVP can apply a motion shift before fetching temporal motion information from the colpicture. Here, the motion shift can be obtained from the motion vector of one of the spatially surrounding blocks of the current block.

[0126] SbTMVP can predict the motion vectors of subblocks (or subCUs) within a current block (or CU) in two steps.

[0127] In the first step, the spatial peripheral blocks can be tested in the order A1, B1, B0, and A0 in Figure 4. If a first spatial peripheral block is identified that has a motion vector using the col picture as its reference picture, the motion vector can be selected as the motion shift to be applied. If no such motion is identified from the spatial peripheral block, the motion shift can be set to (0,0).

[0128] In the second step, the motion shift confirmed in the first step can be applied to obtain subblock-level motion information (motion vectors and reference indices) from the col picture. For example, the motion shift can be added to the coordinates of the current block. For example, the motion shift can be set to the motion of A1 in Figure 4. In this case, for each subblock, the motion information of the corresponding block in the col picture can be used to derive the motion information of the subblock. Temporal motion scaling can be applied to align the reference picture of the temporal motion vector with the reference picture of the current block.

[0129] A combined subblock-based merge list containing all SbTVMP candidates and affine merge candidates can be used for signaling affine merge modes, where affine merge mode may be referred to as subblock-based merge mode. SbTVMP mode may or may not be available depending on the flags included in the SPS (sequence parameter set). If SbTMVP mode is available, an SbTMVP predictor may be added as the first entry in the list of subblock-based merge candidates, followed by affine merge candidates. The maximum allowed size of the affine merge candidate list may be five.

[0130] The size of subCUs (or subblocks) used in SbTMVP may be fixed at 8x8, and, similar to affine merge mode, SbTMVP mode can only be applied to blocks where both width and height are 8 or greater. The encoding logic for additional SbTMVP merge candidates may be the same as for other merge candidates. That is, for each CU in a P or B slice, an RD check with an additional RD (rate-distortion) cost may be performed to determine whether to use an SbTMVP candidate.

[0131] On the other hand, based on the motion information derived by the prediction mode, a predicted block can be derived for the current block. The predicted block may include predicted samples (predicted sample arrays) of the current block. If the motion vector of the current block points to fractional sample units, an interpolation procedure may be performed, through which predicted samples of the current block can be derived based on reference samples in fractional sample units within the reference picture. When affine interpretation (affine prediction mode) is applied to the current block, predicted samples can be generated based on MV at the sample / subblock level. When biprediction is applied, predicted samples derived through a (phase-based) weighted sum or weighted average of predicted samples derived based on L0 prediction (i.e., prediction using reference pictures in reference picture list L0 and MVL0) and predicted samples derived based on L1 prediction (i.e., prediction using reference pictures in reference picture list L1 and MVL1) can be used as predicted samples for the current block. Here, the motion vector in the L0 direction may be called the L0 motion vector or MVL0, and the motion vector in the L1 direction may be called the L1 motion vector or MVL1. When bi-prediction is applied, if the reference picture used for L0 prediction and the reference picture used for L1 prediction are located in different temporal directions relative to the current picture (i.e., it is both bi-prediction and bidirectional prediction), this may be called true bi-prediction.

[0132] Furthermore, as mentioned above, reconstructed samples and reconstructed pictures can be generated based on the derived predicted samples, and subsequent procedures such as in-loop filtering can be performed.

[0133] On the other hand, when bi-prediction is applied to a block, the predicted samples can be derived based on a weighted average. For example, bi-prediction using a weighted average may be called BCW (Bi-prediction with CU-level Weight), BWA (Bi-prediction with Weighted Average), or weighted averaging bi-prediction.

[0134] Previously, a dual prediction signal (i.e., dual prediction sample) could be derived through a simple average of the L0 prediction signal (L0 prediction sample) and the L1 prediction signal (L1 prediction sample). That is, the dual prediction sample was derived as the average of the L0 prediction sample based on the L0 reference picture and MVL0 and the L1 prediction sample based on the L1 reference picture and MVL1. However, when dual prediction is applied, the dual prediction signal (dual prediction sample) can also be derived through a weighted average of the L0 prediction signal and the L1 prediction signal, as follows. For example, the dual prediction signal (dual prediction sample) can be derived as shown in Equation 3.

[0135]

number

[0136] In equation 3, P bi-pred P0 can represent the value of the dual prediction signal, i.e., the predicted sample value derived by applying the dual prediction, and w can represent the weighted value. Furthermore, P0 can represent the value of the L0 prediction signal, i.e., the predicted sample value derived by applying the L0 prediction, and P1 can represent the value of the L1 prediction signal, i.e., the predicted sample value derived by applying the L1 prediction.

[0137] For example, in weighted average biprediction, five weight values ​​may be allowed. For example, the five weight values ​​(w) may include -2, 3, 4, 5, or 10. That is, the weight value (w) can be determined by one of the candidate weight values ​​including -2, 3, 4, 5, or 10. For each CU to which biprediction is applied, the weight value w can be determined by one of two methods. The first method is that for non-merged CUs, the weight value index can be signaled after motion vector difference. The second method is that for merged CUs, the weight value index can be inferred from surrounding blocks based on merge candidate indexes.

[0138] For example, weighted average biprediction can be applied to CUs with 256 or more luma samples. That is, weighted average biprediction can be applied when the product of the width and height of the CU is greater than or equal to 256. For low-delay pictures, five weight values ​​may be used, and for non-low-delay pictures, three weight values ​​may be used. For example, the three weight values ​​may include 3, 4, or 5.

[0139] For example, in an encoding device, a fast search algorithm can be applied to find weighted indexes without significantly increasing the complexity of the encoding device. Such algorithms can be summarized as follows: For example, when combined with AMVR (adaptive motion vector resolution) (when AMVR is used as the interprediction mode), if the current picture is a low-latency picture, non-identical weights can be checked as a condition for the precision of the 1-pel and 4-pel motion vectors. For example, when combined with affine (when affine prediction mode is used as the interprediction mode), if affine prediction mode is selected as the current best mode, affine ME (Motion Estimation) can be performed on non-identical weights. For example, if the two reference pictures in a biprediction are identical, non-identical weights can be checked as a condition. For example, depending on the POC distance between the current picture and the reference picture, the coding QP (quantization parameter), and the temporal level, certain conditions may be met that result in non-identical weighted values ​​not being found.

[0140] For example, a BCW weighted index (or weighted index) can be coded using a context-coded bin followed by a bypass-coded bin. The first context-coded bin can indicate whether identical weights are used or not. Based on the first context-coded bin, if non-identical weights are used, an additional bin can be signaled using bypass coding to indicate which non-identical weights are to be used.

[0141] On the other hand, according to one embodiment of this document, when constructing candidate motion vectors for merge mode, a weighted index for a weighted average can be derived when the candidate temporal motion vectors use bi-prediction. That is, when the inter-prediction type is bi-prediction, weighted index information can be derived for the temporal merge candidates (or temporal motion vector candidates) in the merge candidate list.

[0142] For example, for a candidate time motion vector, the weight index for the weighted average can always be derived as 0. Here, a weight index of 0 may mean that the weights for each reference direction (i.e., the L0 and L1 prediction directions in biprediction) are the same. For example, in this case, the procedure for deriving the motion vector of the luma component for the merge mode is as shown in the following table.

[0143] [Table 1]

[0144] [Table 2]

[0145] [Table 3]

[0146] Tables 1 to 3 above may represent a single procedure, and the procedure may be performed sequentially in the order of the tables. The procedure may include a procedure for deriving the motion vector of the luma component for the merge mode (8.4.2.2).

[0147] Referring to Tables 1 to 3 above, gbiIdx can represent a bipredictive weighted index, gbiIdxCol can represent a bipredictive weighted index for a temporal merge candidate (e.g., a candidate for a temporal motion vector in the merge candidate list), and in the procedure for deriving the motion vector of the luma component for the merge mode (third step in 8.4.2.2), gbiIdxCol can be derived as 0. That is, the weighted index for a candidate for a temporal motion vector can be derived as 0.

[0148] Alternatively, for example, the weighted index for a weighted average of candidate time motion vectors can be derived as the weighted index of a collocated block. Here, a collocated block may be called a col block, same-position block, or same-position reference block, and a col block can represent a block at the same position as the current block on the reference picture. For example, in this case, the procedure for deriving the motion vector of the luma component for the merge mode is as shown in the following table.

[0149] [Table 4]

[0150] [Table 5]

[0151] [Table 6]

[0152] Tables 4 to 6 may represent a single procedure, and the procedure may be performed sequentially in the order of the tables. The procedure may include a procedure for deriving the motion vector of the luma component for the merge mode (8.4.2.2).

[0153] Referring to Tables 4 to 6 above, gbiIdx can represent a bipredictive weighted index, gbiIdxCol can represent a bipredictive weighted index for a temporal merge candidate (e.g., a candidate for a temporal motion vector in the merge candidate list), and in the procedure for deriving the motion vector of the luma component for the merge mode (third step in 8.4.2.2), gbiIdxCol can be derived as 0, but if the slice type or tile group type is B (fourth step in 8.4.2.2), then gbiIdxCol can be derived as gbiIdxCol. That is, the weighted index of a candidate for a temporal motion vector can be derived as the weighted index of a col block.

[0154] On the other hand, according to another embodiment of this document, when constructing motion vector candidates for a subblock-level merge mode, a weighted index for a weighted average can be derived when the temporal motion vector candidates use bi-prediction. Here, the subblock-level merge mode may be called an affine merge mode, and the temporal motion vector candidates may refer to subblock-based temporal motion vector candidates, sometimes called SbTMVP candidates. That is, when the inter-prediction type is bi-prediction, weighted index information can be derived for SbTMVP candidates (or subblock-based temporal motion vector candidates) in the affine merge candidate list or the subblock merge candidate list.

[0155] For example, the weighted index for the weighted average of candidate subblock-based temporal motion vectors can always be derived as 0. Here, a weighted index of 0 may mean that the weights for each reference direction (i.e., the L0 and L1 prediction directions in biprediction) are the same. For example, in this case, the procedure for deriving the motion vectors and reference index in the subblock merge mode or the procedure for deriving subblock-based temporal merge candidates is as shown in the following table.

[0156] [Table 7]

[0157] [Table 8]

[0158] [Table 9]

[0159] [Table 10]

[0160] [Table 11]

[0161] Tables 7 to 11 may show two procedures, which may be performed sequentially in the order shown in the tables. The procedures may include a procedure for deriving motion vectors and reference indices within a subblock merge mode (8.4.4.2) or a procedure for deriving subblock-based temporal merge candidates (8.4.4.3).

[0162] Referring to Tables 7 to 11, gbiIdx can represent a biprediction weighted index, gbiIdxSbCol can represent a biprediction weighted index for a subblock-based temporal merge candidate (e.g., a candidate for a temporal movement vector in a subblock-based merge candidate list), and in the procedure for deriving the subblock-based temporal merge candidate (8.4.4.3), gbiIdxSbCol can be derived as 0. That is, the weighted index for a candidate for a subblock-based temporal movement vector can be derived as 0.

[0163] Alternatively, for example, a weighted index for a weighted average of candidate subblock-based temporal motion vectors can be derived as a weighted index of a temporal center block. For example, the temporal center block may represent a col block or a subblock or sample located in the center of a col block, specifically, a subblock or sample located in the lower right of the four central subblocks or samples of a col block. For example, in this case, the procedure for deriving motion vectors and reference indices in a subblock merge mode, the procedure for deriving subblock-based temporal merge candidates, or the procedure for deriving base motion information for subblock-based temporal merges are as shown in the following table.

[0164] [Table 12]

[0165] [Table 13]

[0166] [Table 14]

[0167] [Table 15]

[0168] [Table 16]

[0169] [Table 17]

[0170] [Table 18]

[0171] [Table 19]

[0172] [Table 20]

[0173] Tables 12 to 20 may show three procedures, which may be performed sequentially in the order shown in the tables. The procedures may include a procedure for deriving motion vectors and reference indices within a subblock merge mode (8.4.4.2), a procedure for deriving subblock-based temporal merge candidates (8.4.4.3), or a procedure for deriving base motion information for subblock-based temporal merge (8.4.4.4).

[0174] Referring to Tables 12 to 20, gbiIdx can represent a bipredictive weighted index, gbiIdxSbCol can represent a bipredictive weighted index for a subblock-based temporal merge candidate (e.g., a temporal motion vector candidate in a subblock-based merge candidate list), and in the procedure for deriving base motion information for subblock-based temporal merges (8.4.4.4), gbiIdxSbCol can be derived as gbiIdxcolCb. That is, the weighted index for a subblock-based temporal motion vector candidate can be derived as a temporal center block. For example, the temporal center block can represent a col block or a subblock or sample located in the center of a col block, specifically a subblock or sample located in the lower right of the four central subblocks or samples in a col block.

[0175] Alternatively, for example, a weighted index for a weighted average of candidate subblock-based temporal motion vectors can be derived as a weighted index for each subblock unit, or, if a subblock is unavailable, as a weighted index for the temporal center block. For example, the temporal center block may refer to a col block or a subblock or sample located in the center of a col block, specifically, a subblock or sample located in the lower right of the four central subblocks or samples of a col block. For example, in this case, the procedure for deriving motion vectors and reference indices within a subblock merge mode, the procedure for deriving subblock-based temporal merge candidates, or the procedure for deriving base motion information for subblock-based temporal merges are as shown in the following table.

[0176] [Table 21]

[0177] [Table 22]

[0178] [Table 23]

[0179] [Table 24]

[0180] [Table 25]

[0181] [Table 26]

[0182] [Table 27]

[0183] [Table 28]

[0184] [Table 29]

[0185] Tables 21 to 29 above may show three procedures, and each of these procedures may be performed sequentially in the order shown in the tables. The procedures may include a procedure for deriving motion vectors and reference indices within a subblock merge mode (8.4.4.2), a procedure for deriving subblock-based temporal merge candidates (8.4.4.3), or a procedure for deriving base motion information for subblock-based temporal merge (8.4.4.4).

[0186] Referring to Tables 21 to 29 above, gbiIdx can represent a biprediction weighted index, gbiIdxSbCol can represent a biprediction weighted index for a subblock-based temporal merge candidate (e.g., a temporal motion vector candidate in a subblock-based merge candidate list), and in the procedure for deriving base motion information for subblock-based temporal merges (8.4.4.3), gbiIdxSbCol can be derived as gbiIdxcolCb. Alternatively, depending on the conditions (for example, if availableFlagL0SbCol and availableFlagL1SbCol are all 0), in the procedure for deriving base motion information for subblock-based temporal merging (8.4.4.3), gbiIdxSbCol can be derived as ctrgbiIdx, and in the procedure for deriving base motion information for subblock-based temporal merging (8.4.4.4), ctrgbiIdx can be derived as gbiIdxSbCol. That is, the weighted index of candidate subblock-based temporal motion vectors can be derived as the weighted index for each subblock unit, and if no subblock is available, it can be derived as the temporal center block. For example, the temporal center block can represent a col block or a subblock or sample located in the center of a col block, specifically, a subblock or sample located in the lower right of the four central subblocks or samples of a col block.

[0187] On the other hand, according to yet another embodiment of this document, when constructing motion vector candidates for a merge mode, a weighted index for pair-wise candidates can be derived. In other words, a pair-wise candidate may be included in the merge candidate list, in which case a weighted index for the weighted average of the pair-wise candidate can be derived. For example, the pair-wise candidate may be derived based on other merge candidates in the merge candidate list, and if the pair-wise candidate uses bi-prediction, a weighted index for the weighted average can be derived. That is, if the inter-prediction type is bi-prediction, weighted index information for pair-wise candidates in the merge candidate list can be derived.

[0188] For example, if the pairwise candidate can be derived based on two merge candidates (e.g., cand0 and cand1) in the merge candidate list, and the pairwise candidate uses biprediction, the weighted index of the pairwise candidate can be derived based on the weighted index of the merge candidates cand0 and / or cand1. In other words, the weighted index of the pairwise candidate can be derived as the weighted index of either of the merge candidates (e.g., merge candidate cand0 or merge candidate cand1) used to derive the pairwise candidate. Alternatively, for example, the weighted index of the pairwise candidate can be derived as a specific ratio of the weighted indices of the merge candidates (e.g., merge candidates cand0 and cand1) used to derive the pairwise candidate, where the specific ratio may be 1:1, but may also be derived as a different ratio. For example, a particular ratio can be determined by a default ratio or a default value, but is not limited to this; the default ratio may be defined as a 1:1 ratio, but may also be defined by a different ratio. Alternatively, for example, when deriving the weighted index of a pairwise candidate based on a particular ratio of the weighted index of a merge candidate, the particular ratio may yield the same result as deriving the weighted index of the pairwise candidate as the weighted index of one of the merge candidates, as described above.

[0189] On the other hand, according to yet another embodiment of this document, when constructing candidate motion vectors for a subblock-level merge mode, a weighted index for a weighted average can be derived when the (representative) motion vector candidate uses bi-prediction. That is, when the inter-prediction type is bi-prediction, weighted index information can be derived for the candidates (or affine merge candidates) in the affine merge candidate list or the subblock merge candidate list.

[0190] For example, among the affine merge candidates, the constructed affine merge candidate can derive CP0, CP1, CP2, or RB candidates based on the motion information of spatially adjacent blocks (or spatially surrounding blocks) or temporally adjacent blocks (or temporally surrounding blocks) of the current block, and indicate a candidate from which the MVF is derived as an affine model. For example, CP0 can indicate a control point located at the upper left sample position of the current block, CP1 can indicate a control point located at the upper right sample position of the current block, and CP2 can indicate a control point located at the lower left sample position of the current block. RB can also indicate a control point located at the lower right sample position of the current block.

[0191] For example, if the candidate for (representative) motion vector is a constructed affine merge candidate (or (current) affine merge candidate), the weighted index of the (current) affine merge candidate can be derived as the weighted index of the block determined by the motion vector at CP0 among the CP0 candidate blocks. Alternatively, the weighted index of the (current) affine merge candidate can be derived as the weighted index of the block determined by the motion vector at CP1 among the CP1 candidate blocks. Alternatively, the weighted index of the (current) affine merge candidate can be derived as the weighted index of the block determined by the motion vector at CP2 among the CP2 candidate blocks. Alternatively, the weighted index of the (current) affine merge candidate can be derived as the weighted index of the block determined by the motion vector at RB among the RB candidate blocks. Alternatively, the weighted index of the (current) affine merge candidate may be derived based on at least one of the weighted indexes of the block determined by the motion vector at CP0, the weighted index of the block determined by the motion vector at CP1, the weighted index of the block determined by the motion vector at CP2, or the weighted index of the block determined by the motion vector at RB. For example, when deriving the weighted index of the (current) affine merge candidate based on multiple weighted indices, a specific ratio to the multiple weighted indices may be used. Here, the specific ratio may be 1:1, 1:1:1, or 1:1:1:1, but may be derived by a different ratio. For example, the specific ratio may be determined by a default ratio or default value, and the default ratio may be defined as a 1:1 ratio, but may be defined by a different ratio.

[0192] Alternatively, for example, the weighted index of the (current) affine merge candidate can be derived as the weighted index of the candidate with the highest frequency of occurrence among the weighted indexes of each candidate. For example, among the CP0 candidate block, the weighted index of the candidate block determined by the motion vector at CP0; among the CP1 candidate block, the weighted index of the candidate block determined by the motion vector at CP1; among the CP2 candidate block, the weighted index of the candidate block determined by the motion vector at CP2; and / or among the RB candidate block, the weighted index that overlaps the most can be derived as the weighted index of the (current) affine merge candidate.

[0193] For example, CP0 and CP1 may be used as the control points, or CP0, CP1 and CP2 may be used, and RB may not be used. However, for example, when trying to utilize RB candidates for affine blocks (blocks coded in affine prediction mode), the method for deriving or deriving weighted indexes in temporal candidate blocks as described in the above embodiment can be used. For example, CP0, CP1, or CP2 can derive candidates based on the spatially surrounding blocks of the current block, and from among the candidates, a block to be used as the motion vector (i.e., CPMV1, CPMV2, or CPMV3) in CP0, CP1, or CP2 can be determined. Alternatively, for example, RB can derive candidates based on the temporally surrounding blocks of the current block, and from among the candidates, a block to be used as the motion vector in RB can be determined.

[0194] Alternatively, for example, if a candidate for (representative) motion vector is an SbTMVP (or ATMVP) candidate, the weighted index of the SbTMVP candidate can be derived as the weighted index of the surrounding block to the left of the current block. That is, if a candidate induced to SbTMVP (or ATMVP) uses bi-prediction, the weighted index of the surrounding block to the left of the current block can be derived as the weighted index for the subblock-based merge mode. That is, if the inter-prediction type is bi-prediction, weighted index information for the SbTMVP candidate in the affine merge candidate list or the subblock merge candidate list can be derived or derived.

[0195] For example, since the SbTMVP candidate can derive a col block based on the spatially adjacent block to the left (or the surrounding block to the left) of the current block, it can be seen that the weighted index of the surrounding block to the left can be trusted. Thus, the weighted index of the SbTMVP candidate can be derived as the weighted index of the surrounding block to the left.

[0196] Figures 10 and 11 schematically show an example of a video / image encoding method and related components according to the embodiments of this document.

[0197] The method disclosed in Figure 10 can be performed by the encoding apparatus disclosed in Figure 2 or Figure 11. Specifically, for example, steps S1000 to S1030 in Figure 10 can be performed by the prediction unit 220 of the encoding apparatus 200 in Figure 11, and step S1040 in Figure 10 can be performed by the entropy encoding unit 240 of the encoding apparatus 200 in Figure 11. Although not shown in Figure 10, the prediction unit 220 of the encoding apparatus 200 in Figure 11 can derive predicted samples or prediction-related information, the residual processing unit 230 of the encoding apparatus 200 can derive residual information from original samples or predicted samples, and the entropy encoding unit 240 of the encoding apparatus 200 can generate a bitstream from the residual information or prediction-related information. The method disclosed in Figure 10 can include the embodiments described above in this document.

[0198] Referring to Figure 10, the encoding device can determine the inter-prediction mode of the current block and generate inter-prediction mode information indicating the inter-prediction mode (S1000). For example, the encoding device can determine a merge mode, an affine (merge) mode, or a sub-block merge mode as the inter-prediction mode to be applied to the current block, and generate inter-prediction mode information indicating this.

[0199] The encoding device can generate a list of merge candidates for the current block based on the inter prediction mode (S1010). For example, the encoding device can generate a merge candidate list based on the determined inter prediction mode. Here, if the determined inter prediction mode is an affine merge mode or a subblock merge mode, the merge candidate list may be called an affine merge candidate list or a subblock merge candidate list, etc., but is sometimes simply called a merge candidate list.

[0200] For example, candidates can be inserted into the merge candidate list until the number of candidates in the merge candidate list reaches the maximum number of candidates. Here, a candidate can represent a candidate or candidate block for deriving motion information (or motion vector) of the current block. For example, a candidate block can be derived through a search of the surrounding blocks of the current block. For example, the surrounding blocks may include spatial and / or temporal surrounding blocks of the current block, and spatial surrounding blocks may be searched preferentially to derive (spatial merge) candidates, then temporal surrounding blocks may be searched to derive (temporal merge) candidates, and the derived candidates can be inserted into the merge candidate list. For example, even after the candidates have been inserted, if the number of candidates in the merge candidate list is less than the maximum number of candidates, additional candidates can be inserted. For example, additional candidates may include at least one of the following: history-based merge candidate(s), pair-wise average merge candidate(s), ATMVP, combined bi-predictive merge candidate(s) (if the slice / tile group type of the current slice / tile group is type B), and / or zero-vector merge candidate(s).

[0201] Alternatively, for example, candidates can be inserted into the affine merge candidate list until the number of candidates in the affine merge candidate list reaches the maximum number of candidates. Here, a candidate may include the CPMV (Control Point Motion Vector) of the current block. Alternatively, the candidate may indicate a candidate or candidate block for deriving the CPMV. The CPMV may indicate the motion vector at the CP (Control Point) of the current block. For example, there may be two, three, or four CPs, and they may be located on at least part of the upper left side (or upper left corner), upper right side (or upper right corner), lower left side (or lower left corner), or lower right side (or lower right corner) of the current block, with only one CP at each position.

[0202] For example, candidates can be derived through a search of the surrounding blocks of the current block (or the surrounding blocks of the current block's CP). For example, the affine merge candidate list can include at least one of the following: inherited affine merge candidates, constructed affine merge candidates, or zero motion vector candidates. For example, the affine merge candidate list can initially insert the inherited affine merge candidates, and thereafter insert constructed affine merge candidates. Also, if the number of candidates in the affine merge candidate list is less than the maximum number of candidates after inserting up to the constructed affine merge candidates, the remainder can be filled with zero motion vector candidates. Here, zero motion vector candidates are sometimes called zero vectors. For example, the affine merge candidate list may be a list of affine merge modes in which motion vectors are derived on a sample-by-sample basis, or it may be a list of affine merge modes in which motion vectors are derived on a subblock-by-subblock basis. In this case, the affine merge candidate list is sometimes called the subblock merge candidate list, and the subblock merge candidate list may also include candidates derived as SbTMVP (or SbTMVP candidates). For example, if an SbTMVP candidate is included in the subblock merge candidate list, it may be located before any inherited affine merge candidates and constructed affine merge candidates within the subblock merge candidate list.

[0203] The encoding device can select one candidate from among those included in the merge candidate list and generate selection information indicating the selected candidate (S1020). For example, the merge candidate list may include at least some of spatial merge candidates, temporal merge candidates, pairwise candidates, or zero vector candidates, and one of these candidates can be selected for inter prediction of the current block. Alternatively, for example, the subblock merge candidate list may include at least some of inherited affine merge candidates, configured affine merge candidates, SbTMVP candidates, or zero vector candidates, and one of these candidates can be selected for inter prediction of the current block.

[0204] For example, the selection information may include index information indicating the selected candidate within the merge candidate list. For example, the selection information may also be called merge index information or subblock merge index information.

[0205] The encoding device can generate inter-prediction type information indicating the current block's inter-prediction type as a bi-prediction (S1030). For example, the current block's inter-prediction type can be determined as a bi-prediction among L0 prediction, L1 prediction, or bi-prediction, and the encoding device can generate inter-prediction type information indicating this. Here, L0 prediction can indicate a prediction based on reference picture list 0, L1 prediction can indicate a prediction based on reference picture list 1, and bi-prediction can indicate a prediction based on both reference picture list 0 and reference picture list 1. For example, the encoding device can generate inter-prediction type information based on the inter-prediction type. For example, the inter-prediction type information may include the syntax element inter_pred_idc.

[0206] The encoding device can encode video information including interprediction mode information, selection information, and interprediction type information (S1040). For example, the video information may also be called video information. The video information may include various types of information relating to the embodiments described above in this document. For example, the video information may include at least a portion of prediction-related information or residual-related information. For example, the prediction-related information may include at least a portion of the interprediction mode information, selection information, and interprediction type information. For example, the encoding device can encode video information including all or part of the aforementioned information (or syntax elements) and generate a bitstream or encoded information. Alternatively, it can output it in the form of a bitstream. The bitstream or encoded information can also be transmitted to a decoding device via a network or storage medium.

[0207] Although not shown in Figure 10, for example, the encoding device can generate predicted samples of the current block. Alternatively, for example, the encoding device can generate predicted samples of the current block based on selected candidates. Alternatively, for example, the encoding device can derive motion information based on selected candidates and generate predicted samples of the current block based on the motion information. For example, the encoding device can generate L0 predicted samples and L1 predicted samples by dual prediction, and generate predicted samples of the current block based on the L0 predicted samples and the L1 predicted samples. In this case, weighted index information (or weighted information) for dual prediction can be used to generate predicted samples of the current block from the L0 predicted samples and the L1 predicted samples. Here, the weighted information can be shown based on the weighted index information.

[0208] In other words, for example, the encoding device can generate L0 and L1 prediction samples for the current block based on the selected candidates. For example, if the interpretation type of the current block is determined to be biprediction, then reference picture list 0 and reference picture list 1 can be used for the prediction of the current block. For example, the L0 prediction sample may represent the prediction sample of the current block derived based on reference picture list 0, and the L1 prediction sample may represent the prediction sample of the current block derived based on reference picture list 1.

[0209] For example, the candidates may include spatial merge candidates. For example, if the selected candidate is a spatial merge candidate, L0 motion information and L1 motion information can be derived based on the spatial merge candidate, and based on this, the L0 prediction sample and the L1 prediction sample can be generated.

[0210] For example, the candidates may include temporal merge candidates. For example, if the selected candidate is a temporal merge candidate, L0 motion information and L1 motion information can be derived based on the temporal merge candidate, and based on this, the L0 prediction sample and the L1 prediction sample can be generated.

[0211] For example, the candidates may include pairwise candidates. For example, if the selected candidate is a pairwise candidate, L0 motion information and L1 motion information can be derived based on the pairwise candidate, and the L0 prediction sample and the L1 prediction sample can be generated based on this. For example, the pairwise candidate may be derived based on two other candidates from among the candidates included in the merge candidate list.

[0212] Alternatively, for example, the merge candidate list may be a list of subblock merge candidates, and affine merge candidates, subblock merge candidates, or SbTMVP candidates may be selected. Here, affine merge candidates at the subblock level are sometimes called subblock merge candidates.

[0213] For example, the candidates may include subblock merge candidates. For example, if the selected candidate is a subblock merge candidate, L0 motion information and L1 motion information can be derived based on the subblock merge candidate, and the L0 prediction sample and the L1 prediction sample can be generated based on this. For example, the subblock merge candidate may include a CPMV (Control Point Motion Vector), and the L0 prediction sample and the L1 prediction sample can be generated by performing predictions on a subblock basis based on the CPMV.

[0214] Here, CPMV can be represented based on one of the surrounding blocks of the current block's CP (Control Point). For example, there may be two, three, or four CPs, and they may be located on at least part of the upper left side (or upper left corner), upper right side (or upper right corner), lower left side (or lower left corner), or lower right side (or lower right corner) of the current block, with only one CP present at each location.

[0215] For example, CP could be CP0, located to the upper left of the current block. In this case, the surrounding blocks could include the surrounding block at the upper left corner of the current block, the surrounding block to the left adjacent to the lower side of the surrounding block at the upper left corner, and the surrounding block to the upper side adjacent to the right side of the surrounding block at the upper left corner. Alternatively, the surrounding blocks could include blocks A2, B2, or B3 in Figure 8.

[0216] Alternatively, for example, the CP could be CP1 located to the upper right of the current block. In this case, the surrounding block could include the surrounding block at the upper right corner of the current block and the upper surrounding block adjacent to the left of the surrounding block at the upper right corner. Alternatively, the surrounding block could include block B0 or ​​block B1 in Figure 8.

[0217] Alternatively, for example, CP could be CP2 located to the lower left of the current block. In this case, the surrounding block could include the surrounding block at the lower left corner of the current block and the surrounding block to the left adjacent to the upper side of the surrounding block at the lower left corner. Alternatively, the surrounding block could include block A0 or block A1 in Figure 8.

[0218] Alternatively, for example, the CP could be CP3 located to the lower right of the current block. Here, CP3 is sometimes called RB. In this case, the surrounding block may include the col block of the current block or the surrounding block at the lower right corner of the col block. Here, the col block may include a block at the same position as the current block in a reference picture different from the current picture in which the current block is located. Alternatively, the surrounding block may include a T block in Figure 8.

[0219] Alternatively, for example, the candidates may include SbTMVP candidates. For example, if the selected candidate is an SbTMVP candidate, L0 motion information and L1 motion information can be derived based on the surrounding blocks to the left of the current block, and based on this, the L0 prediction sample and the L1 prediction sample can be generated. For example, the L0 prediction sample and the L1 prediction sample can be generated by performing predictions on a subblock basis.

[0220] For example, L0 motion information may include the index of the L0 reference picture and the L0 motion vector, and L1 motion information may include the index of the L1 reference picture and the L1 motion vector. The index of the L0 reference picture may include information indicating the reference picture in reference picture list 0, and the index of the L1 reference picture may include information indicating the reference picture in reference picture list 1.

[0221] For example, an encoding device can generate prediction samples for the current block based on L0 prediction samples, L1 prediction samples, and weighted value information. For example, the weighted value information can be indicated based on weighted value index information. The weighted value index information can indicate weighted value index information for biprediction. For example, the weighted value information can include information for a weighted average of L0 prediction samples or L1 prediction samples. That is, the weighted value index information can indicate index information for the weights used in the weighted average, and the weighted value index information can also be generated in a procedure for generating prediction samples based on the weighted average. For example, the weighted value index information can include information indicating the weights of any three or five weights. For example, the weighted average can indicate a weighted average in BCW (Bi-prediction with CU-level Weight) or BWA (Bi-prediction with Weighted Average).

[0222] For example, the candidate may include a temporal merge candidate, and the weighted index information may be indicated as 0. That is, the weighted index information for the temporal merge candidate may be indicated as 0. Here, a weighted index information of 0 indicates that the weights for each reference direction (i.e., the L0 prediction direction and the L1 prediction direction in biprediction) are the same. Alternatively, for example, the candidate may include a temporal merge candidate, and the weighted index information may be indicated based on the weighted index information of a col block. That is, the weighted index information for the temporal merge candidate may be indicated based on the weighted index information of a col block. Here, the col block may include a block at the same position as the current block in a reference picture different from the current picture in which the current block is located.

[0223] Alternatively, for example, the candidates may include pairwise candidates, and the weighted index information may be represented by the weighted index information of one of the other two candidates in the merge candidate list used to derive the pairwise candidate. That is, the weighted index information for the pairwise candidate may be represented by the weighted index information of one of the other two candidates in the merge candidate list used to derive the pairwise candidate. Alternatively, for example, the weighted index information may be represented based on the weighted index information of the two candidates.

[0224] Alternatively, for example, the merge candidate list may be a list of subblock merge candidates, and affine merge candidates, subblock merge candidates, or SbTMVP candidates may be selected. Here, affine merge candidates at the subblock level are sometimes called subblock merge candidates.

[0225] For example, the candidates may include merge candidates for subblocks, and the weighted index information may be indicated based on the weighted index information of a specific block among the surrounding blocks of the current block's CP. That is, the weighted index information for the merge candidates for subblocks may be indicated based on the weighted index information of a specific block among the surrounding blocks of the current block's CP. Here, the specific block may be a block used for deriving the CPMV for the CP, or it may be a block among the surrounding blocks of the current block's CP that has an MV used as the CPMV.

[0226] For example, the CP may be CP0, located on the upper left side of the current block. In this case, the weighted index information can be shown based on the weighted index information of the surrounding block at the upper left corner of the current block, the weighted index information of the surrounding block to the left adjacent to the lower side of the surrounding block at the upper left corner, or the weighted index information of the surrounding block to the upper side adjacent to the right side of the surrounding block at the upper left corner. Alternatively, the weighted index information can be shown in Figure 8 based on the weighted index information of block A2, the weighted index information of block B2, or the weighted index information of block B3.

[0227] Alternatively, for example, the CP may be CP1 located to the upper right of the current block. In this case, the weighted index information can be shown based on the weighted index information of the surrounding blocks at the upper right corner of the current block, or the weighted index information of the upper surrounding block adjacent to the left of the surrounding block at the upper right corner. Alternatively, the weighted index information can be shown in Figure 8 based on the weighted index information of block B0, or the weighted index information of block B1.

[0228] Alternatively, for example, the CP may be CP2 located to the lower left of the current block. In this case, the weighted index information can be shown based on the weighted index information of the surrounding block at the lower left corner of the current block, or the weighted index information of the surrounding block to the left adjacent to the upper side of the surrounding block at the lower left corner. Alternatively, the weighted index information can be shown in Figure 8 based on the weighted index information of block A0, or the weighted index information of block A1.

[0229] Alternatively, for example, the CP could be CP3 located to the lower right of the current block. Here, CP3 is sometimes called RB. In this case, the weighted index information can be shown based on the weighted index information of the col block of the current block, or the weighted index information of the surrounding blocks at the lower right corner of the col block. Here, the col block may contain a block at the same position as the current block in a reference picture different from the current picture in which the current block is located. Alternatively, the weighted index information can be shown based on the weighted index information of the T block in Figure 8.

[0230] Alternatively, for example, the CP may include multiple CPs. For example, the multiple CPs may include at least two of CP0, CP1, CP2, or RB. In this case, the weighted index information may be shown based on the weighted index information that most frequently overlaps among the weighted index information of the specific block used to derive each of the CPMVs. Alternatively, the weighted index information may be shown based on the weighted index information that occurs most frequently among the weighted index information of the specific block. That is, the weighted index information may be shown based on the weighted index information of the specific block used to derive the CPMV for each of the multiple CPs.

[0231] Alternatively, for example, the candidate may include an SbTMVP candidate, and the weighted index information may be indicated based on the weighted index information of the surrounding block to the left of the current block. That is, the weighted index information for the SbTMVP candidate may be indicated based on the weighted index information of the surrounding block to the left. Alternatively, for example, the candidate may include an SbTMVP candidate, and the weighted index information may be indicated as 0. That is, the weighted index information for the SbTMVP candidate may be indicated as 0. Here, weighted index information of 0 can indicate that the weights for each reference direction (i.e., the L0 prediction direction and the L1 prediction direction in biprediction) are the same. Alternatively, for example, the candidate may include an SbTMVP candidate, and the weighted index information may be indicated based on the weighted index information of the center block in the col block. That is, the weighted index information for the SbTMVP candidate may be indicated based on the weighted index information of the center block in the col block. Here, the col block may include a block at the same position as the current block in a reference picture different from the current picture in which the current block is located, and the center block may include the lower right subblock of the four subblocks located in the center of the col block. Alternatively, for example, the candidate may include an SbTMVP candidate, and the weighted index information may be shown based on the weighted index information of each subblock of the col block. That is, the weighted index information for an SbTMVP candidate may be shown based on the weighted index information of each subblock of the col block.

[0232] Alternatively, although not shown in Figure 10, for example, the encoding device can derive a residual sample based on the predicted sample and the original sample. In this case, residual-related information can be derived based on the residual sample. A residual sample can be derived based on the residual-related information. A restored sample can be generated based on the residual sample and the predicted sample. A restored block and a restored picture can be derived based on the restored sample. Alternatively, for example, the encoding device can encode video information that includes residual-related information or prediction-related information.

[0233] For example, an encoding device can encode video information containing all or part of the aforementioned information (or syntax elements) and generate a bitstream or encoded information. Alternatively, it can output it in the form of a bitstream. The bitstream or encoded information can be transmitted to a decoding device via a network or storage medium. Alternatively, the bitstream or encoded information can be stored on a computer-readable storage medium, and the bitstream or encoded information can be generated by the video encoding method described above.

[0234] Figures 12 and 13 schematically show an example of a video / image decoding method and related components according to the embodiment of this document.

[0235] The method disclosed in Figure 12 can be performed by the decoding apparatus disclosed in Figure 3 or Figure 13. Specifically, for example, S1200 in Figure 12 can be performed by the entropy decoding unit 310 of the decoding apparatus 300 in Figure 13, and S1210 to S1260 in Figure 12 can be performed by the prediction unit 330 of the decoding apparatus 300 in Figure 13. Although not shown in Figure 12, in Figure 13, the entropy decoding unit 310 of the decoding apparatus 300 can derive prediction-related information or residual information from the bitstream, the residual processing unit 320 of the decoding apparatus 300 can derive residual samples from the residual information, the prediction unit 330 of the decoding apparatus 300 can derive prediction samples from prediction-related information, and the addition unit 340 of the decoding apparatus 300 can derive a restored block or restored picture from the residual sample or predicted sample. The method disclosed in Figure 12 may include the embodiments described above in this document.

[0236] Referring to Figure 12, the decoding device can receive video information including interprediction mode information and interprediction type information via a bitstream (S1200). For example, the video information may also be called video information. The video information may include various types of information relating to the embodiments described above in this document. For example, the video information may include at least some of the prediction-related information or the residual-related information.

[0237] For example, the prediction-related information may include interpretation mode information or interpretation type information. For example, the interpretation mode information may include information indicating at least some of the various interpretation modes. For example, various modes such as merge mode, skip mode, MVP (motion vector prediction) mode, affine mode, subblock merge mode, or MMVD (merge with MVD) mode may be used. In addition, DMVR (Decoder side motion vector refinement) mode, AMVR (adaptive motion vector resolution) mode, BCW (Bi-prediction with CU-level weight), or BDOF (Bi-directional optical flow) may be used as incidental modes, either further or alternatively. For example, the interpretation type information may include the syntax element inter_pred_idc. Alternatively, the interpretation type information may include information indicating either L0 prediction, L1 prediction, or bi-prediction.

[0238] The decoding device can generate a merge candidate list for the current block based on the inter-prediction mode information (S1210). For example, based on the inter-prediction mode information, the decoding device can determine the inter-prediction mode of the current block to be a merge mode, an affine (merge) mode, or a sub-block merge mode, and generate a merge candidate list based on the determined inter-prediction mode. Here, if the inter-prediction mode is determined to be an affine merge mode or a sub-block merge mode, the merge candidate list may be called an affine merge candidate list or a sub-block merge candidate list, etc., but is sometimes simply called a merge candidate list.

[0239] For example, candidates can be inserted into the merge candidate list until the number of candidates in the merge candidate list reaches the maximum number of candidates. Here, a candidate can indicate a candidate or a candidate block for deriving motion information (or a motion vector) of a current block. For example, a candidate block can be derived through a search for neighboring blocks of the current block. For example, the neighboring blocks can include spatial neighboring blocks and / or temporal neighboring blocks of the current block. The spatial neighboring blocks are preferentially searched, (spatial merge) candidates can be derived, and then the temporal neighboring blocks are searched, and can be derived as (temporal merge) candidates. The derived candidates can be inserted into the merge candidate list. For example, if the number of candidates in the merge candidate list is less than the maximum number of candidates even after inserting the candidates, additional candidates can be inserted. For example, the additional candidates can include at least one of a history based merge candidate(s), a pair-wise average merge candidate(s), an ATMVP, a combined bi-predictive merge candidate (when the type of the slice / tile group of the current slice / tile group is of type B), and / or a zero vector merge candidate.

[0240] Alternatively, for example, candidates can be inserted into the affine merge candidate list until the number of candidates in the affine merge candidate list reaches the maximum number of candidates. Here, a candidate can include a CPMV (Control Point Motion Vector) of a current block. Alternatively, the candidate can also indicate a candidate or a candidate block for deriving the CPMV. The CPMV can indicate a motion vector at a CP (Control Point) of a current block. For example, the number of CPs can be 2, 3, or 4, and can be located at at least a part of the upper left side (or upper left corner), upper right side (or upper right corner), lower left side (or lower left corner), or lower right side (or lower right corner) of the current block, and only one CP can exist for each (position).

[0241] For example, candidate blocks can be derived through searching for surrounding blocks of the current block (or surrounding blocks of the CP of the current block). For example, the affine merge candidate list can include at least one of inherited affine merge candidates, constructed affine merge candidates, or zero motion vector candidates. For example, the affine merge candidate list can first insert the inherited affine merge candidates, and then can insert the constructed affine merge candidates. Also, even if constructed affine merge candidates are inserted into the affine merge candidate list, if the number of candidates in the affine merge candidate list is smaller than the maximum number of candidates, the remainder can be filled with zero motion vector candidates. Here, zero motion vector candidates may also be called zero vectors. For example, the affine merge candidate list can be a list in the affine merge mode in which motion vectors are derived in sample units, but can also be a list in the affine merge mode in which motion vectors are derived in sub-block units. In this case, the affine merge candidate list may also be called the merge candidate list of sub-blocks, and the merge candidate list of sub-blocks may also include candidates derived as SbTMVP (or SbTMVP candidates). For example, if SbTMVP candidates are included in the merge candidate list of sub-blocks, they may be located before the inherited affine merge candidates and the constructed affine merge candidates in the merge candidate list of sub-blocks.

[0242] The decoding device can select one of the candidates included in the merge candidate list (S1220). For example, the merge candidate list may include at least some of spatial merge candidates, temporal merge candidates, pairwise candidates, or zero vector candidates, and one of these candidates can be selected for inter prediction of the current block. Alternatively, for example, the subblock merge candidate list may include at least some of inherited affine merge candidates, configured affine merge candidates, SbTMVP candidates, or zero vector candidates, and one of these candidates can be selected for inter prediction of the current block. For example, the selected candidate can be selected from the merge candidate list based on selection information. For example, the selection information may include index information indicating the selected candidate in the merge candidate list. For example, the selection information may also be called merge index information or subblock merge index information. For example, the selection information may be included in the video information. Alternatively, the selection information may be included in the inter prediction mode information.

[0243] The decoding device can derive the inter-prediction type of the current block as a bi-prediction based on the inter-prediction type information (S1230). For example, the inter-prediction type of the current block can be derived as a bi-prediction from among L0 prediction, L1 prediction, or bi-prediction based on the inter-prediction type information. Here, L0 prediction can represent a prediction based on reference picture list 0, L1 prediction can represent a prediction based on reference picture list 1, and bi-prediction can represent a prediction based on both reference picture list 0 and reference picture list 1. For example, the inter-prediction type information may include the syntax element inter_pred_idc.

[0244] The decoding device can derive motion information for the current block based on the selected candidate (S1240). For example, the decoding device can derive L0 motion information and L1 motion information based on the selected candidate by deriving the interpretation type as biprediction. For example, the L0 motion information may include the index of the L0 reference picture and the L0 motion vector, and the L1 motion information may include the index of the L1 reference picture and the L1 motion vector. The index of the L0 reference picture may include information indicating the reference picture in reference picture list 0, and the L1 reference picture index may include information indicating the reference picture in reference picture list 1.

[0245] The decoding device can generate L0 and L1 prediction samples of the current block based on motion information (S1250). For example, if the interpretation type of the current block is derived as biprediction, reference picture list 0 and reference picture list 1 may be used for the prediction of the current block. For example, the L0 prediction sample may represent the prediction sample of the current block derived based on reference picture list 0, and the L1 prediction sample may represent the prediction sample of the current block derived based on reference picture list 1.

[0246] For example, the candidates may include spatial merge candidates. For example, if the selected candidate is a spatial merge candidate, L0 motion information and L1 motion information can be derived based on the spatial merge candidate, and based on this, the L0 prediction sample and the L1 prediction sample can be generated.

[0247] For example, the candidates may include temporal merge candidates. For example, if the selected candidate is a temporal merge candidate, L0 motion information and L1 motion information can be derived based on the temporal merge candidate, and based on this, the L0 prediction sample and the L1 prediction sample can be generated.

[0248] For example, the candidates may include pairwise candidates. For example, if the selected candidate is a pairwise candidate, L0 motion information and L1 motion information can be derived based on the pairwise candidate, and the L0 prediction sample and the L1 prediction sample can be generated based on this. For example, the pairwise candidate may be derived based on two other candidates from among the candidates included in the merge candidate list.

[0249] Alternatively, for example, the merge candidate list may be a list of subblock merge candidates, and affine merge candidates, subblock merge candidates, or SbTMVP candidates may be selected. Here, affine merge candidates at the subblock level are sometimes called subblock merge candidates.

[0250] For example, the candidates may include subblock merge candidates. For example, if the selected candidate is a subblock merge candidate, L0 motion information and L1 motion information can be derived based on the subblock merge candidate, and the L0 prediction sample and the L1 prediction sample can be generated based on this. For example, the subblock merge candidate may include a CPMV (Control Point Motion Vector), and the L0 prediction sample and the L1 prediction sample can be generated by performing predictions on a subblock basis based on the CPMV.

[0251] Here, CPMV can be derived based on one of the surrounding blocks of the current block's CP (Control Point). For example, there may be two, three, or four CPs, and they may be located in at least part of the upper left side (or upper left corner), upper right side (or upper right corner), lower left side (or lower left corner), or lower right side (or lower right corner) of the current block, with only one CP at each location.

[0252] For example, CP could be CP0, located to the upper left of the current block. In this case, the surrounding blocks could include the surrounding block at the upper left corner of the current block, the surrounding block to the left adjacent to the lower side of the surrounding block at the upper left corner, and the surrounding block to the upper side adjacent to the right side of the surrounding block at the upper left corner. Alternatively, the surrounding blocks could include blocks A2, B2, or B3 in Figure 8.

[0253] Alternatively, for example, the CP could be CP1 located to the upper right of the current block. In this case, the surrounding block could include the surrounding block at the upper right corner of the current block and the upper surrounding block adjacent to the left of the surrounding block at the upper right corner. Alternatively, the surrounding block could include block B0 or ​​block B1 in Figure 8.

[0254] Alternatively, for example, CP could be CP2 located to the lower left of the current block. In this case, the surrounding block could include the surrounding block at the lower left corner of the current block and the surrounding block to the left adjacent to the upper side of the surrounding block at the lower left corner. Alternatively, the surrounding block could include block A0 or block A1 in Figure 8.

[0255] Alternatively, for example, the CP could be CP3 located to the lower right of the current block. Here, CP3 is sometimes called RB. In this case, the surrounding block may include the col block of the current block or the surrounding block at the lower right corner of the col block. Here, the col block may include a block at the same position as the current block in a reference picture different from the current picture in which the current block is located. Alternatively, the surrounding block may include a T block in Figure 8.

[0256] Alternatively, for example, the candidates may include SbTMVP candidates. For example, if the selected candidate is an SbTMVP candidate, L0 motion information and L1 motion information can be derived based on the surrounding blocks to the left of the current block, and based on this, the L0 prediction sample and the L1 prediction sample can be generated. For example, the L0 prediction sample and the L1 prediction sample can be generated by performing predictions on a subblock basis.

[0257] The decoding device can generate prediction samples for the current block based on L0 prediction samples, L1 prediction samples, and weighted value information (S1260). For example, the weighted value information can be derived based on weighted value index information. For example, the weighted value information may include information for a weighted average of L0 prediction samples or L1 prediction samples. That is, the weighted value index information may indicate index information for the weights used in the weighted average, and the weighted average may be performed based on the weighted value index information. For example, the weighted value index information may include information indicating the weights of three or five weights. For example, the weighted average may indicate a weighted average using BCW (Bi-prediction with CU-level Weight) or BWA (Bi-prediction with Weighted Average).

[0258] For example, the candidate may include a temporal merge candidate, and the weighted index information may be derived as 0. That is, the weighted index information for a temporal merge candidate may be derived as 0. Here, a weighted index information of 0 indicates that the weights for each reference direction (i.e., the L0 prediction direction and the L1 prediction direction in biprediction) are the same. Alternatively, for example, the candidate may include a temporal merge candidate, and the weighted index information may be derived based on the weighted index information of a col block. That is, the weighted index information for a temporal merge candidate may be derived based on the weighted index information of a col block. Here, the col block may include a block at the same position as the current block in a reference picture different from the current picture in which the current block is located.

[0259] Alternatively, for example, the candidates may include pairwise candidates, and the weighted index information may be derived as the weighted index information of one of the other two candidates in the merge candidate list used to derive the pairwise candidate. That is, the weighted index information for the pairwise candidate may be derived as the weighted index information of one of the other two candidates in the merge candidate list used to derive the pairwise candidate. Alternatively, for example, the weighted index information may be derived based on the weighted index information of the two candidates.

[0260] Alternatively, for example, the merge candidate list may be a list of subblock merge candidates, and affine merge candidates, subblock merge candidates, or SbTMVP candidates may be selected. Here, affine merge candidates at the subblock level are sometimes called subblock merge candidates.

[0261] For example, the candidate may include a merge candidate for a sub-block, and the weighted value index information can be derived based on the weighted value index information of a specific block among the peripheral blocks of the CP of the current block. That is, the weighted value index information for the merge candidate of the sub-block can be derived based on the weighted value index information of a specific block among the peripheral blocks of the CP of the current block. Here, the specific block may be a block used for the derivation of CPMV with respect to the CP. Alternatively, it may be a block having an MV used as CPMV among the peripheral blocks of the CP of the current block.

[0262] For example, the CP may be CP0 located at the upper left side of the current block. In this case, the weighted value index information can be derived based on the weighted value index information of the peripheral blocks at the upper left corner of the current block, the weighted value index information of the left peripheral block adjacent to the lower side of the peripheral blocks at the upper left corner, or the weighted value index information of the upper peripheral block adjacent to the right side of the peripheral blocks at the upper left corner. Alternatively, the weighted value index information can be derived based on the weighted value index information of block A2, block B2, or block B3 in FIG. 8.

[0263] Alternatively, for example, the CP may be CP1 located at the upper right side of the current block. In this case, the weighted value index information can be derived based on the weighted value index information of the peripheral blocks at the upper right corner of the current block or the weighted value index information of the upper peripheral block adjacent to the lower side of the peripheral blocks at the upper right corner. Alternatively, the weighted value index information can be derived based on the weighted value index information of block B0 or block B1 in FIG. 8.

[0264] Alternatively, for example, the CP could be CP2 located to the lower left of the current block. In this case, the weighted index information can be derived based on the weighted index information of the surrounding block at the lower left corner of the current block or the weighted index information of the surrounding block to the left adjacent to the upper side of the surrounding block at the lower left corner. Alternatively, the weighted index information can be derived based on the weighted index information of block A0 or block A1 in Figure 8.

[0265] Alternatively, for example, the CP could be CP3 located to the lower right of the current block. Here, CP3 is sometimes called RB. In this case, the weighted index information can be derived based on the weighted index information of the col block of the current block or the weighted index information of the surrounding blocks at the lower right corner of the col block. Here, the col block may contain a block at the same position as the current block in a reference picture different from the current picture in which the current block is located. Alternatively, the weighted index information can be derived based on the weighted index information of the T block in Figure 8.

[0266] Alternatively, for example, the CP may include multiple CPs. For example, the multiple CPs may include at least two of CP0, CP1, CP2, or RB. In this case, the weighted index information can be derived based on the weighted index information of the specific block that has the most overlap among the weighted index information of the specific block used to derive each of the CPMVs. Alternatively, the weighted index information can be derived based on the weighted index information of the specific block that occurs most frequently. That is, the weighted index information can be derived based on the weighted index information of the specific block used to derive the CPMV for each of the multiple CPs.

[0267] Alternatively, for example, the candidate may include an SbTMVP candidate, and the weighted index information may be derived based on the weighted index information of the surrounding block to the left of the current block. That is, the weighted index information for the SbTMVP candidate may be derived based on the weighted index information of the surrounding block to the left. Alternatively, for example, the candidate may include an SbTMVP candidate, and the weighted index information may be derived as 0. That is, the weighted index information for the SbTMVP candidate may be derived as 0. Here, weighted index information of 0 can indicate that the weights for each reference direction (i.e., the L0 prediction direction and the L1 prediction direction in biprediction) are the same. Alternatively, for example, the candidate may include an SbTMVP candidate, and the weighted index information may be derived based on the weighted index information of the center block in the col block. That is, the weighted index information for the SbTMVP candidate may be derived based on the weighted index information of the center block in the col block. Here, the col block may include a block at the same position as the current block in a reference picture different from the current picture in which the current block is located, and the center block may include the lower right subblock of the four subblocks located in the center of the col block. Alternatively, for example, the candidate may include an SbTMVP candidate, and the weighted index information may be derived based on the weighted index information of each subblock of the col block. That is, the weighted index information for the SbTMVP candidate may be derived based on the weighted index information of each subblock of the col block.

[0268] Although not shown in Figure 12, for example, the decoding device can derive a residual sample based on residual-related information contained in the video information. Furthermore, the decoding device can generate a restored sample based on the predicted sample and the residual sample. Based on the restored sample, a restored block and a restored picture can be derived.

[0269] For example, a decoding device can decode a bitstream or encoded information to obtain video information that includes all or part of the aforementioned information (or syntax elements). Furthermore, the bitstream or encoded information can be stored on a computer-readable storage medium, which can trigger the aforementioned decoding method.

[0270] In the embodiments described above, the method is explained based on a flowchart as a series of steps or blocks, but the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and that different steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments described herein.

[0271] The methods relating to the embodiments described in this document above can be implemented in software form, and the encoding and / or decoding devices relating to this document may be included in, for example, video processing devices such as TVs, computers, smartphones, set-top boxes, and display devices.

[0272] In this document, when embodiments are implemented in software, the methods described above can be implemented by modules (processes, functions, etc.) that perform the functions described above. These modules are stored in memory and can be executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. The processor may include an ASIC (application-specific integrated circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document may be implemented on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing may be implemented on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.

[0273] Furthermore, the decoding and encoding devices to which the embodiments described in this document apply may include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, customized video (VoD) service providers, OTT video (Over the Top Video) devices, internet streaming service providers, 3D video devices, VR (virtual reality) devices, AR (argumente reality) devices, video telephone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video equipment, and may be used to process video signals or data signals. For example, OTT video (Over the Top Video) devices may include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorders), etc.

[0274] Furthermore, the processing methods to which the embodiments of this document apply can be produced in the form of programs executed by a computer and stored on a computer-readable recording medium. Multimedia data having the data structure relating to the embodiments of this document can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store data to be read by a computer. The computer-readable recording medium may include, for example, Blu-ray discs (BDs), general-purpose serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission over the Internet). Furthermore, a bitstream generated by an encoding method can be stored on a computer-readable recording medium or transmitted over a wired wireless network.

[0275] Furthermore, the embodiments described in this document can be embodied in a computer program product using program code, and the program code can be executed on a computer according to the embodiments described in this document. The program code can be stored on a computer-readable carrier.

[0276] Figure 14 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.

[0277] Referring to Figure 14, the content streaming system to which the embodiments of this document apply can broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.

[0278] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then transmitting this bitstream to the streaming server. As an alternative, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server may be omitted.

[0279] The bitstream can be generated by an encoding method or bitstream generation method to which the embodiments of this document apply, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0280] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server acts as an intermediary to inform users about available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.

[0281] The streaming server can receive content from a media storage and / or encoding server. For example, if it begins receiving content from the encoding server, it can receive the content in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0282] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, and HMDs), digital TVs, desktop computers, and digital signage.

[0283] Each server within the aforementioned content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.

[0284] The claims described herein can be combined in various ways. For example, the technical features of the method claims herein can be combined to embody an apparatus, and the technical features of the apparatus claims herein can be combined to embody a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims herein can be combined to embody an apparatus, and the technical features of the method claims and the technical features of the apparatus claims herein can be combined to embody a method.

Claims

1. In a video decoding method performed by a decoding device, The steps include receiving video information including interpredictive mode information and residual information via a bitstream, The steps include generating a list of merge candidates for the current block based on the aforementioned inter-prediction mode information, The steps include selecting a candidate from among the candidates included in the merge candidate list, The steps include: deriving the movement information of the current block based on the selected candidate; The steps include generating L0 prediction samples and L1 prediction samples based on the derived motion information, A step of generating a prediction sample for the current block based on the L0 prediction sample, the L1 prediction sample, and weighted value information, wherein the weighted value information is derived based on weighted value index information for the selected candidate. The steps include: deriving a conversion coefficient based on the residual information; The steps include: deriving a residual sample by inverse quantizing the aforementioned conversion coefficients; The steps include generating a reconstructed picture based on the predicted sample and the residual sample, The aforementioned candidates include inherited affine merge candidates and configured affine merge candidates, The inherited affine merge candidate is derived based on the CPMV (Control Point Motion Vector) of the surrounding blocks of the current block. The affine merge candidate configured above includes the CPMV of the control point (CP), Based on the fact that the configured affine merge candidate includes CPMV for CP0 and that the biprediction is applied to the current block, the weighted index information for the configured affine merge candidate is fixed to be equal to the weighted index information of a specific block among the surrounding blocks of CP0 in the current block, where CP0 is related to the upper left corner of the current block. The aforementioned particular block is a block used to derive the CPMV for the CP0, The configured affine merge candidate is inserted after the inherited affine merge candidate in the merge candidate list. The configured affine merge candidate corresponds to a 6-parameter affine merge candidate based on a combination of the CPMV of three control points. The CPMV of the control points used for the configured affine merge candidate is derived based on the same reference picture. A method in which, after inserting the configured affine merge candidates, zero motion vector candidates are inserted into the merge candidate list based on the fact that the number of candidates in the affine merge candidate list is less than the maximum number of candidates.

2. In a video encoding method performed by an encoding device, The steps include determining the inter-prediction mode of the current block and generating inter-prediction mode information indicating the inter-prediction mode, The steps include generating a list of merge candidates for the current block based on the inter prediction mode, The steps include selecting one of the candidates included in the merge candidate list and generating selection information indicating the selected candidate, The steps include: deriving residual samples based on the predicted samples associated with the interprediction mode; The process includes a step of encoding video information including the interprediction mode information, the residual information associated with the residual sample, and the selection information, The aforementioned candidates include inherited affine merge candidates and configured affine merge candidates, The inherited affine merge candidate is derived based on the CPMV (Control Point Motion Vector) of the surrounding blocks of the current block. The affine merge candidate configured above includes the CPMV of the control point (CP), Based on the fact that the configured affine merge candidate includes CPMV for CP0 and that the biprediction is applied to the current block, the weighted index information for the configured affine merge candidate is fixed to be equal to the weighted index information of a specific block among the surrounding blocks of CP0 in the current block, where CP0 is related to the upper left corner of the current block. The aforementioned particular block is a block used to derive the CPMV for the CP0, The configured affine merge candidate is inserted after the inherited affine merge candidate in the merge candidate list. The configured affine merge candidate corresponds to a 6-parameter affine merge candidate based on a combination of the CPMV of three control points. The CPMV of the control points used for the configured affine merge candidate is derived based on the same reference picture. A method in which, after inserting the configured affine merge candidates, zero motion vector candidates are inserted into the merge candidate list based on the fact that the number of candidates in the affine merge candidate list is less than the maximum number of candidates.

3. A method for transmitting video data, A step in which a transmitting device acquires encoded video information, wherein the encoded video information is The steps include determining the inter-prediction mode of the current block and generating inter-prediction mode information indicating the inter-prediction mode, The steps include generating a list of merge candidates for the current block based on the inter prediction mode, The steps include selecting one of the candidates included in the merge candidate list and generating selection information indicating the selected candidate, The steps include: deriving residual samples based on the predicted samples associated with the interprediction mode; A step of encoding video information including the interprediction mode information, the residual information associated with the residual sample, and the selection information, which is generated by performing the following steps: The steps include: the transmitting device transmitting the video data relating to the encoded video information, The aforementioned candidates include inherited affine merge candidates and configured affine merge candidates, The inherited affine merge candidate is derived based on the CPMV (Control Point Motion Vector) of the surrounding blocks of the current block. The affine merge candidate configured above includes the CPMV of the control point (CP), Based on the fact that the configured affine merge candidate includes CPMV for CP0 and that the biprediction is applied to the current block, the weighted index information for the configured affine merge candidate is fixed to be equal to the weighted index information of a specific block among the surrounding blocks of CP0 in the current block, where CP0 is related to the upper left corner of the current block. The aforementioned particular block is a block used to derive the CPMV for the CP0, The configured affine merge candidate is inserted after the inherited affine merge candidate in the merge candidate list. The configured affine merge candidate corresponds to a 6-parameter affine merge candidate based on a combination of the CPMV of three control points. The CPMV of the control points used for the configured affine merge candidate is derived based on the same reference picture. A method in which, after inserting the configured affine merge candidates, zero motion vector candidates are inserted into the merge candidate list based on the fact that the number of candidates in the affine merge candidate list is less than the maximum number of candidates.

Citation Information

Patent Citations

  • Affine motion prediction for video coding

    US20170332095A1

  • Motion vector prediction for affine motion models in video coding

    US20180098063A1

  • Method and apparatus for affine inter prediction for video coding system

    US20190028731A1