Inter prediction-based image decoding method and device therefor

By using regression-based affine candidate derivation for inter prediction, the method enhances video/image compression efficiency and prediction accuracy, addressing the challenges of high-resolution and immersive media data size and cost issues.

WO2025150796A1PCT designated stage expired Publication Date: 2025-07-17LX SEMICON CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/000182
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-03
Filing Date
2025-01-03
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images/videos, particularly in immersive media applications like VR, AR, and holograms, has led to higher data sizes and increased transmission and storage costs, necessitating a more efficient image/video compression technology.

Method used

The implementation of a regression-based affine candidate derivation method for inter prediction in video/image coding, which involves deriving motion information of sub-blocks based on surrounding blocks to enhance prediction accuracy and efficiency.

Benefits of technology

This approach improves video/image compression efficiency and prediction performance by accurately determining motion information of sub-blocks, reducing data size and transmission/storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025000182_17072025_PF_FP_ABST
    Figure KR2025000182_17072025_PF_FP_ABST
Patent Text Reader

Abstract

An image decoding method according to an embodiment of the present disclosure includes the steps of: deriving a regression-based affine candidate of a current block on the basis of neighboring sub-blocks of the current block; configuring an affine merge candidate list including the regression-based affine candidate of the current block; deriving motion information of sub-blocks of the current block on the basis of the affine merge candidate list; and deriving prediction samples for the current block on the basis of the motion information of the sub-blocks.
Need to check novelty before this filing date? Find Prior Art

Description

Inter-prediction-based image decoding method and device therefor

[0001] This document relates to a method and device for coding images / videos.

[0002] Image / video coding is used in various applications such as digital storage media, television broadcasting, video streaming services, and real-time communications, and the demand for high-resolution, high-quality images / videos is increasing in various fields.

[0003] As the image / video becomes higher resolution and higher quality, the data size of the image / video increases, and the amount of information or bits transmitted increases relatively. Therefore, when transmitting image data using media such as existing wired or wireless broadband lines or storing image / video data using existing storage media, the transmission and storage costs increase.

[0004] In addition, interest in and demand for immersive media such as VR (virtual reality), AR (artificial reality), MR (mixed reality) content and holograms have been increasing recently, and attempts to provide immersive experiences using immersive media in games, education, medicine, real estate, marketing, etc. are increasing.

[0005] Accordingly, a highly efficient image / video compression technology is required to effectively compress, transmit, store, and play high-resolution, high-quality image / video information having various characteristics as described above.

[0006] According to one embodiment of the present disclosure, a method and device for improving video / image coding efficiency are provided.

[0007] According to one embodiment of the present disclosure, an inter prediction based video / image coding method and device are provided.

[0008] According to one embodiment of the present disclosure, a video decoding method performed by a decoding device is provided. The method is characterized by including the steps of: deriving a regression-based affine candidate of a current block based on surrounding sub-blocks of the current block; constructing an affine merge candidate list including the regression-based affine candidate of the current block; deriving motion information of sub-blocks of the current block based on the affine merge candidate list; and deriving prediction samples for the current block based on the motion information of the sub-blocks.

[0009] According to one embodiment of the present disclosure, a video encoding method performed by an encoding device is provided. The method is characterized by including the steps of: deriving a regression-based affine candidate of a current block based on surrounding sub-blocks of the current block; constructing an affine merge candidate list including the regression-based affine candidate of the current block; deriving motion information of sub-blocks of the current block based on the affine merge candidate list; deriving prediction samples for the current block based on the motion information of the sub-blocks; and encoding image information including prediction-related information of the current block.

[0010] According to one embodiment of the present disclosure, a decoding device for image decoding is provided. The decoding device includes a memory and at least one processor connected to the memory, and the at least one processor is configured to perform a step of deriving a regression-based affine candidate of a current block based on surrounding sub-blocks of the current block, a step of constructing an affine merge candidate list including the regression-based affine candidate of the current block, a step of deriving motion information of sub-blocks of the current block based on the affine merge candidate list, and a step of deriving prediction samples for the current block based on the motion information of the sub-blocks.

[0011] According to one embodiment of the present disclosure, an encoding device for video encoding is provided. The encoding device includes a memory and at least one processor connected to the memory, and the at least one processor is configured to perform a step of deriving a regression-based affine candidate of a current block based on surrounding sub-blocks of the current block, a step of constructing an affine merge candidate list including the regression-based affine candidate of the current block, a step of deriving motion information of sub-blocks of the current block based on the affine merge candidate list, a step of deriving prediction samples for the current block based on the motion information of the sub-blocks, and a step of encoding image information including prediction-related information of the current block.

[0012] According to one embodiment of the present disclosure, a method is provided for transmitting video / video data including a bitstream generated according to a video / video encoding method according to at least one of the embodiments of the present disclosure.

[0013] According to one embodiment of the present disclosure, a device is provided for transmitting video / video data including a bitstream generated according to a video / video encoding method according to at least one of the embodiments of the present disclosure.

[0014] According to one embodiment of the present disclosure, a computer-readable storage medium storing a program for performing a method according to at least one of the embodiments of the present disclosure may be provided.

[0015] According to one embodiment of the present disclosure, a computer-readable digital storage medium storing encoded video / video information generated by a video / video encoding method according to at least one of the embodiments of the present disclosure is provided.

[0016] According to one embodiment of the present disclosure, there is provided a computer-readable digital storage medium storing encoded information or encoded video / image information that causes a decoding device to perform a video / image decoding method according to at least one of the embodiments of the present disclosure.

[0017] According to one embodiment of the present disclosure, the overall video / image compression efficiency can be improved.

[0018] According to one embodiment of the present disclosure, prediction performance for a current block can be improved.

[0019] According to one embodiment of the present disclosure, in affine prediction, motion information of sub-block units of a current block can be derived by considering motion information of surrounding sub-blocks adjacent to the current block, thereby improving the accuracy and coding efficiency of affine prediction.

[0020] FIG. 1 schematically illustrates an example of a video / image coding system to which embodiments of the present disclosure may be applied.

[0021] FIG. 2 is a drawing schematically illustrating the configuration of a video / image encoding device to which embodiments of the present disclosure can be applied.

[0022] FIG. 3 is a drawing schematically illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure can be applied.

[0023] Figure 4 illustrates an example of an inter prediction procedure.

[0024] Figure 5 shows examples of inter prediction based video / image encoding methods.

[0025] Figure 6 shows examples of inter prediction based video / image decoding methods.

[0026] Figure 7 shows an example of a segmentation shape supported by GPM.

[0027] Figure 8 illustrates an intra prediction mode that can be used as an IPM candidate.

[0028] Figure 9 shows a reference template for deriving the TM cost, which is the MVD prediction cost of an MV candidate.

[0029] Figure 10 shows a reference template for deriving the TM cost, which is the MVD prediction cost of an MV candidate in the affine AMVP mode or the affine MMVD mode.

[0030] Figure 11 shows MMVD candidates that are combinations of the available signs and sizes.

[0031] Fig. 12 illustrates an exemplary flowchart of an affine motion prediction method according to one embodiment of the present document.

[0032] Figure 13 shows an example of constructing an affine merge candidate list of the current block.

[0033] Figure 14 illustrates the surrounding blocks of the current block for deriving the inherited affine candidate.

[0034] Figure 15 illustrates blocks surrounding the current block for deriving the constructed affine candidate.

[0035] Figure 16 shows the surrounding sub-blocks that are input into the linear regression process to derive a set of linear model parameters.

[0036] FIG. 17 schematically illustrates a video / image encoding method according to an embodiment(s) of the present disclosure.

[0037] FIG. 18 schematically illustrates a video / image decoding method according to an embodiment(s) of the present disclosure.

[0038] This disclosure is susceptible to various modifications and embodiments, and thus specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit the embodiments of the present disclosure to the specific embodiments. The terminology used herein is only used to describe specific embodiments and is not intended to limit the technical spirit of the present disclosure. The singular forms used herein are intended to include the plural forms as well, unless the context clearly indicates otherwise. The term “and / or” as used herein includes any one or a combination of two or more of the associated listed items. The terms “comprises,” “comprises,” and “contains” as used herein specify the presence of stated features, numbers, operations, elements, components, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, numbers, operations, elements, components, and / or combinations thereof. The use of the term “can” in connection with an example or embodiment (e.g., what the example or embodiment can include or implement) in this disclosure means that there is at least one example or embodiment that includes or implements such feature, but not all examples are limited thereto and such feature or configuration may be omitted.

[0039] Meanwhile, each component in the drawings described in this disclosure is depicted independently for the convenience of explaining different characteristic functions. This does not imply that each component is implemented with separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of the present disclosure, as long as they do not deviate from the essence of the present disclosure.

[0040] In this disclosure, “A or B” can mean “only A,” “only B,” or “both A and B.” In other words, “A or B” in this disclosure can be interpreted as “A and / or B.” For example, “A, B or C” in this disclosure can mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.”

[0041] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."

[0042] In the present disclosure, “at least one of A and B” may mean “only A,” “only B,” or “both A and B.” Additionally, in the present disclosure, the expressions “at least one of A or B” or “at least one of A and / or B” may be interpreted identically to “at least one of A and B.”

[0043] Additionally, in the present disclosure, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”

[0044] Additionally, parentheses used in this disclosure may mean "for example." Specifically, when "prediction (intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction." In other words, "prediction" in this disclosure is not limited to "intra-prediction," and "intra-prediction" may be suggested as an example of "prediction." Furthermore, even when indicated as "prediction (i.e., intra-prediction)", "intra-prediction" may be suggested as an example of "prediction."

[0045] Technical features individually described in one drawing in this disclosure may be implemented individually or simultaneously.

[0046] The present disclosure relates to video / image coding. For example, the methods / embodiments described in this disclosure may be applied to methods disclosed in the enhanced compression model (ECM) or H.267 standards. Furthermore, the methods / embodiments disclosed in this disclosure may be applied to methods disclosed in the AV2 (AOMedia Video 2) standard or next-generation video / image coding standards (e.g., H.268, H.269, etc.).

[0047] In the present disclosure, coding may include encoding and / or decoding. In the present disclosure, image coding may be used interchangeably with video coding.

[0048] In the present disclosure, a video may refer to a set of images over time. A picture generally refers to a unit representing one image at a specific time point, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile may include one or more CTUs (coding tree units). A picture may be composed of one or more slices / tiles. A tile may represent a rectangular area of ​​CTUs within a specific tile row and a specific tile column within a picture.

[0049] Meanwhile, a single picture may be divided into two or more subpictures. A subpicture may be a rectangular region of one or more slices within a picture.

[0050] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can also represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component.

[0051] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.

[0052] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the attached drawings. Hereinafter, identical reference numerals may be used for identical components in the drawings, and redundant descriptions of identical components may be omitted.

[0053] FIG. 1 schematically illustrates an example of a video / image coding system to which embodiments of the present disclosure may be applied.

[0054] Referring to FIG. 1, a video / image coding system may include a first device (encoding device) and a second device (decoding device). The first device may transmit encoded video / image information or data to the second device via a digital storage medium or a network in the form of a file or streaming.

[0055] The video / image coding system may further include a video / image acquisition device and a video / image renderer. The video / image acquisition device may be included in the encoding device, or may be configured as a separate device or external component. The video / image renderer may be included in the decoding device, or may be configured as a separate device or external component.

[0056] The first device may include the transmission unit as an internal component, or as a separate device or external component.

[0057] The second device may include the receiver as an internal component, or as a separate device or external component.

[0058] An encoder may be referred to as an encoding device, and a decoder may be referred to as a decoding device. A transmitting unit may be included in an encoding device. A receiving unit may be included in a decoding device. A renderer may include a display unit, and the display unit may be comprised of a separate device or an external component.

[0059] The decoding device and encoding device to which the embodiment(s) of the present disclosure are applied may be included in a multimedia broadcasting transmitting and receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (agumented reality) device, a video phone video device, a transportation terminal (e.g., a vehicle (including an autonomous vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, and may be used to process a video signal or a data signal. For example, the OTT (Over the top video) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.

[0060] A video / image capture device can capture a video / image source. The video / image capture device can capture the video / image through a process of capturing, synthesizing, or generating the video / image. The video / image capture device can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a camcorder, a computer, a tablet, a smartphone, etc., and can (electronically) generate the video / image. For example, a virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced by a process of generating related data. The video / image source can also perform a video / image preprocessing process to input optimized video / image to an encoder.

[0061] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0062] The transmission unit can transmit encoded video / image information or data output in bitstream form to the reception unit of the receiving device through the network in the form of a file or streaming. The encoded video / image information or data output in bitstream form can also be transmitted to the reception unit through a streaming server. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file through a predetermined file format and an element for transmission through a broadcasting / communication network. The reception unit can receive / extract the bitstream and transmit it to a decoding device.

[0063] The streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream. The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as a medium that informs the user of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server serves to control commands / responses between each device within the content streaming system.

[0064] The streaming server can receive content from a media storage device and / or an encoding device. For example, when receiving content from the encoding device, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0065] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.

[0066] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.

[0067] FIG. 2 is a diagram schematically illustrating the configuration of a video / image encoding device to which embodiments of the present disclosure may be applied. The term "encoding device" hereinafter may include an image encoding device and / or a video encoding device.

[0068] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit and an intra prediction unit. The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processor (230) may further include a subtractor (subtractor) 231. The addition unit (250) may be called a reconstruction unit or a reconstructed block generator. The image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. In addition, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.

[0069] The image segmentation unit (210) can segment an input image (or picture, frame) input to the encoding device (200) into one or more processing units. For example, the processing units may be referred to as coding units (CUs). In this case, the coding units may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a Quad-tree binary-tree ternary-tree (QTBTTT) structure. For example, one coding unit may be segmented into a plurality of coding units of deeper depth based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied first. The coding procedure according to the present disclosure may be performed based on the final coding unit that is no longer segmented. In this case, based on coding efficiency according to image characteristics, etc., the maximum coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units of lower depths, and the coding unit of the optimal size can be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described below. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be divided or partitioned from the final coding unit described above.The above prediction unit may be a unit for sample prediction, and the above transformation unit may be a unit for deriving a transformation coefficient and / or a unit for deriving a residual signal from a transformation coefficient.

[0070] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image).

[0071] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from a prediction unit from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, as illustrated, a unit that subtracts a prediction signal (predicted block, prediction sample array) from an input video signal (original block, original sample array) within the encoder (200) may be called a subtraction unit (231). The prediction unit can perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit can generate various information regarding prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit the information to the entropy encoding unit (240). The information regarding prediction can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0072] An intra prediction unit can predict a current block by referring to samples within a current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from the current block depending on the prediction mode. In intra prediction, prediction modes may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, a DC mode and a planar mode. Directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of detail in the prediction direction. However, this is only an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0073] An inter prediction unit can derive a predicted block for a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in an inter prediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on an inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. A reference picture including the reference block and a reference picture including the temporal neighboring blocks may be the same or different. The above temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and a reference picture including the temporal neighboring blocks may be called a collocated picture (colPic). For example, the inter prediction unit may construct a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit may use the motion information of the neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0074] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for screen content coding (SCC), for example. IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block based on a block vector within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in the present disclosure.

[0075] The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or a residual signal. The transformation unit (232) can apply a transformation technique to the residual signal to generate transform coefficients. For example, the transformation technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT).

[0076] The quantization unit (233) quantizes the transform coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit (240) can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (240) may encode information necessary for video / image restoration (e.g., values ​​of syntax elements, etc.) together or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. In the present disclosure, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream.The above bitstream may be transmitted through a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit (240) may be configured as an internal / external element of the encoding device (200) by a transmitting unit (not shown) and / or a storing unit (not shown), or the transmitting unit may be included in the entropy encoding unit (240).

[0077] The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the prediction unit. When there is no residual for the target block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit (250) may be called a reconstructor or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next target block to be processed within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.

[0078] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.

[0079] The filtering unit (260) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (270), specifically, in the DPB of the memory (270). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.

[0080] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit. Through this, the encoding device can avoid prediction mismatch between the encoding device (200) and the decoding device when inter prediction is applied, and can also improve encoding efficiency.

[0081] The memory (270) DPB can store the modified reconstructed picture to be used as a reference picture in the inter prediction unit. The memory (270) can store motion information of a block from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (270) can store reconstructed samples of reconstructed blocks within the current picture and transfer them to the intra prediction unit.

[0082] FIG. 3 is a diagram schematically illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure may be applied. The term "decoding device" hereinafter may include an image decoding device and / or a video decoding device.

[0083] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-prediction unit and an intra-prediction unit. The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) described above may be configured by a single hardware component (e.g., decoder chipset or processor) depending on the embodiment. In addition, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.

[0084] When a bitstream including video / image information is input, the decoding device (300) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Therefore, the processing unit of decoding may be, for example, a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output through the decoding device (300) can be reproduced through a reproduction device.

[0085] The decoding device (300) can receive a signal output from the encoding device in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device can decode the picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in the present disclosure can be obtained from the bitstream by being decoded through the decoding procedure. For example, the entropy decoding unit (310) can decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements required for image restoration and the quantized values ​​of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of a syntax element to be decoded and decoding information of surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, and predicts the occurrence probability of a bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (310) is provided to the prediction unit (330), and residual values ​​on which entropy decoding is performed by the entropy decoding unit (310), i.e., quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310). Meanwhile, the decoding device according to the present disclosure may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder may include the entropy decoding unit (310), and the sample decoder may include at least one of the inverse quantization unit (321), the inverse transformation unit (322), the addition unit (340), the filtering unit (350), the memory (360), and the prediction unit (330).

[0086] The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.

[0087] In the inverse transform unit (322), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).

[0088] The prediction unit can perform a prediction on the current block and generate a predicted block containing prediction samples for the current block. Based on the information regarding the prediction output from the entropy decoding unit (310), the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block, and can determine a specific intra / inter-prediction mode.

[0089] The prediction unit (330) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for screen content coding (SCC), for example. IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block based on a block vector within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in the present disclosure.

[0090] The intra prediction unit can predict the current block by referencing samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include multiple non-directional modes and multiple directional modes. The intra prediction unit can also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0091] The inter prediction unit can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information about the prediction can include information indicating the mode of inter prediction for the current block.

[0092] The addition unit (340) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (predicted block, prediction sample array) output from the prediction unit (330). In cases where there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restoration block.

[0093] The addition unit (340) may be referred to as a restoration unit or restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture.

[0094] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.

[0095] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can apply various filtering methods to the restored picture to generate a modified restored picture, and transmit the modified restored picture to the memory (360), specifically, to the DPB of the memory (360). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0096] The (modified) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit. The memory (360) can store motion information of a block from which motion information is derived (or decoded) within the current picture and / or motion information of blocks within a picture that has already been reconstructed. The stored motion information can be transmitted to the inter prediction unit to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (360) can store reconstructed samples of reconstructed blocks within the current picture and transmit them to the intra prediction unit.

[0097] In this specification, the embodiments described in the filtering unit (260) and the prediction unit (220) of the encoding device (200) can be applied to the filtering unit (350) and the prediction unit (330) of the decoding device (300) in the same or corresponding manner, respectively.

[0098] As described above, prediction is performed to increase compression efficiency when performing video coding. Through this, a predicted block including prediction samples for a current block, which is a coding target block, can be generated. Here, the predicted block includes prediction samples in a spatial domain (or pixel domain). The predicted block is derived identically from an encoding device and a decoding device, and the encoding device can increase video coding efficiency by signaling information (residual information) about the residual between the original block and the predicted block, rather than the original sample value of the original block itself, to a decoding device. The decoding device can derive a residual block including residual samples based on the residual information, and generate a reconstructed block including reconstructed samples by combining the residual block and the predicted block, and can generate a reconstructed picture including the reconstructed blocks.

[0099] The residual information may be generated through a transformation and quantization procedure. For example, the encoding device may derive a residual block between the original block and the predicted block, perform a transformation procedure on residual samples (a residual sample array) included in the residual block to derive transform coefficients, and perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, thereby signaling the related residual information to a decoding device (via a bitstream). Here, the residual information may include information such as value information, position information, a transformation technique, a transformation kernel, and quantization parameters of the quantized transform coefficients. The decoding device may perform an inverse quantization / inverse transformation procedure based on the residual information to derive residual samples (or residual blocks). The decoding device may generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also inversely quantize / inversely transform the quantized transform coefficients to derive a residual block for reference in inter prediction of a subsequent picture, and generate a restored picture based on the residual block.

[0100] In the present disclosure, at least one of quantization / dequantization and / or transformation / inverse transformation may be omitted. If the quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. If the transformation / inverse transformation is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or may still be referred to as a transform coefficient for consistency of expression.

[0101] In addition, in the present disclosure, the quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information may include information about the transform coefficient(s), and the information about the transform coefficient(s) may be signaled via residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficient(s)), and scaled transform coefficients may be derived through inverse transformation (scaling) on ​​the transform coefficients. Residual samples may be derived based on inverse transformation (transformation) on the scaled transform coefficients. This may be similarly applied / expressed in other parts of the present disclosure.

[0102] As described above, the prediction unit of the encoding device / decoding device can perform inter prediction on a block-by-block basis to derive prediction samples. Inter prediction can refer to a prediction derived in a manner dependent on data elements (i.e., sample values, motion information, etc.) of pictures other than the current picture. When inter prediction is applied to the current block, a predicted block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture pointed to by a reference picture index. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample basis based on the correlation of the motion information between the surrounding blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). When inter prediction is applied, the neighboring blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and the reference pictures including the temporal neighboring blocks may be called collocated pictures (colPic).For example, a motion information candidate list may be constructed based on neighboring blocks of the current block, and flag or index information may be signaled to indicate which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the selected neighboring block may be used as a motion vector predictor, and the motion vector difference may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference.

[0103] The above motion information may include L0 motion information and / or L1 motion information depending on the inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. Prediction based on an L0 motion vector may be called an L0 prediction, prediction based on an L1 motion vector may be called an L1 prediction, and prediction based on both the L0 motion vector and the L1 motion vector may be called a bi-prediction (Bi). Here, an L0 motion vector may represent a motion vector associated with a reference picture list L0 (L0), and an L1 motion vector may represent a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include pictures preceding the current picture in output order as reference pictures, and the reference picture list L1 may include pictures succeeding the current picture in output order. The preceding pictures may be called forward (reference) pictures, and the succeeding pictures may be called backward (reference) pictures. The reference picture list L0 may further include pictures succeeding the current picture in output order as reference pictures. In this case, the preceding pictures may be indexed first and the succeeding pictures may be indexed next within the reference picture list L0. The reference picture list L1 may further include pictures preceding the current picture in output order as reference pictures. In this case, the succeeding pictures may be indexed first and the succeeding pictures may be indexed next within the reference picture list 1. Here, the output order may correspond to a POC (picture order count) order.

[0104] A video / image encoding procedure based on inter prediction and a prediction unit within an encoding device can roughly perform the following for inter prediction, for example.

[0105] Figure 4 illustrates an example of an inter prediction procedure.

[0106] Referring to FIG. 4, the inter prediction procedure as described above may include an inter prediction mode / type determination step, a motion vector derivation / refinement step, and an inter prediction performance (prediction sample generation) step. The inter prediction procedure may be performed in an encoding device and a decoding device as described above. In this document, a coding device may include an encoding device and / or a decoding device.

[0107] The coding device determines the inter prediction mode / type (S400).

[0108] An encoding device can determine an inter-prediction mode / type to be applied to the current block among various inter-prediction modes / types disclosed in this document, and can generate prediction-related information. The prediction-related information can include inter-prediction mode information indicating an inter-prediction mode to be applied to the current block and / or inter-prediction type information indicating an inter-prediction type to be applied to the current block. A decoding device can determine an inter-prediction mode / type to be applied to the current block based on the prediction-related information.

[0109] The coding device derives / refines the motion vector of the current block (S410). The coding device can derive / refine the motion vector of the current block based on the determined inter prediction mode / type. Here, motion information of surrounding blocks of the current block can be used to derive / refine the motion vector.

[0110] For example, when skip mode or merge mode is applied to the current block, the coding device may configure a merge candidate list and select one of the merge candidates included in the merge candidate list. Information indicating the selected merge candidate (e.g., merge index) may be included in the prediction-related information.

[0111] As another example, when the (A)MVP mode is applied to the current block, the decoding device may construct a list of (A)MVP candidates, and use the motion vector of a selected MVP (motion vector predictor) candidate among the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. The selection may be indicated based on selection information (an MVP flag or an MVP index). In this case, information about the MVD as well as the selection information may be included in the prediction-related information.

[0112] Meanwhile, as described below, the motion information of the current block can be derived without constructing a candidate list, in which case the motion information of the current block can be derived according to the procedure disclosed in the prediction mode / type described below. In this case, the candidate list construction described above can be omitted.

[0113] The coding device predicts (generates a prediction sample) the current block based on the derived / refined motion vector (S420). The coding device can derive the prediction sample of the current block using samples of the reference block pointed to by the motion vector in the reference picture.

[0114] An encoding procedure based on inter prediction may roughly include, for example:

[0115] Figure 5 shows examples of inter prediction based video / image encoding methods.

[0116] Referring to FIG. 5, S500 may be performed by a prediction unit of an encoding device, S505 may be performed by a residual processing unit of the encoding device, and S510 or S515 may be performed by an entropy encoding unit of the encoding device. Specifically, the prediction-related information may be derived by the prediction unit and encoded by the entropy encoding unit. The residual information may be derived by the residual processing unit and encoded by the entropy encoding unit. The residual information is information about the residual samples. The residual information may include information about quantized transform coefficients for the residual samples. As described above, the residual samples may be derived as transform coefficients through a transform unit of the encoding device, and the transform coefficients may be derived as quantized transform coefficients through a quantization unit. The information about the quantized transform coefficients may be encoded in the entropy encoding unit through a residual coding procedure.

[0117] An encoding device performs inter prediction on a current block (S500). The encoding device can derive the inter prediction mode / type and motion information of the current block, and generate prediction samples of the current block. Here, the inter prediction mode / type determination, motion information derivation, and prediction sample generation procedures may be performed simultaneously, or one procedure may be performed before the other. For example, the inter prediction unit of the encoding device can search for a block similar to the current block within a certain area (search area) of reference pictures through motion estimation, and derive a reference block whose difference from the current block is minimal or below a certain standard. Based on this, a reference picture index indicating the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine a mode to be applied to the current block among various prediction modes. The encoding device can compare RD costs for the various prediction modes and determine an optimal prediction mode for the current block.

[0118] For example, when the skip mode or merge mode is applied to the current block, the encoding device may configure a merge candidate list described below, and derive a reference block among the reference blocks indicated by the merge candidates included in the merge candidate list, the difference between the current block and the current block being at least or below a certain standard. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to a decoding device. Motion information of the current block may be derived using motion information of the selected merge candidate.

[0119] As another example, when the (A)MVP mode is applied to the current block, the encoding device may configure an (A)MVP candidate list described below, and use the motion vector of an mvp candidate selected from among mvp (motion vector predictor) candidates included in the (A)MVP candidate list as the mvp of the current block. In this case, for example, a motion vector indicating a reference block derived by the above-described motion estimation may be used as the motion vector of the current block, and an mvp candidate having a motion vector with the smallest difference from the motion vector of the current block among the mvp candidates may become the selected mvp candidate. A motion vector difference (MVD), which is a difference obtained by subtracting the mvp from the motion vector of the current block, may be derived. In this case, information about the MVD may be signaled to the decoding device. In addition, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and signaled separately to the decoding device.

[0120] The encoding device may perform residual processing based on the predicted samples (S505). The encoding device may derive residual samples based on the predicted samples. The encoding device may derive the residual samples by comparing the original samples of the current block with the predicted samples. Residual information may be generated based on the residual samples. The residual information may include information regarding quantized transform coefficients as described above.

[0121] An encoding device encodes image information including prediction-related information and / or residual information (S510 or S515). The encoding device can output the encoded image information in the form of a bitstream. The prediction-related information may include information related to the prediction procedure, such as prediction mode information (e.g., skip flag, merge flag, or mode index) and information about motion information. The information about the motion information may include candidate selection information (e.g., merge index, mvp flag, or mvp index), which is information for deriving a motion vector. In addition, the information about the motion information may include information about the above-described MVD and / or reference picture index information. In addition, the information about the motion information may include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information about the residual samples. The residual information may include information about quantized transform coefficients for the residual samples.

[0122] The output bitstream can be stored on a (digital) storage medium and transmitted to a decoding device, or can be transmitted to a decoding device via a network.

[0123] Meanwhile, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is to derive the same prediction result as that performed by the decoding device from the encoding device, thereby improving coding efficiency. Accordingly, the encoding device can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in memory and use it as a reference picture for inter prediction. As described above, an in-loop filtering procedure, etc. can be further applied to the reconstructed picture.

[0124] The decoding device can perform operations corresponding to those performed by the encoding device. A video / image decoding procedure based on inter prediction may include, for example, the following.

[0125] Figure 6 shows examples of inter prediction based video / image decoding methods.

[0126] Referring to FIG. 6, S600 may be performed by an entropy decoding unit of a decoding device, S610 may be performed by a prediction unit of the decoding device, S615 may be performed by a residual processing unit of the decoding device, and S620 may be performed by an adder or restoration unit of the decoding device.

[0127] Specifically, the decoding device obtains image / video information from the bitstream (S600). The image / video information may include prediction-related information and / or residual information.

[0128] The decoding device performs inter prediction based on prediction-related information (S610). The decoding device may derive an inter prediction mode / type for the current block based on the prediction-related information, derive / refine motion information of the current block, and generate prediction samples within the current block based on the intra prediction mode / type and / or the motion information. In this case, the decoding device may perform a prediction sample filtering procedure. The prediction sample filtering procedure may be referred to as post-filtering. Some or all of the prediction samples may be filtered by the prediction sample filtering procedure. In some cases, the prediction sample filtering procedure may be omitted.

[0129] The decoding device performs residual processing based on the residual information (S615). The decoding device can derive residual samples for the current block based on the residual information. Specifically, the inverse quantization unit of the residual processing unit performs inverse quantization based on the quantized transform coefficients derived based on the residual information to derive transform coefficients, and the inverse transform unit of the residual processing unit performs inverse transformation on the transform coefficients to derive residual samples for the current block.

[0130] The decoding device generates a reconstructed block / picture (S620). The decoding device can generate reconstructed samples for the current block based on the prediction samples and / or the residual samples, and derive a reconstructed block including the reconstructed samples. A reconstructed picture for the current picture can be generated based on the reconstructed block. As described above, an in-loop filtering procedure, etc., can be further applied to the reconstructed picture.

[0131] The above prediction-related information can be encoded / decoded using the binarization and coding methods described in this document. For example, the prediction-related information can be binarized using fixed-length binarization, truncated Rice binarization, truncated unary binarization, etc. For example, the prediction-related information can be encoded / decoded using entropy coding (e.g., CABAC, CAVLC) coding.

[0132] Meanwhile, for example, according to this document, as an embodiment of inter prediction, GPM (Geometric Partitioning Mode) can be applied. GPM can be viewed as a prediction technique or a type of prediction. When GPM is applied to a current block, the current block can be divided into two partitions, motion information for each of the two partitions can be derived, and inter prediction for each partition can be performed based on the motion information for each of the two partitions, so that a prediction sample of the current block can be derived.

[0133] Figure 7 shows an example of a segmentation shape supported by GPM.

[0134] Referring to FIG. 7, GPM can support 64 segmentation shapes from combinations of 20 angles and 4 distances. For example, the 64 segmentation shapes can include 32 segmentation shapes that are combinations of 8 angles and 4 distances, 24 segmentation shapes that are combinations of 8 angles and 3 distances, and 8 segmentation shapes that are combinations of 4 angles and 2 distances.

[0135] For example, a GPM partition index indicating the partition shape of the current block may be signaled. For example, the partition shape of the current block may be derived from the partition shape indicated by the GPM partition index. The following table may indicate the partition shape indicated by the GPM partition index.

[0136]

[0137] Here, gpm_partition_idx may represent the GPM partition index, angleIdx may represent the angle index, and distanceIdx may represent the distance index. The partition shape of the current block may be derived based on the signaled GPM partition index, and the current block may be divided into partitions based on the derived partition shape.

[0138] Additionally, GPM can be applied to non-square blocks. For example, Fig. 7 (b) can represent the partitioning shape of GPM applied to a non-square block with W / H of 2. Here, W can represent the width of the block, and H can represent the height of the block. For example, if the current block is a non-square block with W / H of 2, the signaled GPM partitioning index can represent the partitioning shape as shown in the following table.

[0139]

[0140] For example, if the current block is a non-square block with W / H of 2, the segmentation shape of the current block can be derived based on Table 2 and the GPM segmentation index instead of Table 1.

[0141] Also, for example, (c) of Fig. 7 can represent the partition shape of GPM applied to a non-square block with W / H of 4. For example, if the current block is a non-square block with W / H of 2, the signaled GPM partition index can represent the partition shape as shown in the following table.

[0142]

[0143] For example, if the current block is a non-square block with W / H of 4, the segmentation shape of the current block can be derived based on Table 3 and the GPM segmentation index.

[0144] For example, a table for a segmentation shape may be derived based on the size of the current block, and the segmentation shape of the current block may be derived based on the derived table and the GPM segmentation index. Alternatively, for example, a table for a segmentation shape may be derived based on the width and / or height of the current block, and the segmentation shape of the current block may be derived based on the derived table and the GPM segmentation index. Alternatively, for example, a table for a segmentation shape may be derived based on the ratio of the width and height of the current block, and the segmentation shape of the current block may be derived based on the derived table and the GPM segmentation index.

[0145] Specifically, for example, when the W / H of the current block is 1, the above Table 1 can be derived as a table for the partition shape of the current block, and the partition shape of the current block can be derived based on the Table 1 and the GPM partition index of the current block. Or, for example, when the W / H of the current block is 2, the above Table 2 can be derived as a table for the partition shape of the current block, and the partition shape of the current block can be derived based on the Table 2 and the GPM partition index of the current block. Or, for example, when the W / H of the current block is 4, the above Table 3 can be derived as a table for the partition shape of the current block, and the partition shape of the current block can be derived based on the Table 3 and the GPM partition index of the current block.

[0146] Meanwhile, for example, there may be a constraint on the size of the block to which the GPM is applied. That is, for example, whether the GPM is applied may be determined based on the size of the current block. For example, the minimum width or minimum height of the block to which the GPM is applied may be 8 or 4. That is, if the width or height of the current block is less than 8 or 4, the GPM may not be applied. In addition, for example, the maximum width or maximum height of the block to which the GPM is applied may be 64 or 128. That is, if the width or height of the current block is greater than 64 or 128, the GPM may not be applied. In addition, for example, the maximum ratio of the width and height of the block to which the GPM is applied may be 1:4 or 4:1. That is, if the maximum ratio of the width and height of the current block is greater than 1:4 or 4:1, the GPM may not be applied. Alternatively, for example, the maximum ratio of the width and height of the block to which the GPM is applied may be 1:8 or 8:1. That is, if the maximum ratio of the width and height of the current block is greater than 1:8 or 8:1, the GPM may not be applied.

[0147] Also, for example, according to this document, GPM with inter and intra prediction can be applied. GPM with inter and intra prediction can be regarded as one of the prediction techniques or prediction types. GPM with inter and intra prediction can be called inter-intra prediction GPM. When GPM with inter and intra prediction is applied to the current block, the current block can be divided into two partitions, and one of the two partitions can apply inter prediction and the other can apply intra prediction.

[0148] For example, in a GPM including inter prediction and intra prediction, a final prediction sample can be generated by assigning weights to inter prediction samples and intra prediction samples for each separated partition. That is, a weight for each of an inter prediction sample, which is a prediction sample of a partition to which inter prediction is applied, and an intra prediction sample, which is a prediction sample of a partition to which intra prediction is applied, can be derived, and a final prediction sample can be generated based on a weighted sum of the inter prediction sample and the intra prediction sample. The inter prediction sample can be derived based on the inter GPM, and the intra prediction sample can be derived based on an IPM (Intra Prediction Mode) candidate list and an index signaled from an encoding device. An inter prediction mode applied to a partition of a current block can be derived by an inter GPM, and an intra prediction mode applied to another partition of the current block can be derived as an IPM candidate indicated by a signaled index among IPM candidates of a configured IPM candidate list. For example, the size of the IPM candidate list can be predefined as 3.

[0149] Figure 8 illustrates an intra prediction mode that can be used as an IPM candidate.

[0150] Referring to Fig. 8, intra prediction modes that can be used as IPM candidates, i.e., available IPM candidates, can be a parallel angle mode (Parallel mode) with respect to the GPM block boundary, a perpendicular angle mode (Perpendicular mode) with respect to the GPM block boundary, and / or a planar mode. Fig. 8 (a) can represent the parallel mode, Fig. 8 (b) can represent the perpendicular mode, and Fig. 8 (c) can represent the planar mode.

[0151] In addition, for example, in the process of constructing an IPM candidate list, the Decoder-side Intra Mode Derivation (DIMD) method and / or the intra prediction mode derived from the surrounding blocks may be derived as an IPM candidate. When the intra prediction mode derived from the DIMD and / or the surrounding blocks is derived as an IPM candidate, the above-described parallel mode may be derived as an IPM candidate first. For example, when the size of the IPM candidate list is 3, the intra prediction mode derived from the DIMD and / or the surrounding blocks may be derived as an IPM candidate after the parallel mode is derived as an IPM candidate, and therefore, up to two IPM candidates derived from the DIMD and / or the surrounding blocks may be derived as IPM candidates if there is no identical IPM candidate in the IPM candidate list. In the case of deriving the surrounding intra prediction mode (i.e., the intra prediction mode derived from the surrounding blocks), the positions of the available surrounding blocks are up to 5, but this may be limited by the GPM block boundary angle as shown in the table below.

[0152]

[0153] Here, Angle of GPM may mean a GPM block boundary angle indicated by the index, 1st partition may mean an available peripheral block location of the first partition, and 2nd partition may mean an available peripheral block location of the second partition. For example, if the current block is divided by the GPM block boundary angle of index 0, the available peripheral blocks of the first partition may include an upper peripheral block, and the available peripheral blocks of the second partition may include a left peripheral block and an upper peripheral block. Or, for example, if the current block is divided by the GPM block boundary angle of index 5, the available peripheral blocks of the first partition may include a left peripheral block and an upper peripheral block, and the available peripheral blocks of the second partition may include a left peripheral block.

[0154] For example, in a GPM including inter prediction and intra prediction, motion information of a partition to which inter prediction is applied can be derived based on a regular GPM MV candidate list. For example, the regular GPM MV candidate list can be a merge candidate list derived based on neighboring blocks of a partition to which inter prediction is applied. For example, a merge candidate list can be derived based on neighboring blocks of a partition to which inter prediction is applied, and motion information of a partition to which inter prediction is applied can be derived based on a merge candidate indicated by a merge index for the partition to which the inter prediction is applied among the merge candidates in the merge candidate list.

[0155] Additionally, GPM including inter-prediction and intra-prediction can be combined with GPM-MMVD (GPM with merge with motion vector difference). That is, GPM-MMVD can be applied to partitions where inter-prediction is applied in GPM including inter-prediction and intra-prediction.

[0156] For example, MMVD information for a partition to which the inter prediction is applied can be signaled, and an MMVD of a partition to which the inter prediction is applied can be derived based on the MMVD information, and motion information of a partition to which the inter prediction is applied can be derived based on a merge candidate of the partition to which the inter prediction is applied and the MMVD. The MMVD information can include an MMVD distance index and an MMVD direction index. The MMVD distance index can indicate a distance of the MMVD, and the MMVD direction index can indicate a direction of the MMVD.

[0157] Additionally, for example, a regression-based GPM can be applied to a GPM including inter-prediction and intra-prediction. For example, a pair list for partitions of a current block to which a GPM including inter-prediction and intra-prediction is applied can be constructed, and a pair index indicating a selected pair candidate can be signaled. The pair candidate can include an intra-prediction mode candidate for a partition to which intra-prediction of the current block is applied, and an MV candidate for a partition to which inter-prediction of the current block is applied.

[0158] For example, a flag indicating whether a regression-based GPM including inter prediction and intra prediction is applied may be signaled. The flag may be signaled at the CU level. Here, the regression-based GPM including inter prediction and intra prediction may be referred to as a regression-based inter-intra prediction GPM, and the flag may be referred to as a regression-based inter-intra prediction GPM flag. In addition, for example, a flag indicating whether the regression-based GPM including inter prediction and intra prediction is available may be signaled in a higher level syntax (e.g., SPS, PPS). When the value of the available flag is 0, signaling of the flag indicating whether the regression-based GPM including inter prediction and intra prediction is applied may be omitted. The available flag may be referred to as a regression-based inter-intra prediction GPM available flag.

[0159] When the regression-based GPM with inter and intra prediction is applied to the current block, a pair list for partitions of the current block can be constructed. For example, the intra prediction mode candidate can be one of the first six intra prediction modes of the MPM list, and the MV candidates can be one of the regular GPM MV candidates. The pair including the intra prediction mode candidate and the MV candidate can be derived based on two integer blending matrices derived as a regression model for the template of the current block. Here, the integer blending matrices can be as follows.

[0160]

[0161] Here, parameters a, b, and c can be derived as parameters that minimize the mean square error (MSE) for the template of the current block.

[0162] A pair index pointing to one of the pair candidates in the above pair list can be signaled, and a pair candidate for the current block can be derived based on the pair index. Based on the MV candidate of the derived pair candidate, motion information of the partition to which inter prediction of the current block is applied can be derived.

[0163] In addition, to further improve coding performance, TIMD can be used as an IPM candidate for a partition to which intra prediction is applied in a GPM including inter prediction and intra prediction. For example, the IPM candidate list for the partition to which intra prediction is applied can be constructed by first deriving the parallel mode as an IPM candidate, and then deriving the TIMD, DIMD, and the surrounding blocks as IPM candidates in that order. That is, for example, when the IPM candidate list is constructed, the IPM candidates can be derived in the following order: parallel mode, TIMD, DIMD, and the intra prediction mode derived from the surrounding blocks.

[0164] Alternatively, for example, in a GPM including inter prediction and intra prediction, the IPM candidates of a partition to which intra prediction is applied may include at least one of a vertical intra prediction mode, a horizontal intra prediction mode, a planar mode, a DC mode, or a directional planar mode.

[0165] Meanwhile, signaling for GPM, including inter prediction and intra prediction, can be performed as follows.

[0166]

[0167] Referring to Table 5, the GPM segmentation index, inter-intra prediction GPM flag, and / or inter-intra prediction GPM mode flag for the current block may be signaled.

[0168] For example, gpm_partition_idx may represent the GPM partition index. For example, the GPM partition index may represent the partitioning shape of the current block to which GPM is applied. For example, based on gpm_partition_idx, the block boundary angle and the distance between the center point of the current block and the block boundary may be derived.

[0169] gpm_partition_interintra_flag can indicate whether the inter-intra prediction GPM is applied. For example, if the value of gpm_partition_interintra_flag is 1, gpm_partition_interintra_flag can indicate that the inter-intra prediction GPM is applied to the current block, and if the value of gpm_partition_interintra_flag is 0, gpm_partition_interintra_flag can indicate that the inter-intra prediction GPM is not applied to the current block. Here, gpm_partition_interintra_flag can be indicated as the inter-intra prediction GPM flag.

[0170] Additionally, for example, if gpm_partition_interintra_flag indicates that inter-intra prediction GPM is applied to the current block, gpm_prediction_interintra_mode_flag may be signaled. gpm_prediction_interintra_mode_flag may indicate a partition to which inter prediction is applied and a partition to which intra prediction is applied. Here, a partition to which inter prediction is applied may be indicated as an inter partition, and a partition to which intra prediction is applied may be indicated as an inter partition. Additionally, gpm_prediction_interintra_mode_flag may be indicated as an inter-intra prediction GPM mode flag.

[0171] For example, if the value of gpm_prediction_interintra_mode_flag is 1, gpm_prediction_interintra_mode_flag may indicate that inter prediction is applied to the first partition of the current block and intra prediction is applied to the second partition, and if the value of gpm_prediction_interintra_mode_flag is 0, gpm_prediction_interintra_mode_flag may indicate that intra prediction is applied to the first partition of the current block and inter prediction is applied to the second partition. That is, when the value of gpm_prediction_interintra_mode_flag is 1, gpm_prediction_interintra_mode_flag can indicate that the first partition of the current block is an inter-partition and the second partition of the current block is an intra-partition, and when the value of gpm_prediction_interintra_mode_flag is 0, gpm_prediction_interintra_mode_flag can indicate that the first partition of the current block is an intra-partition and the second partition of the current block is an inter-partition. Alternatively, for example, if the value of gpm_prediction_interintra_mode_flag is 0, gpm_prediction_interintra_mode_flag may indicate that inter prediction is applied to the first partition of the current block and intra prediction is applied to the second partition, and if the value of gpm_prediction_interintra_mode_flag is 1, gpm_prediction_interintra_mode_flag may indicate that intra prediction is applied to the first partition of the current block and inter prediction is applied to the second partition.That is, when the value of gpm_prediction_interintra_mode_flag is 0, gpm_prediction_interintra_mode_flag may indicate that the first partition of the current block is an inter-partition and the second partition of the current block is an intra-partition, and when the value of gpm_prediction_interintra_mode_flag is 1, gpm_prediction_interintra_mode_flag may indicate that the first partition of the current block is an intra-partition and the second partition of the current block is an inter-partition. Here, gpm_prediction_interintra_mode_flag may be indicated as an inter-intra prediction GPM mode flag.

[0172] Additionally, for partitions of the current block, an index indicating a motion information candidate for inter prediction and / or an index indicating an IPM candidate may be signaled. For example, if the first partition of the current block is an inter partition and the second partition is an intra partition, an index gpm_inter_idx0 indicating a motion information candidate of the first partition may be signaled, and an index gpm_intra_idx1 indicating an IPM candidate of the second partition may be signaled. Additionally, for example, if the first partition of the current block is an intra partition and the second partition is an inter partition, an index gpm_intra_idx0 indicating an IPM candidate of the first partition may be signaled, and an index gpm_inter_idx1 indicating a motion information candidate of the second partition may be signaled.

[0173] Alternatively, for example, signaling for GPM with inter and intra prediction can be performed as follows.

[0174]

[0175] Referring to Table 6, the GPM split index and inter-intra prediction GPM indicator for the current block can be signaled.

[0176] For example, gpm_partition_interintra_idc may indicate whether a GPM in which the first partition of the current block is an inter-partition and the second partition of the current block is an intra-partition, a GPM in which the first partition of the current block is an intra-partition and the second partition of the current block is an inter-partition, or a GPM with inter and intra prediction is not applied to the current block. The gpm_partition_interintra_idc may be binarized based on truncated rice (or truncated unary). Here, gpm_partition_interintra_idc may indicate the inter-intra prediction GPM indicator.

[0177] For example, the inter-intra prediction GPM mode and binarization indicated by gpm_partition_interintra_idc can be as follows.

[0178]

[0179] Referring to Table 7, when the value of gpm_partition_interintra_idc is 0, gpm_partition_interintra_idc may indicate that the inter-intra prediction GPM is not applied, when the value of gpm_partition_interintra_idc is 1, it may indicate that the first partition is an inter partition and the second partition of the current block is an intra partition and the GPM is applied, and when the value of gpm_partition_interintra_idc is 2, it may indicate that the first partition is an intra partition and the second partition of the current block is an inter partition and the GPM is applied.

[0180] Additionally, for partitions of the current block, an index indicating a motion information candidate for inter prediction and / or an index indicating an IPM candidate may be signaled. For example, if the first partition of the current block is an inter partition and the second partition is an intra partition, an index gpm_inter_idx0 indicating a motion information candidate of the first partition may be signaled, and an index gpm_intra_idx1 indicating an IPM candidate of the second partition may be signaled. Additionally, for example, if the first partition of the current block is an intra partition and the second partition is an inter partition, an index gpm_intra_idx0 indicating an IPM candidate of the first partition may be signaled, and an index gpm_inter_idx1 indicating a motion information candidate of the second partition may be signaled.

[0181] In addition, for example, the partition to which inter prediction is applied and the partition to which intra prediction is applied can be determined based on the partition type of the current block, i.e., the partitioning shape. That is, for example, the partition to which inter prediction is applied and the partition to which intra prediction is applied can be determined based on the partitioning shape of the current block. For example, if one of the partitions of the current block is divided into a partition that is not adjacent to a neighboring sample, the partition that is not adjacent to the neighboring sample can be determined as the partition to which inter prediction is applied, and the other partition can be determined as the partition to which intra prediction is applied. In this case, the syntax indicating the partition to which inter prediction is applied and the partition to which intra prediction is applied (e.g., the gpm_prediction_interintra_mode_flag described above) may not be signaled. The neighboring samples may include an upper neighboring sample and a left neighboring sample. Partitions that are not adjacent to surrounding samples may have less dependency on surrounding blocks and thus are more likely to not be subject to intra prediction. Therefore, determining the partitions to which inter prediction is applied and the partitions to which intra prediction is applied based on the partition shape can improve coding efficiency and reduce the amount of bits for inter-intra prediction GPM.

[0182] Meanwhile, when the (A)MVP mode is applied to the current block as described above, the decoding device can construct an (A)MVP candidate list, derive the motion vector of an MVP candidate selected from among the MVP (motion vector predictor) candidates included in the (A)MVP candidate list as the MVP of the current block, derive the MVD of the current block based on information about the signaled MVD, and derive motion information of the current block based on the MVP and the MVD.

[0183] In addition, according to this document, MVD prediction can be applied. According to MVD prediction, a syntax indicating an MVD prediction index pointing to an MVD and a syntax indicating an MVD magnitude can be parsed, MVD candidates can be derived by combining possible signs and the MVD magnitude derived based on the syntax, MVs combined with the derived MVD candidates can be reordered with a TM (Template matching) cost, and among the reordered candidates, a candidate indicated by the MVD prediction index can be derived as the MVD of the current block.

[0184] Specifically, according to the MVD prediction, the decoding device can derive the MVD as follows.

[0185] For example, 1) the decoding device can parse the magnitude of the MVD components. That is, the MVD magnitude syntax indicating the size of the MVD components can be signaled. Thereafter, 2) the decoding device can parse the context-coded MVD prediction index. The MVD prediction index can point to one of the MVD candidates. 3) The decoding device can generate MVD candidates that are combinations of available magnitudes derived based on available signs and the MVD magnitude syntax, and can configure the MV candidates by adding the MVD candidates to the MV predictor (MVP) of the current block. 4) The decoding device can derive an MVD prediction cost for each of the MV candidates, and can sort the MV candidates based on the MVD prediction costs of the MV candidates. Here, the MVD prediction cost of the MV candidate may be the template matching (TM) cost of the MV candidate. 5) The decoding device may select a true MVD from among the sorted MV candidates based on the signaled MVD prediction index.

[0186] Additionally, a bilinear filter can be used to generate a reference template for deriving, for example, the MVD prediction cost of an MV candidate.

[0187] Figure 9 shows a reference template for deriving the TM cost, which is the MVD prediction cost of an MV candidate.

[0188] As illustrated in Fig. 9, the TM cost of an MV candidate can be derived as the Sum of Absolute Differences (SAD) between a current template including surrounding samples of the current block and a reference template including surrounding reference samples of the reference block pointed to by the MV candidate. For example, the TM cost of the MV candidate can be derived based on the following mathematical equation.

[0189]

[0190] Here, i, j represent the location (i, j) of the sample within the template, and Costdistortion is the cost, Temp ref Temp is the sample value of the reference template of the reference block pointed to by the MV candidate. cur represents a sample value of the current template of the current block. The differences between corresponding samples between the reference template and the current template can be accumulated, and the accumulation of the differences can be used as a cost function for sorting MV candidates of the current block.

[0191] Additionally, for example, MVD prediction can be applied not only to AMVP mode but also to affine AMVP mode, MMVD, and affine MMVD modes. Furthermore, if wraparound motion compensation is available, MV candidates can be clipped to account for the wraparound offset.

[0192] For example, when MVD prediction is applied to the affine AMVP mode or the affine MMVD mode, sub-block-based templates may be used. For example, the template matching cost for each sub-block of the current block may be accumulated to derive the final cost for the MV candidate of the current block.

[0193] Figure 10 shows a reference template for deriving the TM cost, which is the MVD prediction cost of an MV candidate in the affine AMVP mode or the affine MMVD mode.

[0194] As illustrated in FIG. 10, in the affine AMVP mode or the affine MMVD mode, the TM cost of each sub-block of the current block can be derived, and the TM cost of each sub-block can be accumulated to derive the final cost for the MV candidate of the current block.

[0195] Additionally, for example, when coding the MVD size syntax for the size of an MVD, the first six valid suffix bins of the MVD size syntax may be context coded. The valid suffix bins may include a sign bin.

[0196] Also, for example, the number of valid suffix bins of the MVD size syntax to be context-coded can be derived based on the block size. For example, a block with a width and height greater than N can have up to 6 valid suffix bins including a sine bin context-coded, and a block with a width or height less than or equal to N can have up to 2 valid suffix bins coded. For example, N can be 4. Or, for example, a block with a width and height greater than N can have up to 6 valid suffix bins including a sine bin context-coded, and a block with a width or height less than or equal to N can have up to 4 valid suffix bins coded. For example, N can be 4.

[0197] Additionally, for example, the number of sign and size combinations of the above MMVD can be 16.

[0198] Figure 11 shows MMVD candidates that are combinations of the available signs and sizes.

[0199] As illustrated in Fig. 11, there can be 16 MMVD candidates. In addition, for example, when the MVD prediction is applied to 16 MMVDs, TM costs for the MMVD candidates can be derived, and the MMVD candidates can be sorted in the order of the derived TM costs so that only the eight candidates in the earlier order can be derived as MMVD candidates for the current block. The TM cost for the MMVD candidate can be derived as the SAD (Sum of Absolute Differences, SAD) between the reference template of the reference block pointed to by the MV candidate derived based on the MMVD candidate and the MVP of the current block and the template of the current block.

[0200] In addition, this paper proposes an affine motion model that efficiently derives motion vectors for sub-blocks or sample points of a current block and improves the accuracy of inter prediction despite deformations such as rotation, zoom-in, or zoom-out of an image. In other words, an affine motion model that derives motion vectors for sub-blocks or sample points of a current block can be proposed. Prediction using the above affine motion model can be called affine inter prediction or affine motion prediction.

[0201] The encoding device / decoding device can predict the distortion form of the image based on the motion vectors at the control points (CPs) of the current block through the affine inter prediction, thereby improving the compression performance of the image by increasing the accuracy of the prediction. In addition, since the motion vector for at least one control point of the current block can be derived using the motion vector of the surrounding blocks of the current block, the data burden for the additional information can be reduced, and the inter prediction efficiency can be significantly improved.

[0202] For example, inter prediction using the above-described affine motion model, i.e., affine motion prediction, may have an affine merge mode (AF_MERGE) and an affine inter mode (AF_INTER). Here, the affine inter mode may also be expressed as an affine MVP mode (affine motion vector prediction mode, AF_MVP).

[0203] The above affine merge mode is similar to the existing merge mode in that it does not transmit the MVD for the motion vector of the control points. That is, the affine merge mode can represent an encoding / decoding method that performs prediction by deriving CPMV for each of two or three control points from the surrounding blocks of the current block without coding the MVD (motion vector difference), similar to the existing skip / merge mode. Here, if the top-left sample position in the current block is (0,0), the sample positions (0,0), (w, 0), and (0, h) can be determined as the control points. Hereinafter, the control point at the (0,0) sample position can be represented as CP0, the control point at the (w, 0) sample position can be represented as CP1, and the control point at the (0, h) sample position can be represented as CP2.

[0204] For example, when the AF_MRG mode is applied to the current block, MVs (i.e., CPMV0, CPMV1 or CPMV0, CPMV1, CPMV2) for CP0 and CP1 (or CP0, CP1, and CP2) can be derived from the surrounding blocks of the current block to which the affine mode is applied. That is, CPMV0 and CPMV1 (or CPMV0, CPMV1, and CPMV2) of the surrounding blocks to which the affine mode is applied can be derived as merge candidates, and the merge candidates can be derived as CPMV0 and CPMV1 (or CPMV0, CPMV1, and CPMV2) for the current block.

[0205] Here, when the affine merge mode is applied to the current block, the encoding device / decoding device can construct an affine merge candidate list based on the surrounding blocks of the current block.

[0206] Additionally, each affine merge candidate can mean a combination of CPMVs of CP0 and CP1 in a four-parameter affine motion model, and can mean a combination of CPMVs of CP0, CP1, and CP2 in a six-parameter affine motion model.

[0207] The above affine inter mode may represent inter prediction that derives an MVP (motion vector predictor) for the motion vectors of the control points, derives the motion vectors of the control points based on the received MVD (motion vector difference) and the MVP, derives an affine MVF of the current block based on the motion vectors of the control points, and performs prediction based on the affine MVF. Here, the motion vector of the control point may be expressed as CPMV (Control Point Motion Vector), the MVP of the control point may be expressed as CPMVP (Control Point Motion Vector Predictor), and the MVD of the control point may be expressed as CPMVD (Control Point Motion Vector Difference). Specifically, for example, the encoding device can derive a control point point motion vector predictor (CPMVP) and a control point point motion vector (CPMV) for each of CP0 and CP1 (or CP0, CP1, and CP2), and transmit or store information about the CPMVP and / or a CPMVD which is a difference between the CPMVP and the CPMV.

[0208] Here, when the above affine inter mode is applied to the current block, the encoding device / decoding device can construct an affine MVP candidate list based on the surrounding blocks of the current block, and the affine MVP candidate can be referred to as a CPMVP pair candidate, and the affine MVP candidate list can also be referred to as a CPMVP candidate list.

[0209] Additionally, each affine MVP candidate can mean a combination of CPMVPs of CP0 and CP1 in a four-parameter affine motion model, and can mean a combination of CPMVPs of CP0, CP1, and CP2 in a six-parameter affine motion model.

[0210] Fig. 12 illustrates an exemplary flowchart of an affine motion prediction method according to one embodiment of the present document.

[0211] Referring to Fig. 12, the affine motion prediction method can be broadly expressed as follows. When the affine motion prediction method starts, CPMVs can first be acquired (S1200). Here, the CPMVs can include CPMV0 and CPMV1 when using a 4-parameter affine model, and can include CPMV0, CPMV1, and CPMV2 when using a 6-parameter affine model.

[0212] Afterwards, affine motion compensation can be performed based on CPMVs (S1210), and affine motion prediction can be terminated.

[0213] Additionally, there may be two affine prediction modes to determine the CPMVs. Here, the two affine prediction modes may include affine inter mode and affine merge mode. The affine inter mode can clearly determine CPMVs by signaling motion vector difference (MVD) information for CPMVs. On the other hand, the affine merge mode can derive CPMVs without signaling MVD information.

[0214] In other words, the affine merge mode can derive the CPMV of the current block using the CPMV of the surrounding blocks coded in the affine mode, and when the motion vector is determined in units of sub-blocks, the affine merge mode can also be referred to as a sub-block merge mode.

[0215] In affine merge mode, the encoding device can signal an index of a neighboring block coded in affine mode for deriving a CPMV of a current block to a decoding device, and can also signal a difference value between the CPMV of the neighboring block and the CPMV of the current block. Here, the affine merge mode can construct an affine merge candidate list based on the neighboring blocks, and the index of the neighboring block can indicate a neighboring block to be referenced for deriving the CPMV of the current block among the affine merge candidate list. The affine merge candidate list may also be referred to as a subblock merge candidate list.

[0216] The affine inter mode may also be referred to as the affine MVP mode. In the affine MVP mode, the CPMV of the current block can be derived based on the CPMVP (Control Point Motion Vector Predictor) and the CPMVD (Control Point Motion Vector Difference). In other words, the encoding device can determine the CPMVP for the CPMV of the current block, derive the CPMVD, which is the difference between the CPMV and the CPMVP of the current block, and signal information about the CPMVP and information about the CPMVD to the decoding device. Here, the affine MVP mode can construct an affine MVP candidate list based on neighboring blocks, and the information about the CPMVP can indicate neighboring blocks to be referenced to derive the CPMVP for the CPMV of the current block among the affine MVP candidate list. The affine MVP candidate list may also be referred to as a control point motion vector predictor candidate list.

[0217] For example, if the affine merge mode is applied to the current block, the current block can be coded as described below.

[0218] The encoding device / decoding device can construct an affine merge candidate list including affine merge candidates for the current block, and can derive CPMVs (Control Point Motion Vectors) for CPs (Control Points) of the current block based on one of the affine merge candidates in the affine merge candidate list. The encoding device / decoding device can derive prediction samples for the current block based on the CPMVs, and can generate a reconstructed picture for the current block based on the derived prediction samples.

[0219] Specifically, the above affine merge candidate list can be composed as follows.

[0220] Figure 13 shows an example of constructing an affine merge candidate list of the current block.

[0221] Referring to FIG. 13, the encoding device may add a subblock-based temporal merging candidate to an affine merge candidate list (S1300). Specifically, the encoding device / decoding device may derive the candidate based on collocated sub-blocks of a collocated block in a reference picture. For example, the sub-block-based temporal merge candidate may include sub-block unit motion information derived based on motion information of the collocated sub-blocks. The sub-block-based temporal merge candidate may be referred to as a SbTMVP (subblock-based temporal motion vector prediction candidate). In addition, a reference picture including the collocated block may be referred to as a collocated picture (colPic). Meanwhile, a specific method for deriving the sub-block-based temporal merge candidate will be described later.

[0222] Thereafter, the encoding device / decoding device can add the inherited affine candidate to the affine merge candidate list (S1310).

[0223] Specifically, the encoding device / decoding device can derive an inherited affine candidate based on the surrounding blocks of the current block. Here, the surrounding blocks can include a lower left corner surrounding block A0, a left surrounding block A1, an upper left corner surrounding block B0, an upper right corner surrounding block B1, and an upper left corner surrounding block B2 of the current block.

[0224] Fig. 14 exemplarily illustrates the surrounding blocks of the current block for deriving the inherited affine candidate. Referring to Fig. 14, the surrounding blocks of the current block may include a lower left corner surrounding block A0 of the current block, a left surrounding block A1 of the current block, an upper surrounding block B0 of the current block, an upper right corner surrounding block B1 of the current block, and an upper left corner surrounding block B2 of the current block.

[0225] For example, if the size of the current block is WxH and the x component of the top-left sample position of the current block is 0 and the y component is 0, the left peripheral block may be a block including a sample with coordinates (-1, H-1), the upper peripheral block may be a block including a sample with coordinates (W-1, -1), the upper-right corner peripheral block may be a block including a sample with coordinates (W, -1), the lower-left corner peripheral block may be a block including a sample with coordinates (-1, H), and the upper-left corner peripheral block may be a block including a sample with coordinates (-1, -1).

[0226] The above-described inherited affine candidates can be derived based on valid surrounding reconstructed blocks coded in affine mode. For example, the encoding device / decoding device can sequentially check surrounding blocks A0, A1, B0, B1, and B2, and if the surrounding blocks are coded in affine mode (i.e., if the surrounding blocks are validly reconstructed using an affine motion model), two CPMVs or three CPMVs for the current block can be derived based on the affine motion model of the surrounding blocks, and the CPMVs can be derived as inherited affine candidates for the current block. For example, up to five inherited affine candidates can be added to the affine merge candidate list. That is, up to five inherited affine candidates can be derived based on the surrounding blocks.

[0227] Thereafter, the encoding device / decoding device can add the constructed affine candidate to the affine merge candidate list (S1320).

[0228] For example, when the number of affine candidates in the affine merge candidate list is less than 5, the constructed affine candidate may be added to the affine merge candidate list. The constructed affine candidate may represent an affine candidate generated by combining surrounding motion information (i.e., motion vectors and reference picture indices of surrounding blocks) for each of the CPs of the current block. The motion information for each CP may be derived based on spatial surrounding blocks or temporal surrounding blocks for the CP. The motion information for each of the CPs may be represented as a candidate motion vector for the CP.

[0229] Figure 15 illustrates blocks surrounding the current block for deriving the constructed affine candidate.

[0230] Referring to FIG. 15, the surrounding blocks may include spatial surrounding blocks and temporal surrounding blocks. The spatial surrounding blocks may include surrounding block A0, surrounding block A1, surrounding block A2, surrounding block B0, surrounding block B1, surrounding block B2, and surrounding block B3. The surrounding block T illustrated in FIG. 15 may represent the temporal surrounding block.

[0231] Here, the peripheral block B2 may represent a peripheral block located at the upper left of the upper left sample position of the current block, the peripheral block B3 may represent a peripheral block located at the upper left of the upper left sample position of the current block, and the peripheral block A2 may represent a peripheral block located at the left end of the upper left sample position of the current block. In addition, the peripheral block B1 may represent a peripheral block located at the upper right of the upper right sample position of the current block, and the peripheral block B0 may represent a peripheral block located at the upper right of the upper right of the upper right sample position of the current block. In addition, the peripheral block A1 may represent a peripheral block located at the left end of the lower left sample position of the current block, and the peripheral block A0 may represent a peripheral block located at the lower left end of the lower left sample position of the current block.

[0232] In addition, referring to FIG. 15, the CPs of the current block may include CP0, CP1, CP2, and / or CP3. The CP0 may indicate a top-left position of the current block, the CP1 may indicate a top-right position of the current block, the CP2 may indicate a bottom-left position of the current block, and the CP3 may indicate a bottom-right position of the current block. For example, when the size of the current block is WxH and the x-component of the top-left sample position of the current block is 0 and the y-component is 0, the CP0 may indicate a position of coordinates (0, 0), the CP1 may indicate a position of coordinates (W, 0), the CP2 may indicate a position of coordinates (0, H), and the CP3 may indicate a position of coordinates (W, H).

[0233] Candidate motion vectors for each of the above-described CPs can be derived as follows.

[0234] For example, the encoding device / decoding device can check whether the neighboring blocks in the first group are available in a first order, and can derive the motion vector of the available neighboring block confirmed for the first time in the checking process as a candidate motion vector for CP1. That is, the candidate motion vector for CP1 may be the motion vector of the available neighboring block confirmed for the first time by checking the neighboring blocks in the first group in a first order. The availability may indicate that the motion vector of the neighboring block exists. That is, the available neighboring block may be a block coded with inter prediction (i.e., a block to which inter prediction is applied). Here, for example, the first group may include the neighboring block B2, the neighboring block B3, and the neighboring block A2. The first order may be an order from the neighboring block B2 to the neighboring block B3 to the neighboring block A2 in the first group. For example, if the surrounding block B2 is available, the motion vector of the surrounding block B2 can be derived as a candidate motion vector for CP1; if the surrounding block B2 is not available and the surrounding block B3 is available, the motion vector of the surrounding block B3 can be derived as a candidate motion vector for CP1; and if the surrounding block B2 and the surrounding block B3 are not available and the surrounding block A2 is available, the motion vector of the surrounding block A2 can be derived as a candidate motion vector for CP1.

[0235] In addition, for example, the encoding device / decoding device can check whether the neighboring blocks in the second group are available in a second order, and can derive the motion vector of the available neighboring block confirmed for the first time in the checking process as a candidate motion vector for CP2. That is, the candidate motion vector for CP2 may be the motion vector of the available neighboring block confirmed for the first time by checking the neighboring blocks in the second group in a second order. The availability may indicate that the motion vector of the neighboring block exists. That is, the available neighboring block may be a block coded with inter prediction (i.e., a block to which inter prediction is applied). Here, the second group may include the neighboring block B1 and the neighboring block B0. The second order may be an order from the neighboring block B1 to the neighboring block B0 in the second group. For example, if the surrounding block B1 is available, the motion vector of the surrounding block B1 can be derived as a candidate motion vector for CP2, and if the surrounding block B1 is not available and the surrounding block B0 is available, the motion vector of the surrounding block B0 can be derived as a candidate motion vector for CP2.

[0236] In addition, for example, the encoding device / decoding device can check whether the neighboring blocks in the third group are available in a third order, and can derive the motion vector of the available neighboring block that is first confirmed in the checking process as a candidate motion vector for CP3. That is, the candidate motion vector for CP3 may be the motion vector of the available neighboring block that is first confirmed by checking the neighboring blocks in the third group in a third order. The availability may indicate that the motion vector of the neighboring block exists. That is, the available neighboring block may be a block coded with inter prediction (i.e., a block to which inter prediction is applied). Here, the third group may include the neighboring block A1 and the neighboring block A0. The third order may be an order from the neighboring block A1 to the neighboring block A0 in the third group. For example, if the surrounding block A1 is available, the motion vector of the surrounding block A1 can be derived as a candidate motion vector for CP3, and if the surrounding block A1 is not available and the surrounding block A0 is available, the motion vector of the surrounding block A0 can be derived as a candidate motion vector for CP3.

[0237] Additionally, for example, the encoding device / decoding device can check whether the temporal peripheral block (i.e., the peripheral block T) is available, and if the temporal peripheral block (i.e., the peripheral block T) is available, the motion vector of the temporal peripheral block (i.e., the peripheral block T) can be derived as a candidate motion vector for the CP4.

[0238] A combination of a candidate motion vector for the CP0, a candidate motion vector for the CP1, a candidate motion vector for the CP2, and / or a candidate motion vector for the CP3 can be derived as a constructed affine candidate.

[0239] For example, as described above, the 6-affine model requires motion vectors of three CPs. Three CPs among CP0, CP1, CP2, and CP3 for the 6-affine model can be selected. For example, the CPs can be selected as one of {CP0, CP1, CP3}, {CP0, CP1, CP2}, {CP1, CP2, CP3}, and {CP0, CP2, CP3}. For example, the 6-affine model can be configured using CP0, CP1, and CP2. In this case, the CPs can be represented as {CP1, CP1, CP2}.

[0240] Also, for example, the 4-affine model as described above requires motion vectors of two CPs. Two CPs among CP0, CP1, CP2, and CP3 for the 4-affine model can be selected. For example, the CPs can be selected as one of {CP0, CP3}, {CP1, CP2}, {CP0, CP1}, {CP1, CP3}, {CP0, CP2}, {CP2, CP3}. For example, the 4-affine model can be configured using CP0 and CP1. In this case, the CPs can be represented as {CP0, CP1}.

[0241] Constructed affine candidates, which are combinations of candidate motion vectors, can be added to the affine merge candidate list in the following order. That is, after candidate motion vectors for the CPs are derived, constructed affine candidates can be derived in the following order.

[0242] {CP0, CP1, CP2}, {CP0, CP1, CP3}, {CP0, CP2, CP3}, {CP1, CP2, CP3}, {CP0, CP1}, {CP0, CP2}, {CP1, CP2}, {CP0, CP3}, {CP1, CP3}, {CP2, CP3}

[0243] That is, for example, a constructed affine candidate including a candidate motion vector for CP0, a candidate motion vector for CP1, and a candidate motion vector for CP2, a constructed affine candidate including a candidate motion vector for CP0, a candidate motion vector for CP1, and a candidate motion vector for CP3, a constructed affine candidate including a candidate motion vector for CP0, a candidate motion vector for CP2, and a candidate motion vector for CP3, a constructed affine candidate including a candidate motion vector for CP1, a candidate motion vector for CP2, and a candidate motion vector for CP3, a constructed affine candidate including a candidate motion vector for CP0, a candidate motion vector for CP1, a constructed affine candidate including a candidate motion vector for CP0, a candidate motion vector for CP2, a constructed affine candidate including a candidate motion vector for CP1, a candidate motion vector for CP2, a constructed affine candidate including a candidate motion vector for CP0, a candidate motion vector for CP3, A constructed affine candidate including a candidate motion vector for CP1, a constructed affine candidate including a candidate motion vector for CP3, a constructed affine candidate including a candidate motion vector for CP2, and a constructed affine candidate including a candidate motion vector for CP3 may be added to the affine merge candidate list in this order.

[0244] Afterwards, the encoding device / decoding device can add 0 motion vectors as affine candidates to the affine merge candidate list (S1330).

[0245] For example, if the number of affine candidates in the above affine merge candidate list is less than 5, affine candidates including 0 motion vectors may be added to the above affine merge candidate list until the above affine merge candidate list is configured with the maximum number of affine candidates. The maximum number of affine candidates may be 5. In addition, the above 0 motion vector may indicate a motion vector with a vector value of 0.

[0246] Furthermore, this paper proposes a method for deriving regression-based affine candidates, either as affine merge candidates or affine MVP candidates. For example, regression-based affine candidate derivation can be viewed as a technique or type of derivation for affine candidates in inter prediction.

[0247] For example, when a regression-based affine candidate derivation method is applied, the motion vectors and center positions of surrounding sub-blocks can be input into a linear regression process to derive a set of linear model parameters, and the predicted CPMVs (control point motion vectors) can be derived as regression-based affine candidates.

[0248] Figure 16 shows the surrounding sub-blocks that are input into the linear regression process to derive a set of linear model parameters.

[0249] The 4x4 sized surrounding sub-blocks of the current block illustrated in Fig. 16 can be input into a linear regression process to derive a linear model parameter set. For example, the motion vectors and center positions of the 4x4 sized surrounding sub-blocks can be input into the linear regression process. The linear regression process can be expressed as a regression model. In addition, the linear model parameter set can be parameters of an affine motion model.

[0250] For example, the linear model parameter set for a six parameter affine motion model might be:

[0251]

[0252] Referring to Equation 3, the linear model parameter set is linear model parameters a xx , a xy , a yx , a yy , b x , b y may include. The linear model parameter set may be represented as an affine model parameter set, and the linear model parameters may be represented as affine model parameters. The regression model may be a linear regression model of the mean square error (MSE) method. That is, for example, the affine model parameters may be calculated by solving a regression model in which motion vectors and center positions of surrounding sub-blocks are input from the perspective of the mean square error (MSE).

[0253] Referring to Fig. 16, W may represent the width of the current block, and H may represent the height of the current block. The input to the regression model may include the center positions (x, y) and motion vectors (mvx and mvy) of available surrounding 4Х4 sub-blocks as described above. In addition, for example, as illustrated in Fig. 16, the sub-blocks used as input may include lower left surrounding sub-blocks, left surrounding sub-blocks, upper left surrounding sub-blocks, upper surrounding sub-blocks, and / or upper right surrounding sub-blocks. In addition, as illustrated in FIG. 16, the reference area including the upper right peripheral sub-blocks may be an area whose width is half (W / 2) of the width of the current block, the reference area including the upper left peripheral sub-blocks may be an area whose width is half (W / 2) of the width of the current block, and the reference area including the lower left peripheral sub-blocks may be an area whose height is half (H / 2) of the height of the current block.

[0254] In addition, for example, a method for deriving surrounding sub-blocks used to derive regression-based affine candidates based on the shape of the current block may be proposed. For example, if the current block is a non-square block, a regression-based affine candidate may be derived based only on surrounding sub-blocks in some areas among the lower left surrounding sub-blocks, left surrounding sub-blocks, upper left surrounding sub-blocks, upper surrounding sub-blocks, and upper right surrounding sub-blocks of the current block.

[0255] For example, if the current block is a non-square block whose width is greater than its height, the motion vectors and center positions of the available sub-blocks among the upper-left surrounding sub-blocks, upper-right surrounding sub-blocks, and upper-right surrounding sub-blocks of the current block can be input into a linear regression process to derive predicted CPMVs as regression-based affine candidates. That is, for example, if the current block is a non-square block whose width is greater than its height, the motion vectors and center positions of the available sub-blocks among the upper-left surrounding sub-blocks, upper-right surrounding sub-blocks, and upper-right surrounding sub-blocks of the current block can be input into a regression model to be solved in terms of mean square error (MSE), so that the affine model parameters can be calculated, and the CPMVs derived based on the affine model parameters can be derived as regression-based affine candidates.

[0256] Alternatively, for example, if the current block is a non-square block whose width is greater than twice its height, the motion vectors and center positions of the available sub-blocks among the upper-left surrounding sub-blocks, the upper-right surrounding sub-blocks, and the upper-right surrounding sub-blocks of the current block may be input into a linear regression process, and the predicted CPMVs may be derived as regression-based affine candidates. That is, for example, if the current block is a non-square block whose width is greater than twice its height, the motion vectors and center positions of the available sub-blocks among the upper-left surrounding sub-blocks, the upper-right surrounding sub-blocks, and the upper-right surrounding sub-blocks of the current block may be input into a regression model, and the affine model parameters may be calculated by solving the regression model in terms of the mean square error (MSE), and the CPMVs derived based on the affine model parameters may be derived as regression-based affine candidates.

[0257] In addition, for example, when the current block is a non-square block whose height is greater than its width, the motion vectors and center positions of the available sub-blocks among the lower-left surrounding sub-blocks, the left surrounding sub-blocks, and the upper-left surrounding sub-blocks of the current block can be input into a linear regression process, and the predicted CPMVs can be derived as regression-based affine candidates. That is, for example, when the current block is a non-square block whose height is greater than its width, the motion vectors and center positions of the available sub-blocks among the lower-left surrounding sub-blocks, the left surrounding sub-blocks, and the upper-left surrounding sub-blocks of the current block can be input into a regression model, and the affine model parameters can be calculated by solving the regression model in terms of the mean square error (MSE), and the CPMVs derived based on the affine model parameters can be derived as regression-based affine candidates.

[0258] Alternatively, for example, if the current block is a non-square block whose height is greater than twice its width, the motion vectors and center positions of the available sub-blocks among the lower-left surrounding sub-blocks, the left surrounding sub-blocks, and the upper-left surrounding sub-blocks of the current block may be input into a linear regression process, and the predicted CPMVs may be derived as regression-based affine candidates. That is, for example, if the current block is a non-square block whose height is greater than twice its width, the motion vectors and center positions of the available sub-blocks among the lower-left surrounding sub-blocks, the left surrounding sub-blocks, and the upper-left surrounding sub-blocks of the current block may be input into a regression model, and the affine model parameters may be calculated by solving the regression model in terms of the mean square error (MSE), and the CPMVs derived based on the affine model parameters may be derived as regression-based affine candidates.

[0259] Additionally, for example, the predicted CPMVs (control point motion vectors) can be derived as regression-based affine candidates by inputting the subblock motion field of a previously coded affine block (e.g., a previously coded affine CU) and / or the motion vectors of surrounding subblocks adjacent to the current block into a linear regression process. The derived regression-based affine candidates can be added to an affine merge candidate list. The previously coded affine blocks can be derived by scanning a location that is not adjacent to the current block and an affine HMVP (History-based Motion Vector Prediction) table.

[0260] In addition, for example, up to two regression-based affine candidates can be derived. For example, one regression-based affine candidate can be derived based on motion information of neighboring sub-blocks adjacent to the current block, and one regression-based affine candidate can be derived based on motion information of blocks other than the neighboring neighboring sub-blocks. In addition, when ARMC-TM (Adaptive reordering of merge candidates with template matching) is applied, the linear regression affine candidates can be included in one sub-group and the ARMC-TM can be applied. In addition, for example, the number of affine candidates to which ARMC-TM is applied can be 30, and a reordered candidate list including 15 candidates can be derived by reordering them based on the TM cost.

[0261] FIG. 17 schematically illustrates a video / image encoding method according to an embodiment(s) of the present disclosure. The method disclosed in FIG. 17 may be performed by the encoding device disclosed in FIG. 2. Specifically, for example, S1700 to S1730 of FIG. 17 may be performed by the prediction unit (220) of the encoding device (200), and S1740 of FIG. 17 may be performed by the entropy encoding unit (240) of the encoding device (200). The method disclosed in FIG. 17 may include the embodiments described above in the present disclosure.

[0262] Referring to FIG. 17, the encoding device derives a regression-based affine candidate of the current block based on surrounding sub-blocks of the current block (S1700).

[0263] The encoding device can derive the prediction mode applied to the current block from among various prediction modes / types as the affine merge mode or the affine MVP mode. For example, if the affine merge mode or the affine MVP mode is applied to the current block, the encoding device can derive a regression-based affine candidate for the current block based on surrounding sub-blocks of the current block.

[0264] For example, a regression-based affine candidate of the current block may be derived based on the surrounding sub-blocks of the current block. For example, a set of affine model parameters may be derived by inputting the motion vectors and center positions of the surrounding sub-blocks into a linear regression process, and CPMVs (Control Point Motion Vectors) of the current block derived as an affine model according to the set of affine model parameters may be derived as the regression-based affine candidates. That is, for example, a set of affine model parameters may be derived by inputting the motion vectors and center positions of the surrounding sub-blocks into a linear regression process, and the regression-based affine candidates may be CPMVs derived based on the derived set of affine model parameters. The affine model parameters may be calculated by solving a regression model into which the motion vectors and center positions of the surrounding sub-blocks are input from a mean square error (MSE) perspective, and CPMVs derived based on the affine model parameters may be derived as regression-based affine candidates.

[0265] For example, the peripheral sub-blocks may be peripheral sub-blocks included in an upper peripheral row and a left peripheral column of the current block. The size of the peripheral sub-blocks may be 4x4. The upper peripheral row may include upper left peripheral sub-blocks, upper peripheral sub-blocks, and upper right peripheral sub-blocks of the current block, and the left peripheral column may include lower left peripheral sub-blocks and left peripheral sub-blocks of the current block. That is, for example, the peripheral sub-blocks may include lower left peripheral sub-blocks, left peripheral sub-blocks, upper left peripheral sub-blocks, upper peripheral sub-blocks, and upper right peripheral sub-blocks. For example, the reference area including the upper right peripheral sub-blocks and the reference area including the upper left peripheral sub-blocks may be areas having a width equal to half the width of the current block (i.e., W / 2), and the reference area including the lower left peripheral sub-blocks may be an area having a height equal to half the height of the current block (i.e., H / 2).

[0266] Alternatively, as another example, the surrounding sub-blocks may be derived based on at least one of the size, width, height, width-to-height ratio, shape, type, or prediction mode of the current block.

[0267] For example, if the current block is a non-square block whose width is greater than its height, the peripheral sub-blocks may be peripheral sub-blocks included in an upper peripheral row of the current block. The size of the peripheral sub-blocks may be 4x4. The upper peripheral row may include upper left peripheral sub-blocks, upper peripheral sub-blocks, and upper right peripheral sub-blocks of the current block. That is, for example, if the current block is a non-square block whose width is greater than its height, the peripheral sub-blocks may include upper left peripheral sub-blocks, upper peripheral sub-blocks, and upper right peripheral sub-blocks of the current block. For example, the reference region including the upper right peripheral sub-blocks and the reference region including the upper left peripheral sub-blocks may be regions having a width equal to half of the width of the current block (i.e., W / 2).

[0268] Or, for example, if the current block is a non-square block whose width is greater than twice its height, the peripheral sub-blocks may be peripheral sub-blocks included in an upper peripheral row of the current block. The size of the peripheral sub-blocks may be 4x4. The upper peripheral row may include upper-left peripheral sub-blocks, upper-right peripheral sub-blocks, and upper-right peripheral sub-blocks of the current block. That is, for example, if the current block is a non-square block whose width is greater than twice its height, the peripheral sub-blocks may include upper-left peripheral sub-blocks, upper-right peripheral sub-blocks, and upper-right peripheral sub-blocks of the current block. For example, the reference region including the upper-right peripheral sub-blocks and the reference region including the upper-left peripheral sub-blocks may be regions having a width equal to half of the width of the current block (i.e., W / 2).

[0269] Or, for example, if the current block is a non-square block whose height is greater than its width, the peripheral sub-blocks may be peripheral sub-blocks included in a left peripheral column of the current block. The size of the peripheral sub-blocks may be 4x4. The left peripheral column may include lower left peripheral sub-blocks and left peripheral sub-blocks of the current block. That is, for example, if the current block is a non-square block whose height is greater than its width, the peripheral sub-blocks may include lower left peripheral sub-blocks, left peripheral sub-blocks, and upper left peripheral sub-blocks of the current block. For example, the reference region including the lower left peripheral sub-blocks may be a region whose height is half of the height of the current block (i.e., H / 2).

[0270] Alternatively, for example, if the current block is a non-square block whose height is greater than twice its width, the peripheral sub-blocks may be peripheral sub-blocks included in a left peripheral column of the current block. The size of the peripheral sub-blocks may be 4x4. The left peripheral column may include lower left peripheral sub-blocks and left peripheral sub-blocks of the current block. That is, for example, if the current block is a non-square block whose height is greater than twice its width, the peripheral sub-blocks may include lower left peripheral sub-blocks, left peripheral sub-blocks, and upper left peripheral sub-blocks of the current block. For example, the reference region including the lower left peripheral sub-blocks may be a region whose height is half of the height of the current block (i.e., H / 2).

[0271] The encoding device constructs an affine merge candidate list including the regression-based affine candidate of the current block (S1710).

[0272] The encoding device may construct an affine merge candidate list for the current block. The affine merge candidate list may include at least one candidate. The encoding device may construct an affine merge candidate list for the current block that includes the regression-based affine candidate.

[0273] Additionally, for example, the encoding device may add an inherited affine candidate and / or a constructed affine candidate to the affine merge candidate list. That is, the affine merge candidate list may include the inherited affine candidate and / or the constructed affine candidate.

[0274] Additionally, for example, the encoding device can derive a second regression-based affine candidate of the current block based on sub-blocks of an affine block coded before the current block, and add the second regression-based affine candidate to the affine merge candidate list. That is, the affine merge candidate list can include the second regression-based affine candidate. Here, the previously coded affine block can mean a coding block coded based on affine prediction before the current block is coded. For example, the previously coded affine block can be derived by scanning a location that is not adjacent to the current block and an affine HMVP table.

[0275] For example, an affine model parameter set can be derived by inputting the motion vectors and center positions of the sub-blocks of the previously coded affine block into a linear regression process, and CPMVs (Control Point Motion Vectors) of the current block derived as an affine model according to the affine model parameter set can be derived as the regression-based affine candidates. That is, for example, an affine model parameter set can be derived by inputting the motion vectors and center positions of the sub-blocks of the previously coded affine block into a linear regression process, and the regression-based affine candidates can be CPMVs derived based on the derived affine model parameter set. The affine model parameters can be calculated by solving a regression model into which the motion vectors and center positions of the sub-blocks of the previously coded affine block are input from a mean square error (MSE) perspective, and CPMVs derived based on the affine model parameters can be derived as the regression-based affine candidates.

[0276] Additionally, for example, up to two regression-based affine candidates can be derived. For example, the regression-based affine candidate can be derived based on surrounding sub-blocks of the current block, and one regression-based affine candidate (i.e., a second regression-based affine candidate) can be derived based on sub-blocks of the previously coded affine block.

[0277] Also, for example, whether or not to derive the second regression-based affine candidate may be determined based on the shape of the current block. Or, for example, whether or not to derive the second regression-based affine candidate may be determined based on at least one of the size, width, height, width-to-height ratio, shape, type, or prediction mode of the current block. For example, if the size of the current block is larger than a specific size, the second regression-based affine candidate may not be derived. Or, for example, if the size of the current block is smaller than a specific size, the second regression-based affine candidate may not be derived. Or, for example, if the current block is a non-square block, the second regression-based affine candidate may not be derived.

[0278] Additionally, for example, the affine merge candidate list may be reordered based on template matching costs. For example, template matching costs for candidates in the affine merge candidate list may be derived, and the candidates may be reordered based on the template matching costs. For example, the candidates may be reordered in descending order of template matching costs. Alternatively, the candidates in the affine merge candidate list may be divided into multiple subgroups, and the candidates may be reordered within the subgroups based on the template matching costs. In this case, the regression-based affine candidate derived based on the surrounding subblocks and the second regression-based affine candidate may be included in the same subgroup.

[0279] The encoding device derives motion information of sub-blocks of the current block based on the above affine merge candidate list (S1720).

[0280] The encoding device can derive motion information of sub-blocks of the current block based on the affine merge candidate list. For example, the encoding device can select one of the candidates in the affine merge candidate list, and derive motion information of sub-blocks of the current block based on the selected candidate. For example, the encoding device can derive motion vectors of sub-blocks of the current block based on CPMVs of the selected candidate. That is, the encoding device can derive motion vectors of each sub-block of the current block based on the CPMVs. The motion vectors can be represented as an affine motion vector field (MVF) or a motion vector array.

[0281] The encoding device derives prediction samples for the current block based on the motion information of the sub-blocks (S1730). The encoding device can derive prediction samples for the current block based on the motion information of the sub-blocks. The encoding device can derive prediction samples for the current block by performing prediction based on the motion information of the sub-blocks. That is, the encoding device can derive a reference region within a reference picture based on the motion information of the sub-blocks, and can generate prediction samples for the sub-blocks of the current block based on reconstructed samples within the reference region. In this case, as described above, a prediction sample filtering procedure may be further performed on all or part of the prediction samples of the current block, as the case may be.

[0282] The encoding device encodes image information including prediction-related information of the current block (S1740). The encoding device can encode image information including prediction-related information of the current block and signal it via a bitstream. That is, the encoding device can output image information including prediction-related information for the current block in the form of a bitstream.

[0283] For example, the prediction-related information may include an affine merge flag indicating whether the affine merge mode is applied to the current block. When the value of the affine merge flag is 1, the affine merge flag may indicate that the affine merge mode is applied to the current block, and when the value of the affine merge flag is 0, the affine merge flag may indicate that the affine MVP mode is applied to the current block. The affine merge flag may also be referred to as a merge subblock flag or a subblock merge flag.

[0284] Additionally, for example, the prediction-related information may include a candidate index for the current block. The candidate index may indicate one of the candidates included in the affine merge candidate list. In other words, the candidate index may indicate a selected candidate among the candidates included in the affine merge candidate list.

[0285] Additionally, the image information may include various information according to embodiments of the present disclosure. For example, the image information may include information disclosed in at least one of the tables described above.

[0286] Meanwhile, the image information may include residual information. The residual information is information about residual samples. The residual information may include information about quantized transform coefficients for the residual samples.

[0287] Encoded image information can be output in the form of a bitstream. The bitstream can be transmitted to a decoding device via a network or storage medium.

[0288] In addition, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and a reconstructed block) based on the reference samples and the residual samples. This is to derive the same prediction result as that performed by the decoding device from the encoding device, thereby increasing coding efficiency. Accordingly, the encoding device can store the reconstructed picture (or reconstructed samples, reconstructed block) in memory and use it as a reference picture for inter prediction. As described above, an in-loop filtering procedure, etc. can be further applied to the reconstructed picture.

[0289] According to the above-described embodiment(s), in affine prediction, motion information of a sub-block unit of a current block can be derived by considering motion information of surrounding sub-blocks adjacent to the current block, thereby improving the accuracy and coding efficiency of affine prediction.

[0290] In addition, when deriving a regression-based affine candidate based on the movement information of surrounding sub-blocks, the shape of the current block can be taken into consideration to derive the regression-based affine candidate only from surrounding blocks with high correlation, thereby improving the accuracy of the regression-based affine candidate.

[0291] FIG. 18 schematically illustrates a video / image decoding method according to an embodiment(s) of the present disclosure. The method disclosed in FIG. 18 may be performed by the decoding device disclosed in FIG. 3 . Specifically, for example, steps S1800 to S1830 of FIG. 18 may be performed by the prediction unit (330) of the decoding device (300). The method disclosed in FIG. 18 may include the embodiments described above in the present disclosure.

[0292] Referring to FIG. 18, the decoding device derives a regression-based affine candidate of the current block based on surrounding sub-blocks of the current block (S1800).

[0293] The decoding device can derive a regression-based affine candidate for the current block based on the surrounding sub-blocks of the current block. For example, if the affine merge mode or the affine MVP mode is applied to the current block, the decoding device can derive a regression-based affine candidate for the current block based on the surrounding sub-blocks of the current block.

[0294] For example, a decoding device can obtain image information including prediction-related information through a bitstream. The image information may further include residual information as described above. The prediction-related information may include an affine merge flag indicating whether an affine merge mode is applied to the current block. When the value of the affine merge flag is 1, the affine merge flag may indicate that the affine merge mode is applied to the current block, and when the value of the affine merge flag is 0, the affine merge flag may indicate that the affine MVP mode is applied to the current block. The affine merge flag may also be referred to as a merge subblock flag or a subblock merge flag.

[0295] For example, a regression-based affine candidate of the current block may be derived based on the surrounding sub-blocks of the current block. For example, a set of affine model parameters may be derived by inputting the motion vectors and center positions of the surrounding sub-blocks into a linear regression process, and CPMVs (Control Point Motion Vectors) of the current block derived as an affine model according to the set of affine model parameters may be derived as the regression-based affine candidates. That is, for example, a set of affine model parameters may be derived by inputting the motion vectors and center positions of the surrounding sub-blocks into a linear regression process, and the regression-based affine candidates may be CPMVs derived based on the derived set of affine model parameters. The affine model parameters may be calculated by solving a regression model into which the motion vectors and center positions of the surrounding sub-blocks are input from a mean square error (MSE) perspective, and CPMVs derived based on the affine model parameters may be derived as regression-based affine candidates.

[0296] For example, the peripheral sub-blocks may be peripheral sub-blocks included in an upper peripheral row and a left peripheral column of the current block. The size of the peripheral sub-blocks may be 4x4. The upper peripheral row may include upper left peripheral sub-blocks, upper peripheral sub-blocks, and upper right peripheral sub-blocks of the current block, and the left peripheral column may include lower left peripheral sub-blocks and left peripheral sub-blocks of the current block. That is, for example, the peripheral sub-blocks may include lower left peripheral sub-blocks, left peripheral sub-blocks, upper left peripheral sub-blocks, upper peripheral sub-blocks, and upper right peripheral sub-blocks. For example, the reference area including the upper right peripheral sub-blocks and the reference area including the upper left peripheral sub-blocks may be areas having a width equal to half the width of the current block (i.e., W / 2), and the reference area including the lower left peripheral sub-blocks may be an area having a height equal to half the height of the current block (i.e., H / 2).

[0297] Alternatively, as another example, the surrounding sub-blocks may be derived based on at least one of the size, width, height, width-to-height ratio, shape, type, or prediction mode of the current block.

[0298] For example, if the current block is a non-square block whose width is greater than its height, the peripheral sub-blocks may be peripheral sub-blocks included in an upper peripheral row of the current block. The size of the peripheral sub-blocks may be 4x4. The upper peripheral row may include upper left peripheral sub-blocks, upper peripheral sub-blocks, and upper right peripheral sub-blocks of the current block. That is, for example, if the current block is a non-square block whose width is greater than its height, the peripheral sub-blocks may include upper left peripheral sub-blocks, upper peripheral sub-blocks, and upper right peripheral sub-blocks of the current block. For example, the reference region including the upper right peripheral sub-blocks and the reference region including the upper left peripheral sub-blocks may be regions having a width equal to half of the width of the current block (i.e., W / 2).

[0299] Or, for example, if the current block is a non-square block whose width is greater than twice its height, the peripheral sub-blocks may be peripheral sub-blocks included in an upper peripheral row of the current block. The size of the peripheral sub-blocks may be 4x4. The upper peripheral row may include upper-left peripheral sub-blocks, upper-right peripheral sub-blocks, and upper-right peripheral sub-blocks of the current block. That is, for example, if the current block is a non-square block whose width is greater than twice its height, the peripheral sub-blocks may include upper-left peripheral sub-blocks, upper-right peripheral sub-blocks, and upper-right peripheral sub-blocks of the current block. For example, the reference region including the upper-right peripheral sub-blocks and the reference region including the upper-left peripheral sub-blocks may be regions having a width equal to half of the width of the current block (i.e., W / 2).

[0300] Or, for example, if the current block is a non-square block whose height is greater than its width, the peripheral sub-blocks may be peripheral sub-blocks included in a left peripheral column of the current block. The size of the peripheral sub-blocks may be 4x4. The left peripheral column may include lower left peripheral sub-blocks and left peripheral sub-blocks of the current block. That is, for example, if the current block is a non-square block whose height is greater than its width, the peripheral sub-blocks may include lower left peripheral sub-blocks, left peripheral sub-blocks, and upper left peripheral sub-blocks of the current block. For example, the reference region including the lower left peripheral sub-blocks may be a region whose height is half of the height of the current block (i.e., H / 2).

[0301] Alternatively, for example, if the current block is a non-square block whose height is greater than twice its width, the peripheral sub-blocks may be peripheral sub-blocks included in a left peripheral column of the current block. The size of the peripheral sub-blocks may be 4x4. The left peripheral column may include lower left peripheral sub-blocks and left peripheral sub-blocks of the current block. That is, for example, if the current block is a non-square block whose height is greater than twice its width, the peripheral sub-blocks may include lower left peripheral sub-blocks, left peripheral sub-blocks, and upper left peripheral sub-blocks of the current block. For example, the reference region including the lower left peripheral sub-blocks may be a region whose height is half of the height of the current block (i.e., H / 2).

[0302] The decoding device constructs an affine merge candidate list including the regression-based affine candidate of the current block (S1810).

[0303] The decoding device may construct an affine merge candidate list for the current block. The affine merge candidate list may include at least one candidate. The decoding device may construct an affine merge candidate list for the current block that includes the regression-based affine candidate.

[0304] Additionally, for example, the decoding device may add an inherited affine candidate and / or a constructed affine candidate to the affine merge candidate list. That is, the affine merge candidate list may include the inherited affine candidate and / or the constructed affine candidate.

[0305] Additionally, for example, the decoding device can derive a second regression-based affine candidate of the current block based on sub-blocks of an affine block coded before the current block, and add the second regression-based affine candidate to the affine merge candidate list. That is, the affine merge candidate list can include the second regression-based affine candidate. Here, the previously coded affine block can mean a coding block coded based on affine prediction before the current block is coded. For example, the previously coded affine block can be derived by scanning a location that is not adjacent to the current block and an affine HMVP table.

[0306] For example, an affine model parameter set can be derived by inputting the motion vectors and center positions of the sub-blocks of the previously coded affine block into a linear regression process, and CPMVs (Control Point Motion Vectors) of the current block derived as an affine model according to the affine model parameter set can be derived as the regression-based affine candidates. That is, for example, an affine model parameter set can be derived by inputting the motion vectors and center positions of the sub-blocks of the previously coded affine block into a linear regression process, and the regression-based affine candidates can be CPMVs derived based on the derived affine model parameter set. The affine model parameters can be calculated by solving a regression model into which the motion vectors and center positions of the sub-blocks of the previously coded affine block are input from a mean square error (MSE) perspective, and CPMVs derived based on the affine model parameters can be derived as the regression-based affine candidates.

[0307] Additionally, for example, up to two regression-based affine candidates can be derived. For example, the regression-based affine candidate can be derived based on surrounding sub-blocks of the current block, and one regression-based affine candidate (i.e., a second regression-based affine candidate) can be derived based on sub-blocks of the previously coded affine block.

[0308] Also, for example, whether or not to derive the second regression-based affine candidate may be determined based on the shape of the current block. Or, for example, whether or not to derive the second regression-based affine candidate may be determined based on at least one of the size, width, height, width-to-height ratio, shape, type, or prediction mode of the current block. For example, if the size of the current block is larger than a specific size, the second regression-based affine candidate may not be derived. Or, for example, if the size of the current block is smaller than a specific size, the second regression-based affine candidate may not be derived. Or, for example, if the current block is a non-square block, the second regression-based affine candidate may not be derived.

[0309] Additionally, for example, the affine merge candidate list may be reordered based on template matching costs. For example, template matching costs for candidates in the affine merge candidate list may be derived, and the candidates may be reordered based on the template matching costs. For example, the candidates may be reordered in descending order of template matching costs. Alternatively, the candidates in the affine merge candidate list may be divided into multiple subgroups, and the candidates may be reordered within the subgroups based on the template matching costs. In this case, the regression-based affine candidate derived based on the surrounding subblocks and the second regression-based affine candidate may be included in the same subgroup.

[0310] The decoding device derives motion information of sub-blocks of the current block based on the above affine merge candidate list (S1820).

[0311] The decoding device can derive motion information of sub-blocks of the current block based on the affine merge candidate list. For example, the decoding device can select one of the candidates in the affine merge candidate list, and derive motion information of sub-blocks of the current block based on the selected candidate. For example, the decoding device can derive motion vectors of sub-blocks of the current block based on CPMVs of the selected candidate. That is, the decoding device can derive motion vectors of each sub-block of the current block based on the CPMVs. The motion vectors can be represented as an affine motion vector field (MVF) or a motion vector array.

[0312] Additionally, for example, the decoding device can obtain a candidate index for the current block from the bitstream, and select a candidate indicated by the candidate index from among the candidates included in the affine merge candidate list. The decoding device can obtain image information from the bitstream, and the image information can include prediction-related information for the current block. The prediction-related information can include the candidate index.

[0313] The decoding device derives prediction samples for the current block based on the motion information of the sub-blocks (S1830). The decoding device can derive prediction samples for the current block based on the motion information of the sub-blocks. The decoding device can derive prediction samples for the current block by performing prediction based on the motion information of the sub-blocks. That is, the decoding device can derive a reference region within a reference picture based on the motion information of the sub-blocks, and can generate prediction samples for the sub-blocks of the current block based on reconstructed samples within the reference region. In this case, as described above, a prediction sample filtering procedure may be further performed on all or part of the prediction samples of the current block, depending on the case.

[0314] The decoding device can generate reconstructed samples based on prediction samples of the current block. For example, the decoding device can generate the reconstructed samples for the current block based on residual samples for the current block and the prediction samples. The residual samples for the current block can be generated based on received residual information. In addition, the decoding device can generate a reconstructed picture including the reconstructed samples, for example. As described above, an in-loop filtering procedure, etc. can be further applied to the reconstructed picture.

[0315] According to the above-described embodiment(s), in affine prediction, motion information of a sub-block unit of a current block can be derived by considering motion information of surrounding sub-blocks adjacent to the current block, thereby improving the accuracy and coding efficiency of affine prediction.

[0316] In addition, when deriving a regression-based affine candidate based on the movement information of surrounding sub-blocks, the shape of the current block can be taken into consideration to derive the regression-based affine candidate only from surrounding blocks with high correlation, thereby improving the accuracy of the regression-based affine candidate.

[0317] In the embodiments described above, the methods are described based on a flowchart as a series of steps or blocks. However, the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will understand that the steps depicted in the flowchart are not exclusive, and other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of this document.

[0318] The method according to the embodiments of the present document described above can be implemented in the form of software, and the encoding device and / or decoding device according to the present document can be included in a device that performs image processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.

[0319] The embodiments of the present disclosure described above may also be implemented in the form of a recording medium containing computer-executable (program) instructions, such as program modules, executed by a computer. The modules may be stored in a memory and executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. Computer-readable media may be any available media that can be accessed by a computer, and includes both volatile and nonvolatile media, removable and non-removable media. Furthermore, computer-readable media may include both computer storage media and communication media. Computer storage media includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Communication media typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, or other transport mechanism, and includes any information delivery media.

[0320] In addition, the embodiments of the present disclosure described above may be implemented as a computer program (or computer program product) including computer-executable instructions. The computer program includes programmable machine instructions processed by the processor, and may be implemented in a high-level programming language, an object-oriented programming language, assembly language, or machine language. In addition, the computer program may be recorded on a tangible computer-readable recording medium (e.g., memory, a hard disk, a magnetic / optical medium, or a solid-state drive (SSD), etc.).

[0321] Accordingly, the embodiments of the present disclosure described above can be implemented by executing the computer program described above on a computing device. The computing device may include a processor, memory, a storage device, a high-speed interface connecting the memory and a high-speed expansion port, and at least some of a low-speed interface connecting the low-speed bus and the storage device. Each of these components is connected to one another using various buses and may be mounted on a common motherboard or in another suitable manner.

[0322] Here, the processor can process instructions within the computing device, such as instructions stored in a memory or storage device to display graphical information for providing a graphical user interface (GUI) on an external input / output device, such as a display connected to a high-speed interface. In another embodiment, multiple processors and / or multiple buses may be utilized, as appropriate, together with multiple memories and memory types. The processor may also be implemented as a chipset comprising multiple independent analog and / or digital processors.

[0323] Memory also stores information within a computing device. For example, memory may consist of volatile memory units or a collection of volatile memory units. For another example, memory may consist of nonvolatile memory units or a collection of nonvolatile memory units. Memory may also be another form of computer-readable media, such as magnetic or optical disks.

[0324] A storage device can provide a large amount of storage space to a computing device. The storage device can be a computer-readable medium or a configuration including such a medium, and can include, for example, devices within a storage area network (SAN) or other configurations, and can be a floppy disk device, a hard disk device, an optical disk device, a tape device, flash memory, or other similar semiconductor memory device or device array.

[0325] Additionally, the network may be implemented as a wired network such as a Local Area Network (LAN), a Wide Area Network (WAN), or a Value Added Network (VAN), or as a wireless network of various types such as a mobile radio communication network or a satellite communication network.

[0326] Although the present disclosure has been described above with reference to the embodiments illustrated in the drawings, these are merely exemplary, and those skilled in the art will understand that various modifications and variations of the embodiments are possible from the above-described embodiments. In other words, the scope of the present disclosure is not limited to the above-described embodiments, and various modifications and improvements made by those skilled in the art using the basic concepts of the embodiments defined in the following claims also fall within the scope of the embodiments. Therefore, the true technical protection scope of the present disclosure should be determined by the technical spirit of the appended claims.

Claims

1. In a video decoding method performed by a decoding device, A step of deriving a regression-based affine candidate of the current block based on surrounding sub-blocks of the current block; A step of constructing an affine merge candidate list including the regression-based affine candidates of the current block; A step of deriving motion information of sub-blocks of the current block based on the above affine merge candidate list; and An image decoding method, characterized by comprising a step of deriving prediction samples for the current block based on motion information of the sub-blocks.

2. In paragraph 1, The motion vectors and center positions of the above surrounding sub-blocks are input into a linear regression process to derive a set of affine model parameters. An image decoding method, characterized in that the above regression-based affine candidates are CPMVs derived based on the above derived affine model parameter set.

3. In paragraph 1, An image decoding method, characterized in that the above peripheral sub-blocks include lower left peripheral sub-blocks, left peripheral sub-blocks, upper left peripheral sub-blocks, upper right peripheral sub-blocks, and upper right peripheral sub-blocks.

4. In paragraph 3, The reference area including the upper right peripheral sub-blocks and the reference area including the upper left peripheral sub-blocks are areas whose width is half the width of the current block, An image decoding method, characterized in that the reference area including the lower left peripheral sub-blocks is an area having a height equal to half the height of the current block.

5. In paragraph 3, An image decoding method, characterized in that the size of the above surrounding sub-blocks is 4x4.

6. In paragraph 1, An image decoding method, characterized in that the surrounding sub-blocks are derived based on the shape of the current block.

7. In paragraph 6, An image decoding method, characterized in that when the current block is a non-square block whose width is greater than its height, the surrounding sub-blocks include upper left surrounding sub-blocks, upper right surrounding sub-blocks, and upper right surrounding sub-blocks of the current block.

8. In paragraph 6, An image decoding method, characterized in that when the current block is a non-square block whose height is greater than its width, the surrounding sub-blocks include lower left surrounding sub-blocks, left surrounding sub-blocks, and upper left surrounding sub-blocks of the current block.

9. In paragraph 1, A step of deriving a second regression-based affine candidate of the current block based on the sub-blocks of the affine block coded before the current block; and An image decoding method, characterized by further comprising the step of adding the second regression-based affine candidate to the affine merge candidate list.

10. In paragraph 9, An image decoding method, characterized in that it is determined whether the second regression-based affine candidate is derived based on the shape of the current block.

11. In a video encoding method performed by an encoding device, A step of deriving a regression-based affine candidate of the current block based on surrounding sub-blocks of the current block; A step of constructing an affine merge candidate list including the regression-based affine candidates of the current block; A step of deriving motion information of sub-blocks of the current block based on the above affine merge candidate list; A step of deriving prediction samples for the current block based on the motion information of the above sub-blocks; and An image encoding method, characterized by comprising a step of encoding image information including prediction-related information of the current block.

12. In paragraph 11, The motion vectors and center positions of the above surrounding sub-blocks are input into a linear regression process to derive a set of affine model parameters. An image encoding method, characterized in that the above regression-based affine candidates are CPMVs derived based on the above derived affine model parameter set.

13. In paragraph 11, An image encoding method, characterized in that the above peripheral sub-blocks include lower left peripheral sub-blocks, left peripheral sub-blocks, upper left peripheral sub-blocks, upper right peripheral sub-blocks, and upper right peripheral sub-blocks.

14. In paragraph 13, The reference area including the upper right peripheral sub-blocks and the reference area including the upper left peripheral sub-blocks are areas whose width is half the width of the current block, An image encoding method, characterized in that the reference area including the lower left peripheral sub-blocks is an area having a height equal to half the height of the current block.

15. In paragraph 13, An image encoding method, characterized in that the size of the above surrounding sub-blocks is 4x4.

16. In paragraph 11, An image encoding method, characterized in that the surrounding sub-blocks are derived based on the shape of the current block.

17. In paragraph 16, An image encoding method, characterized in that when the current block is a non-square block whose width is greater than its height, the surrounding sub-blocks include upper left surrounding sub-blocks, upper right surrounding sub-blocks, and upper right surrounding sub-blocks of the current block.

18. In paragraph 16, An image encoding method, characterized in that when the current block is a non-square block whose height is greater than its width, the surrounding sub-blocks include lower left surrounding sub-blocks, left surrounding sub-blocks, and upper left surrounding sub-blocks of the current block.

19. In paragraph 1, A step of deriving a second regression-based affine candidate of the current block based on the sub-blocks of the affine block coded before the current block; and An image encoding method, characterized in that it further comprises a step of adding the second regression-based affine candidate to the affine merge candidate list.

20. In a transmission method for video data, A method for obtaining a bitstream generated by an image encoding method, the method comprising: a step of deriving a regression-based affine candidate of a current block based on surrounding sub-blocks of the current block; a step of configuring an affine merge candidate list including the regression-based affine candidate of the current block; a step of deriving motion information of sub-blocks of the current block based on the affine merge candidate list; a step of deriving prediction samples for the current block based on the motion information of the sub-blocks; and a step of encoding image information including prediction-related information of the current block; and A transmission method, characterized by comprising a step of transmitting image data including the bitstream.

Citation Information

Patent Citations

  • Image prediction method and related device

    KR1020170045264A

  • Thermal request mediating device, thermal request mediating method and vehicle

    KR1020220011198A

  • Cream composition for paper incense and paper incense using thereof

    KR102323824B1

  • Method and apparatus for video coding

    US20200389663A1

  • Method and apparatus for regression-based affine merge mode motion vector derivation in video coding systems

    WO2023202713A1