Video or image coding based on time-motion vector predictor candidates at the subblock level.

By deriving subblock-based temporal motion vectors from reference subblocks, the method addresses inefficient motion vector prediction in high-resolution video, improving compression efficiency and reducing complexity in video encoding and decoding processes.

JP2026104986APending Publication Date: 2026-06-25LG ELECTRONICS INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2026-04-16
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality video content, including VR and AR, necessitates a highly efficient video compression technology to reduce transmission and storage costs, while existing methods struggle with efficient motion vector prediction in sub-block units.

Method used

The method involves deriving subblock-based temporal motion vectors using reference subblocks on a collocated reference picture, integrating corresponding positions at subcoding and coding block levels, and applying these techniques in video/image encoding and decoding processes.

Benefits of technology

This approach improves video compression efficiency, reduces computational complexity, and simplifies hardware implementation by efficiently calculating subblock positions for motion vector prediction, enhancing overall coding efficiency and prediction performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026104986000001_ABST
    Figure 2026104986000001_ABST
Patent Text Reader

Abstract

This invention provides a method and apparatus for improving the efficiency of video / image coding. [Solution] A video decoding method performed by a decoding device, comprising: deriving a reference subblock in a collated reference picture for a subblock in the current block; deriving motion information for a subblock in the current block based on candidate time motion information based on the subblock derived based on the reference subblock; generating a reconstructed sample based on a predicted sample generated based on the motion information and a residual sample generated based on residual information obtained from a bitstream; the reference subblock is derived from the collated reference picture by applying a motion shift to the center sample position of the right lower end of each subblock in the current block; the motion shift is performed based on a motion vector derived from a spatially adjacent block of the current block; and the spatially adjacent block is the left adjacent block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This technology relates to video or video coding, for example, to video or video coding technology based on a temporal motion vector predictor candidate in sub-block units.

Background Art

[0002] Recently, the demand for high-resolution and high-quality video / video such as 4K or UHD (Ultra High Definition) video / video of 8K or higher has been increasing in various fields. As the video / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to existing video / video data. Therefore, when transmitting video data using media such as existing wired or wireless broadband lines, or storing (storing) video / video data using existing storage media, the transmission cost (cost) and storage cost increase.

[0003] Also, recently, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcasting of videos / videos having video characteristics different from real-world videos, such as game videos, has been increasing.

[0004] Therefore, in order to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality videos / videos having various characteristics as described above, a highly efficient video / video compression technology is required.

[0005] Also, in order to improve the video / video coding efficiency, there has been discussion on the temporal motion vector prediction technology in sub-block units. For this purpose, a method for efficiently performing the process of patching the motion vector in sub-block units by temporal motion vector prediction in sub-block units is required.

Summary of the Invention

[0006] The technical objective of this document is to provide methods and apparatus for improving the efficiency of video / image coding.

[0007] Another technical challenge of this document is to provide an efficient interpretation method and apparatus.

[0008] Another technical problem of the present invention is to provide a method and apparatus for deriving time-motion vectors based on subblocks to improve prediction performance.

[0009] A further technical problem of the present invention is to provide a method and apparatus for efficiently deriving the corresponding positions of subblocks for deriving (inducing) time motion vectors based on subblocks.

[0010] A further technical problem of the present invention is to provide a method and apparatus for integrating corresponding positions at the subcoding block level and corresponding positions at the coding block level for deriving a time motion vector based on subblocks. [Means for solving the problem]

[0011] According to one embodiment of this paper, in subblock-based temporal motion vector prediction (sbTMVP), the subblock unit motion vector relative to the current block can be derived based on a reference subblock on a collocated reference picture.

[0012] According to one embodiment of this document, reference subblocks on a collated reference picture for multiple subblocks within the current block can be derived based on the center sample position of each subblock within the current block.

[0013] According to one embodiment of this document, a base motion vector can be used for the subblock unit motion vector for unusable reference subblocks in a reference subblock.

[0014] According to one embodiment of this document, the base motion vector can be derived on (from) the collated reference picture based on the center sample position of the current block.

[0015] According to one embodiment of this document, a video / image decoding method is provided that is performed by a decoding (decoding, decryption) device. The video / image decoding method may have the method disclosed in the embodiments of this document.

[0016] According to one embodiment of this document, a decoding device for performing video / image decoding is provided. The decoding device can perform the methods disclosed in the embodiments of this document.

[0017] According to one embodiment of this document, a video / image encoding method is provided that is performed by an encoding device. The video / image encoding method may have the method disclosed in the embodiments of this document.

[0018] According to one embodiment of this document, an encoding device for performing video / image encoding is provided. The encoding device can perform the methods disclosed in the embodiments of this document.

[0019] According to one embodiment of this document, a computer-readable digital storage medium is provided which stores encoded video / image information generated according to a video / image encoding method disclosed in at least one embodiment of this document.

[0020] According to one embodiment of this document, a computer-readable digital storage medium is provided which stores encoded information or encoded video / image information, which triggers a decoding device to perform a video / image decoding method disclosed in at least one embodiment of this document. [Effects of the Invention]

[0021] This document can achieve a variety of effects. For example, it can improve overall video compression efficiency. It can also reduce computational complexity through efficient interpretation, thereby improving overall coding efficiency. Furthermore, by efficiently calculating the corresponding subblock positions for deriving subblock-based temporal motion vectors in subblock-based temporal motion vector prediction (sbTMVP), it can improve efficiency in terms of complexity and prediction performance. In addition, by integrating methods for calculating corresponding subcoding block-level positions and coding block-level positions for deriving subblock-based temporal motion vectors, it can simplify hardware implementation.

[0022] The effects obtained through the specific embodiments of this document are not limited to those listed above. For example, there may be a variety of technical effects that a person having ordinary skill in the related art can understand or derive from this document. Thus, the specific effects of this document are not limited to those explicitly stated herein, but may include a variety of effects that can be understood or derive from the technical features of this document. [Brief explanation of the drawing]

[0023] [Figure 1] This figure schematically illustrates an example of a video / image coding system applicable to the embodiments described herein. [Figure 2] It is a diagram schematically explaining the configuration of a video / video encoding device to which an embodiment of this document can be applied. [Figure 3] It is a diagram schematically explaining the configuration of a video / video decoding device to which an embodiment of this document can be applied. [Figure 4] It is a diagram showing an example of a schematic video / video encoding method to which an embodiment of this document is applicable. [Figure 5] It is a diagram showing an example of a schematic video / video decoding method to which an embodiment of this document is applicable. [Figure 6] It is a diagram showing an example of a schematic video / video encoding method based on inter prediction to which an embodiment of this document is applicable. [Figure 7] It is a diagram showing an example of a schematic video / video decoding method based on inter prediction to which an embodiment of this document is applicable. [Figure 8] It is a diagram exemplarily showing an inter prediction procedure. [Figure 9] It is a diagram exemplarily showing spatial adjacent blocks and temporal adjacent blocks of a current block. [Figure 10] It is a diagram exemplarily showing spatial adjacent blocks that can be used to derive sub-block based temporal motion information candidates (sbTMVP candidates). [Figure 11] It is a diagram schematically explaining the process of deriving sub-block based temporal motion information candidates (sbTMVP candidates). [Figure 12] It is a diagram schematically explaining a method of calculating corresponding positions for deriving default MVs and sub-block MVs according to block sizes in the sbTMVP derivation process. [Figure 13] It is a diagram schematically explaining a method of calculating corresponding positions for deriving default MVs and sub-block MVs according to block sizes in the sbTMVP derivation process. [Figure 14]This diagram schematically illustrates the method for calculating the corresponding position for deriving the default MV and sub-block MV based on the block size during the sbTMVP derivation process. [Figure 15] This diagram schematically illustrates the method for calculating the corresponding position for deriving the default MV and sub-block MV based on the block size during the sbTMVP derivation process. [Figure 16] This is an illustrative diagram illustrating a schematic method for integrating corresponding positions for deriving the default MV and sub-block MV based on the block size during the sbTMVP derivation process. [Figure 17] This is an illustrative diagram illustrating a schematic method for integrating corresponding positions for deriving the default MV and sub-block MV based on the block size during the sbTMVP derivation process. [Figure 18] This is an illustrative diagram illustrating a schematic method for integrating corresponding positions for deriving the default MV and sub-block MV based on the block size during the sbTMVP derivation process. [Figure 19] This is an illustrative diagram illustrating a schematic method for integrating corresponding positions for deriving the default MV and sub-block MV based on the block size during the sbTMVP derivation process. [Figure 20] This diagram schematically illustrates a pipeline configuration that allows for the integrated calculation of corresponding positions for deriving the default MV and subblock MV during the sbTMVP derivation process. [Figure 21] This diagram schematically illustrates a pipeline configuration that allows for the integrated calculation of corresponding positions for deriving the default MV and subblock MV during the sbTMVP derivation process. [Figure 22]This figure schematically illustrates an example of a video / image encoding method according to the embodiment of this document. [Figure 23] This figure schematically illustrates an example of a video / image decoding method according to the embodiment of this document. [Figure 24] This figure shows an example of a content streaming system to which the embodiments disclosed in this document may be applied. [Modes for carrying out the invention]

[0024] This document is subject to various modifications and may have many different embodiments. Specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit this document to any particular embodiment. Terms used in this document are used solely to describe specific embodiments and are not intended to limit the technical ideas of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "includes" or "has" in this document are intended to indicate the existence of features, figures, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood not to preemptively exclude the existence or possibility of adding one or more other features, figures, steps, actions, components, parts, or combinations thereof.

[0025] On the other hand, each configuration shown in the diagrams described in this document is illustrated independently for the purpose of explaining its distinct characteristic functions, and does not mean that each configuration is embodied in separate hardware or separate software. For example, two or more configurations may be combined to form a single configuration, and a single configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included within the scope of this document as long as they do not deviate from the essence of this document.

[0026] In this document, “A or B” can mean “only A,” “only B,” or “both A and B.” Furthermore, “A or B” can be interpreted as “A and / or B.” For example, in this document, “A, B or C” can mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.”

[0027] The slashes ( / ) and commas used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "just A", "just B", or "both A and B". For example, "A, B, C" can mean "A, B or C".

[0028] In this document, "at least one of A and B" can mean "just A," "just B," or "both A and B." Furthermore, in this document, the expressions "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B."

[0029] Furthermore, in this document, "at least one of A, B and C" can mean "just A," "just B," "just C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" can mean "at least one of A, B and C."

[0030] Furthermore, the parentheses used in this document can mean "for example." Specifically, when displayed as "prediction (intra prediction)," "intra prediction" is proposed as an example of "prediction." In other words, "prediction" in this document is not limited to "intra prediction," but rather "intra prediction" is proposed as an example of "prediction." Similarly, when displayed as "prediction (i.e., intra prediction)," "intra prediction" is proposed as an example of "prediction."

[0031] This document relates to video / image coding. For example, the methods / examples disclosed in this document can be applied to methods disclosed in the VVC (Versatile Video Coding) standard. Furthermore, the methods / examples disclosed in this document can be applied to methods disclosed in the EVC (Essential Video Coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd Generation Of Audio Video Coding Standard), or next-generation video / image coding standards (e.g., H.267 or H.267).

[0032] This document presents various examples of video / image coding, and unless otherwise noted, these examples can be combined and implemented in conjunction with each other.

[0033] In this document, "video" can mean a collection of images over time. "Picture" generally refers to a unit representing a single image at a specific time point in time, and "slice" or "tile" is a unit that constitutes part of a picture in coding. A slice or tile can contain one or more Coding Tree Units (CTUs). A single picture can consist of one or more slices or tiles. A tile is a rectangular region of CTUs within a particular tile column and particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture, and a width specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a height equal to the height of the picture.A tile scan can represent a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice includes an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that may be exclusively contained in a single NAL unit.

[0034] On the other hand, a single picture can be divided into two or more subpictures. A subpicture is a rectangular region of one or more slices within a picture.

[0035] A pixel or pel can refer to the smallest unit that makes up a picture (or image). Alternatively, the term "sample" can be used as a counterpart to pixel. A sample can generally represent a pixel or a pixel value, or it can represent only the luma component pixel / pixel value, or only the chroma component pixel / pixel value. Alternatively, a sample can refer to a pixel value in the spatial domain, and when such a pixel value is converted to the frequency domain, it can also refer to the conversion coefficient in the frequency domain.

[0036] A unit can represent a basic unit of image processing. A unit can contain at least one of a specific region of a picture and information associated with that region. A unit can contain one luma block and two chroma (e.g., cb, cr) blocks. The term unit may sometimes be used interchangeably with terms such as block or area. In general, an M×N block can contain a sample (or sample array) consisting of M columns and N rows, or a set (or array) of transform coefficients.

[0037] Furthermore, in this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. If quantization / inverse quantization is omitted, the quantized transformation coefficients may be called transformation coefficients. If transformation / inverse transformation is omitted, the transformation coefficients may also be called coefficients or residual coefficients, or for consistency of expression, they may still be called transformation coefficients.

[0038] In this document, quantized transformation coefficients and transformation coefficients can be referred to as transformation coefficients and scaled transformation coefficients, respectively. In this case, residual information can include information about one or more transformation coefficients, and information about one or more transformation coefficients can be signaled via residual coding syntax. Transformation coefficients can be derived based on residual information (or information about one or more transformation coefficients), and scaled transformation coefficients can be derived via inverse transformation (scaling) of the transformation coefficients. Residual samples can be derived based on inverse transformation (transformation) of the scaled transformation coefficients. This can be applied / expressed similarly in other parts of this document.

[0039] In this document, technical features described individually within a single drawing may be embodied individually or simultaneously.

[0040] Preferred embodiments of this document will be described in more detail below with reference to the attached drawings. Hereafter, the same reference numerals will be used for the same components in the drawings, and redundant descriptions of the same components may be omitted.

[0041] Figure 1 schematically shows an example of a video / image coding system that can be applied to the embodiments described in this document.

[0042] Referring to Figure 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or network.

[0043] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoder, and a renderer. The encoding device may be called a video / image encoding device, and the decoder may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may consist of a separate device or external component.

[0044] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, or video / image archives containing previously captured video / images. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images can be generated via a computer, in which case the video / image capture process may be replaced by the process of generating the associated data.

[0045] An encoding device can encode input video / image data. For compression and coding efficiency, the encoding device can perform a series of steps, including prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in bitstream format.

[0046] The transmitting unit can transmit encoded video / image information or data output in bitstream format to the receiving unit of a receiving device via a digital storage medium or network in file or streaming format. The digital storage medium can include a variety of storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitting unit may include elements for generating media files via a predetermined file format and elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0047] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of the encoding device.

[0048] A renderer can render decoded video / images. The rendered video / images can then be displayed via the display unit.

[0049] Figure 2 is a schematic diagram illustrating the configuration of a video / image encoding device to which the embodiments of this document can be applied. Hereinafter, the term "encoding device" may include an image encoding device and / or a video encoding device.

[0050] Referring to Figure 2, the encoding device 200 may be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be called a rebuilder or a reconstructed block generator. The aforementioned video splitting unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 can be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. Furthermore, the memory 270 may include a DPB (Decoded Picture Buffer) and may be configured by a digital storage medium. The hardware components may also further include the memory 270 as an internal / external component.

[0051] The video splitting unit 210 can split the input video (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, coding units can be recursively split from a coding tree unit (CTU) or a largeest coding unit (LCU) using a QTBTTT (Quad-Tree Binary-Tree Ternary-Tree) structure. For example, a single coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary-tree structure and / or the ternary-tree structure. Alternatively, the binary-tree structure may be applied first. The coding procedure according to this document can be executed based on the final coding unit that cannot be further split. In this case, based on coding efficiency due to video characteristics, the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the above processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can each be separated or partitioned from the final coding unit described above.The above prediction unit is the unit of sample prediction, and the above conversion unit is the unit for deriving the conversion coefficients and / or the unit for deriving the residual signal from the conversion coefficients.

[0052] The term "unit" can sometimes be used interchangeably with terms such as "block" or "area." Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, or it can represent only the pixel / pixel value of the lumen component, or only the pixel / pixel value of the chroma component. A sample can be used as the term corresponding to a single picture (or image) as a pixel or pel.

[0053] The encoding device 200 can generate a residual signal (residual block, residual sample array) by subtracting the predicted signal (predicted block, predicted sample array) output from the inter-prediction unit 221 or intra-prediction unit 222 from the input video signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the predicted signal (predicted block, predicted sample array) from the input video signal (original block, original sample array) within the encoder 200 can be called the subtraction unit 231. The prediction unit can perform predictions on the block to be processed (hereinafter referred to as the current block) and generate a predicted block that includes predicted samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied on a current block or CU basis. The prediction unit can generate various information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each prediction mode. Prediction information can be encoded by the entropy encoding unit 240 and output in bitstream format.

[0054] The intra-prediction unit 222 can predict the current block by referring to a sample in the current picture. The referenced sample may be located adjacent to the current block or at a distance, depending on the prediction mode. The prediction mode in intra-prediction can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Depending on the degree of fineness of prediction direction, the directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 222 can also determine the prediction mode to be applied to the current block by utilizing the prediction modes applied to adjacent blocks.

[0055] The interprediction unit 221 can derive predicted blocks relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, adjacent blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the above reference block and the reference picture containing the above temporal neighboring block may be the same or different. The above temporal neighboring block may be called a collocated reference block or collocated CU (colCU), and the reference picture containing the above temporal neighboring block may be called a collocated picture (colPic). For example, the interpretation unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted.In Motion Vector Prediction (MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0056] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also be based on an intra-block copy (IBC) prediction mode or a palette mode for predictions on blocks. The above IBC prediction mode or palette mode can be used for content video / moving video coding such as in games, for example, as in SCC (Screen Content Coding). IBC basically performs predictions within the current picture, but can be performed in a manner similar to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, sample values ​​within the picture can be signaled based on information about the palette table and palette index.

[0057] The prediction signal generated via the above prediction unit (including the inter-prediction unit 221 and / or the intra-prediction unit 222) can be used to generate a reconstructed signal or a residual signal. The transformation unit 232 can generate transformation coefficients by applying a transformation technique to the residual signal. For example, the transformation technique may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented by this graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels and based on that. The transformation process can also be applied to pixel blocks of the same size and square shape, or to blocks of various (variable) sizes that are not square.

[0058] The quantization unit 233 quantizes the conversion coefficients and transmits them to the entropy encoding unit 240, which can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients can be called residual information. The quantization unit 233 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients. The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding). In addition to the quantized conversion coefficients, the entropy encoding unit 240 can also encode information necessary for video / image restoration (e.g., the values ​​of syntax elements) together with or separately. Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form in units of Network Abstraction Layer (NAL) units. The video / image information may further include information about various parameter sets, such as the Adaptation Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). The video / image information may also further include general constraint information. In this document, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded via the encoding procedure described above and included in the bitstream.The bitstream described above can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include a variety of storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) that transmits the signal output from the entropy encoding unit 240 and / or a storage unit (not shown) that stores it can be configured as an internal / external element of the encoding device 200, or the transmitting unit may be included in the entropy encoding unit 240.

[0059] The quantized conversion coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 234 and the inverse transformation unit 235. The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222. If there are no residuals for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 250 can be called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, and can also be used for inter-prediction of the next picture after filtering, as described later.

[0060] On the other hand, LMCS (Luma Mapping with Chroma Scaling) can also be applied during the picture encoding and / or restoration process.

[0061] The filtering unit 260 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 260 can generate various filtering information and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each filtering method. The filtering information can be encoded by the entropy encoding unit 240 and output in bitstream form.

[0062] The corrected restored picture sent to memory 270 can be used as a reference picture in the interpretation unit 221. When interpretation is applied via this, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve encoding efficiency.

[0063] The DPB in memory 270 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 221. Memory 270 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in already restored pictures. The stored motion information can be transmitted to the inter-prediction unit 221 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 270 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 222.

[0064] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which the embodiments of this document can be applied. Hereinafter, the term "decoding device" may include an image decoding device and / or a video decoding device.

[0065] Referring to Figure 3, the decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-predictor 331 and an intra-predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. The aforementioned entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering device 350 can be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. The memory 360 may include a DPB (Decoded Picture Buffer) and may be configured by a digital storage medium. The above hardware components may also include Memory 360 as an internal / external component.

[0066] When a bitstream containing video / image information is input, the decoding device 300 can reconstruct the image in accordance with the process by which the video / image information was processed in the encoding device shown in Figure 2. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Thus, the decoding processing unit is, for example, a coding unit, which can be divided from a coding tree unit or a maximum coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transformation units can be derived from the coding unit. The reconstructed video signal decoded and output via the decoding device 300 can then be reproduced via a playback device.

[0067] The decoding device 300 can receive the signal output from the encoding device shown in Figure 2 in bitstream form, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can purge the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can decode the picture based on the parameter set information and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of the syntax elements necessary for image reconstruction and the quantized values ​​of the conversion coefficients related to the residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information (and its neighbors) to be decoded, the decoding information of the block to be decoded, or the symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element.In this case, the CABAC entropy decoding method can update the context model after determining the context model by utilizing the information of the decoded symbol / bin for the context model of the next symbol / bin. Of the information decoded by the entropy decoding unit 310, information related to prediction is provided to the prediction unit (inter-prediction unit 332 and intra-prediction unit 331), and the residual values ​​from which entropy decoding has been performed in the entropy decoding unit 310, i.e., quantized conversion coefficients and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, of the information decoded by the entropy decoding unit 310, information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. On the other hand, the decoding device described in this document may be called a video / image / picture decoding device, and the decoding device may be divided (classified) into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, inverse transform unit 322, adder 340, filtering unit 350, memory 360, inter-prediction unit 332, and intra-prediction unit 331.

[0068] The inverse quantization unit 321 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 321 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients.

[0069] The inverse transform unit 322 performs an inverse transform on the transformation coefficients to obtain the residual signal (residual block, residual sample array).

[0070] The prediction unit can perform a prediction on the current block and generate a predicted block containing the predicted sample for the current block. Based on the prediction information output from the entropy decoding unit 310, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block, and can determine a specific intra / inter-prediction mode.

[0071] The prediction unit 320 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for prediction of a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also be based on an intra-block copy (IBC) prediction mode or a palette mode for prediction of a block. The above IBC prediction mode or palette mode can be used for content video / moving image coding such as games, for example, as in SCC (Screen Content Coding). IBC basically performs prediction within the current picture, but can be performed in a manner similar to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, information about the palette table and palette index can be included in the above video / moving image information and signaled.

[0072] The intra-prediction unit 331 can predict the current block by referring to a sample within the current picture. The referenced sample may be located adjacent to the current block or at a distance, depending on the prediction mode. The prediction mode in intra-prediction can include multiple non-directional modes and multiple directional modes. The intra-prediction unit 331 can also determine the prediction mode to be applied to the current block by utilizing the prediction modes applied to adjacent blocks.

[0073] The interprediction unit 332 can derive predicted blocks for the current block based on a reference block (reference sample array) identified by motion vectors on a reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, adjacent blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the interprediction unit 332 can construct a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes, and the prediction information may include information indicating the mode of interprediction for the current block.

[0074] The summing unit 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-prediction unit 332 and / or intra-prediction unit 331). If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block.

[0075] The summing unit 340 may be called the restoration unit or restoration block generation unit. The generated restoration signal can be used for intra-prediction of the next block to be processed in the current picture, and can be output after filtering as described later, or it can be used for intra-prediction of the next picture.

[0076] On the other hand, LMCS (Luma Mapping with Chroma Scaling) can also be applied during the picture decoding process.

[0077] The filtering unit 350 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.

[0078] The restored picture stored (modified) in the DPB of memory 360 can be used as a reference picture by the inter-prediction unit 332. Memory 360 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in already restored pictures. The stored motion information can be transmitted to the inter-prediction unit 260 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 360 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 331.

[0079] In this document, the embodiments described for the filtering unit 260, the inter-prediction unit 221, and the intra-prediction unit 222 of the encoding device 200 can be applied identically or in a corresponding manner to the filtering unit 350, the inter-prediction unit 332, and the intra-prediction unit 331 of the decoding device 300, respectively.

[0080] As mentioned above, prediction is performed to improve compression efficiency when performing video coding. Through this, a predicted block containing predicted samples for the current block, which is the block to be coded, can be generated. Here, the predicted block contains predicted samples in the spatial domain (or pixel domain). The predicted block is derived in both the encoding and decoding devices, and the encoding device can improve video coding efficiency by signaling the decoding device information about the residuals between the original block and the predicted block (residual information), which is not the original sample value of the original block itself. The decoding device can derive a residual block containing residual samples based on the residual information, and can combine the residual block and the predicted block to generate a restored block containing restored samples, and can generate a restored picture containing the restored block.

[0081] The residual information described above can be generated through transformation and quantization procedures. For example, an encoding device can signal the relevant residual information to a decoding device (via a bitstream) by deriving a residual block between the original block and the predicted block, performing a transformation procedure on the residual samples (residual sample array) contained in the residual block to derive transformation coefficients, and then performing a quantization procedure on the transformation coefficients to derive quantized transformation coefficients. Here, the residual information may include information such as the value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transformation coefficients. The decoding device can derive residual samples (or residual blocks) by performing an inverse quantization / inverse transformation process based on the residual information. The decoding device can generate a reconstructed picture based on the predicted block and the residual block. Alternatively, the encoding device can derive a residual block by inverse quantization / inverse transformation of the quantized transformation coefficients for reference in subsequent picture interpredictions, and generate a reconstructed picture based on this.

[0082] Figure 4 shows an example of a schematic video / image encoding method to which the embodiments described in this document can be applied.

[0083] The method disclosed in Figure 4 can be performed by the encoding device 200 shown in Figure 2. Specifically, S400 can be performed by the inter-prediction unit 221 or intra-prediction unit 222 of the encoding device 200, and S410, S420, S430, and S440 can be performed by the subtraction unit 231, conversion unit 232, quantization unit 233, and entropy coding unit 240 of the encoding device 200, respectively.

[0084] Referring to Figure 4, the encoding device can derive predicted samples through predictions for the current block (S400). The encoding device can decide whether to perform inter-prediction or intra-prediction on the current block, and can determine a specific inter-prediction mode or a specific intra-prediction mode based on the RD cost. Depending on the determined mode, the encoding device can derive predicted samples for the current block.

[0085] The encoding device can derive a residual sample by comparing the original sample and the predicted sample for the current block (S410).

[0086] The encoding device can derive conversion coefficients through a conversion procedure for residual samples (S420), quantize the derived conversion coefficients, and derive quantized conversion coefficients (S430).

[0087] The encoding device can encode video information including prediction information and residual information, and output the encoded video information in bitstream format (S440). The prediction information is information related to the prediction procedure and may include information about the prediction mode and motion information (for example, if interpretation is applied). The residual information may include information about the quantized transformation coefficients. The residual information may be entropic coded.

[0088] The output bitstream can be transmitted to a decoding device via a storage medium or network.

[0089] Figure 5 shows an example of a schematic video / image decoding method to which the embodiments described in this document can be applied.

[0090] The method disclosed in Figure 5 can be performed by the decoding device 300 shown in Figure 3. Specifically, S500 can be performed by the inter-prediction unit 332 or the intra-prediction unit 331 of the decoding device 300. In S500, the procedure of decoding the prediction information contained in the bitstream and deriving the values ​​of the related syntax elements can be performed by the entropy decoding unit 310 of the decoding device 300. S510, S520, S530, and S540 can be performed by the entropy decoding unit 310, the inverse quantization unit 321, the inverse transform unit 322, and the adder unit 340 of the decoding device 300, respectively.

[0091] Referring to Figure 5, the decoding device can perform operations corresponding to those performed by the encoding device. Based on the received prediction information, the decoding device can perform inter-prediction or intra-prediction for the current block to derive prediction samples (S500).

[0092] The decoding device can derive quantized conversion coefficients for the current block based on the received residual information (S510). The decoding device can derive quantized conversion coefficients from the residual information via entropy decoding.

[0093] The decoding device can derive the conversion coefficients by inverse quantization of the quantized conversion coefficients (S520).

[0094] The decoding device derives the residual sample through an inverse transformation process with respect to the transformation coefficients (S530).

[0095] The decoding device can generate a reconstructed sample for the current block based on the predicted sample and residual sample, and generate a reconstructed picture based on this (S540). As previously mentioned, an in-loop filtering procedure can be further applied to the reconstructed picture thereafter.

[0096] On the other hand, as mentioned above, when performing predictions on the current block, intra-prediction or inter-prediction can be applied. The following describes the case where inter-prediction is applied to the current block.

[0097] The prediction unit (more specifically, the inter-prediction unit) of the encoding / decoding device can perform inter-prediction on a block-by-block basis to derive predicted samples. Inter-prediction can represent predictions derived in a manner that depends on data elements of pictures other than the current picture (e.g., sample values ​​or motion information). When inter-prediction is applied to the current block, a predicted block (predicted sample array) for the current block can be derived based on the reference block (reference sample array) identified by the motion vector on the reference picture pointed to by the reference picture index. In this case, in order to reduce the amount of motion information transmitted in inter-prediction mode, the motion information of the current block can be predicted on a block, subblock, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information can include motion vectors and reference picture indices. Motion information can further include inter-prediction type information (L0 prediction, L1 prediction, Bi prediction, etc.). When inter-prediction is applied, neighboring blocks can include spatial neighboring blocks that exist within the current picture and temporal neighboring blocks that exist in the reference picture. The reference picture containing the above-mentioned reference block and the reference picture containing the above-mentioned time-adjacent block may be the same or different. The above-mentioned time-adjacent block may be called a collocated reference block, collocated CU (colCU), etc., and the reference picture containing the above-mentioned time-adjacent block may be called a collocated picture (colPic). For example, a list of motion information candidates may be constructed based on the adjacent blocks of the current block, and flags or index information may be signaled to indicate which candidate is selected (used) in order to derive the motion vector and / or reference picture index of the current block.Interpretation can be performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected adjacent block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In Motion Vector Prediction (MVP) mode, the motion vector of the selected adjacent block can be used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0098] The above motion information may include L0 motion information and / or L1 motion information depending on the interpretation type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. A prediction based on an L0 motion vector may be called an L0 prediction, a prediction based on an L1 motion vector may be called an L1 prediction, and a prediction based on both an L0 motion vector and an L1 motion vector may be called a bi prediction. Here, an L0 motion vector may represent a motion vector associated with a reference picture list L0 (L0), and an L1 motion vector may represent a motion vector associated with a reference picture list L1 (L1). A reference picture list L0 may include prior pictures in the output order of the current picture as reference pictures, and a reference picture list L1 may include pictures in the output order of the current picture. Previous pictures may be called forward (reference) pictures, and subsequent pictures may be called backward (reference) pictures. The reference picture list L0 may further include pictures that are later in the output order than the current picture as reference pictures. In this case, the previous picture may be indexed first in the reference picture list L0, and the subsequent picture may be indexed after that. The reference picture list L1 may further include pictures that are earlier in the output order than the current picture as reference pictures. In this case, the subsequent picture may be indexed first in the reference picture list L1, and the previous picture may be indexed after that. Here, the output order may correspond to the POC (Picture Order Count) order.

[0099] Furthermore, various interpretation modes can be used to predict the current block within a picture. For example, various modes such as merge mode, skip mode, MVP (Motion Vector Prediction) mode, affine mode, subblock merge mode, MMVD (Merge with MVD) mode, and HMVP (Historical Motion Vector Prediction) mode can be used. Additional modes such as DMVR (Decoder side Motion Vector Refinement) mode, AMVR (Adaptive Motion Vector Resolution) mode, Bi-prediction with CU-level Weight (BCW), and Bi-Directional Optical Flow (BDOF) can be used. The affine mode may also be called the affine motion prediction mode. The MVP mode may also be called the AMVP (Advanced Motion Vector Prediction) mode. In this document, motion information candidates derived from some modes and / or some modes may also be included as motion information related candidates from other modes. For example, an HMVP candidate may be added as a merge candidate in merge / skip mode, or as an mvp candidate in MVP mode. When an HMVP candidate is used as a motion information candidate in merge mode or skip mode, it may be called an HMVP merge candidate.

[0100] Prediction mode information indicating the inter-prediction mode of the current block can be signaled from the encoding device to the decoding device. In this case, the prediction mode information can be received by the decoding device as part of the bitstream. The prediction mode information may include index information indicating one of many candidate modes. Alternatively, the inter-prediction mode can be indicated through hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether or not the skip mode can be applied, and if the skip mode cannot be applied, a merge flag may be signaled to indicate whether or not the merge mode can be applied, and if the merge mode cannot be applied, an additional flag may be signaled to indicate that the MVP mode is applied, or an additional flag for distinction may be added. Affine modes may be signaled as independent modes, or as modes dependent on the merge mode or MVP mode, etc. For example, affine modes may include affine merge mode and affine MVP mode.

[0101] Furthermore, when applying interpretation to the current block, motion information of the current block can be used. The encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding device can use the original block in the original picture for the current block to search for highly correlated similar reference blocks in fractional pixel units within a predetermined search range in the reference picture, thereby deriving motion information. Block similarity can be derived based on the difference in sample values ​​based on phase. For example, block similarity can be calculated based on the SAD (Sum of Absolute Differences) between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, motion information can be derived based on the reference block with the smallest SAD in the search space. The derived motion information can be signaled to the decoding device according to various methods based on the interpretation mode.

[0102] As described above, a predicted block can be derived for the current block based on the motion information derived by the interprediction mode. The predicted block may include predicted samples (predicted sample arrays) of the current block. If the motion vector (MV) of the current block indicates fractional sample units, an interpolation procedure can be performed to derive predicted samples of the current block based on reference samples of fractional sample units in the reference picture. When affine interprediction is applied to the current block, predicted samples can be generated based on sample / subblock unit MVs. When biprediction is applied, predicted samples derived through a (phase-weighted) sum or weighted average of predicted samples derived based on L0 prediction (i.e., prediction using reference pictures in reference picture list L0 and MVL0) and predicted samples derived based on L1 prediction (i.e., prediction using reference pictures in reference picture list L1 and MVL1) can be used as predicted samples of the current block. When biprediction is applied, if the reference picture used for L0 prediction and the reference picture used for L1 prediction are located in different time directions relative to the current picture (i.e., it is both biprediction and bidirectional prediction), then it can be called true biprediction.

[0103] As mentioned above, based on the predicted samples derived as described above, reconstructed samples and reconstructed pictures can be generated, and then procedures such as in-loop filtering can be performed.

[0104] Figure 6 shows an example of a schematic interpretation-based video / image encoding method to which the embodiments described in this document can be applied.

[0105] The method disclosed in Figure 6 can be performed by the encoding device 200 shown in Figure 2. Specifically, S600 can be performed by the interpretation unit 221 of the encoding device 200, S610 can be performed by the subtraction unit 231 of the encoding device 200, and S620 can be performed by the entropy encoding unit 240 of the encoding device 200.

[0106] Referring to Figure 6, the encoding device can perform interpretation for the current block (S600). The encoding device can derive the interpretation mode and motion information of the current block and generate prediction samples for the current block. Here, the processes of determining the interpretation mode, deriving motion information, and generating prediction samples may be performed simultaneously, or one process may be performed before the others. For example, the interpretation unit of the encoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit can determine the prediction mode for the current block, the motion information derivation unit can derive motion information for the current block, and the prediction sample derivation unit can derive prediction samples for the current block. For example, the interpretation unit of the encoding device can search for blocks similar to the current block within a predetermined area (search area) of the reference picture through motion estimation and derive reference blocks whose difference from the current block is minimal or less than a predetermined standard. Based on this, a reference picture index indicating the reference picture in which the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine which mode to apply to the current block from among various prediction modes. The encoding device can determine the optimal prediction mode for the current block by comparing the RD costs (cost) of various prediction modes.

[0107] For example, when skip mode or merge mode is applied to the current block, the encoding device can configure a merge candidate list and derive a reference block from among the reference blocks indicated by the merge candidates included in the merge candidate list that has the smallest difference from the current block or is below a predetermined standard. In this case, a merge candidate associated with the derived reference block is selected, and merge index information indicating the selected merge candidate is generated and signaled to the decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.

[0108] As another example, when (A)MVP mode is applied to the current block, the encoding device can construct an (A)MVP candidate list and use the motion vector of the selected mvp candidate from among the mvp (motion vector predictor) candidates included in the (A)MVP candidate list as the mvp of the current block. In this case, for example, the motion vector indicating the reference block derived by the motion estimation described above may be used as the motion vector of the current block, and the mvp candidate having the motion vector with the smallest difference from the motion vector of the current block may become the selected mvp candidate. The MVD (Motion Vector Difference), which is the difference obtained by subtracting the mvp from the motion vector of the current block, can be derived. In this case, information regarding the MVD can be signaled to the decoding device. Also, when (A)MVP mode is applied, the value of the reference picture index can be composed of reference picture index information and separately signaled to the decoding device.

[0109] The encoding device can derive residual samples based on predicted samples (S610). The encoding device can derive residual samples by comparing the original samples of the current block with the predicted samples.

[0110] The encoding device can encode video information including prediction information and residual information (S620). The encoding device can output the encoded video information in bitstream format. Here, the prediction information may include information about the prediction procedure, such as prediction mode information (e.g., skip flag, merge flag, or mode index) and information about motion information. The information about motion information may include candidate selection information (e.g., merge index, mvp flag, or mvp index), which is information for deriving the motion vector. The information about motion information may also include the aforementioned MVD information and / or reference picture index information. Furthermore, the information about motion information may include information indicating whether L0 prediction, L1 prediction, or bi prediction is applied. The residual information is information about residual samples. The residual information may include information about quantized transformation coefficients for residual samples.

[0111] The output bitstream can be stored in a (digital) storage medium and transmitted to a decoding device, or it can be transmitted to a decoding device via a network.

[0112] On the other hand, as mentioned above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference sample and residual sample. This is because the encoding device can derive the same prediction results as the decoding device, thereby improving coding efficiency. Therefore, the encoding device can store the reconstructed picture (or reconstructed sample, reconstructed block) in memory and use it as a reference picture for interpretation. As mentioned above, further procedures such as in-loop filtering can be applied to the reconstructed picture.

[0113] Figure 7 shows an example of a schematic interpretation-based video / image decoding method to which the embodiments described in this document can be applied.

[0114] The method disclosed in Figure 7 can be performed by the decoding device 300 shown in Figure 3. Specifically, S700 can be performed by the interpretation unit 332 of the decoding device 300. The procedure in S700 to decode the prediction information contained in the bitstream and derive the values ​​of the associated syntax elements can be performed by the entropy decoding unit 310 of the decoding device 300. S710 and S720 can be performed by the interpretation unit 332 of the decoding device 300, S730 can be performed by the residual processing unit 320 of the decoding device 300, and S740 can be performed by the addition unit 340 of the decoding device 300.

[0115] Referring to Figure 7, the decoding device can perform operations corresponding to those performed by the encoding device. The decoding device can perform predictions on the current block based on the received prediction information and derive prediction samples.

[0116] Specifically, the decoding device can determine the prediction mode for the current block based on the received prediction information (S700). Based on the prediction mode information in the prediction information, the decoding device can determine which inter-prediction mode is applied to the current block.

[0117] For example, based on the merge flag, it can be determined whether a merge mode is applied to the current block, or whether (A)MVP mode is determined. Alternatively, based on the mode index, one of several inter-prediction mode candidates can be selected. Inter-prediction mode candidates may include skip mode, merge mode and / or (A)MVP mode, or may include a variety of inter-prediction modes.

[0118] The decoding device can derive motion information for the current block based on a predetermined interpretation mode (S710). For example, if a skip mode or merge mode is applied to the current block, the decoding device can construct a merge candidate list and select one of the merge candidates included in the merge candidate list. The above selection can be performed based on the selection information (merge index) mentioned above. The motion information for the current block can be derived using the motion information for the selected merge candidate. The motion information for the selected merge candidate can be used as the motion information for the current block.

[0119] As another example, when the (A)MVP mode is applied to the current block, the decoding device can construct an (A)MVP candidate list and use the motion vector of the selected mvp candidate from among the mvp (motion vector predictor) candidates included in the (A)MVP candidate list as the mvp of the current block. The above selection can be performed based on the aforementioned selection information (mvp flag or mvp index). In this case, the MVD of the current block can be derived based on the MVD information, and the motion vector of the current block can be derived based on the current block's mvp and MVD. In addition, the reference picture index of the current block can be derived based on the reference picture index information. The picture indicated by the reference picture index in the reference picture list for the current block can be derived as the reference picture referenced for inter prediction of the current block.

[0120] On the other hand, the movement information of the current block can be derived without constructing a candidate list. In this case, the movement information of the current block can be derived according to the procedure disclosed in the prediction mode described later. In this case, the construction of the candidate list as described above can be omitted.

[0121] The decoding device can generate predicted samples for the current block based on the motion information of the current block (S720). In this case, the reference picture can be derived based on the reference picture index of the current block, and the predicted samples for the current block can be derived using the reference block sample indicated on the reference picture by the motion vector of the current block. In this case, as described later, a prediction sample filtering procedure may be further performed on all or part of the predicted samples for the current block.

[0122] For example, the interpretation unit of a decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit determines the prediction mode for the current block based on the prediction mode information received, the motion information derivation unit derives motion information for the current block (such as motion vectors and / or reference picture indices) based on the motion information received, and the prediction sample derivation unit derives prediction samples for the current block.

[0123] The decoding device can generate residual samples for the current block based on the received residual information (S730). The decoding device can generate reconstructed samples for the current block based on the predicted samples and residual samples, and generate a reconstructed picture based on these (S740). As mentioned above, further procedures such as in-loop filtering can then be applied to the reconstructed picture.

[0124] Figure 8 illustrates an interpretation procedure. The interpretation procedure disclosed in Figure 8 can be applied to the interpretation processes disclosed in Figures 6 and 7 above (when the interpretation mode is applied).

[0125] Referring to Figure 8, as described above, the interpretation procedure may include an interpretation mode determination step, a motion information derivation step based on the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information. The interpretation procedure may be performed in an encoding device and a decoding device, as described above. In this document, the coding device may include an encoding device and / or a decoding device.

[0126] The coding device can determine the interpretation mode for the current block (S800). A variety of interpretation modes can be used to predict the current block in the picture. For example, various modes can be used, such as merge mode, skip mode, MVP (Motion Vector Prediction) mode, affine mode, subblock merge mode, and MMVD (Merge with MVD) mode. Additional modes such as DMVR (Decoder side Motion Vector Refinement) mode, AMVR (Adaptive Motion Vector Resolution) mode, Bi-prediction with CU-level Weight (BCW), and Bi-Directional Optical Flow (BDOF) can be used as supplementary modes or alternatives. The affine mode may also be called the affine motion prediction mode. The MVP mode may also be called the AMVP (Advanced Motion Vector Prediction) mode. In this document, motion information candidates derived by some modes and / or some modes may also be included as one of the motion information related candidates for other modes. For example, an HMVP candidate may be added as a merge candidate in merge / skip mode, or as an MVP candidate in MVP mode. When an HMVP candidate is used as a motion information candidate in merge mode or skip mode, it may be called an HMVP merge candidate.

[0127] Prediction mode information indicating the inter-prediction mode of the current block can be signaled from the encoding device to the decoding device. The prediction mode information can be included in the bitstream and received by the decoding device. The prediction mode information may include index information indicating one of a number of candidate modes. Alternatively, the inter-prediction mode can be indicated through hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether the skip mode can be applied, and if the skip mode is not applied, a merge flag may be signaled to indicate whether the merge mode can be applied, and if the merge mode is not applied, it may indicate that the MVP mode has been applied, or additional flags may be signaled for further differentiation. Affine modes may be signaled as independent modes, or as modes dependent on the merge mode or MVP mode, etc. For example, affine modes may include affine merge mode and affine MVP mode.

[0128] The coding device can derive motion information for the current block (S810). The motion information can be derived based on the interprediction mode.

[0129] The coding device can perform interpretation using motion information of the current block. The encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding device can use the original block in the original picture for the current block to search for highly correlated similar reference blocks in fractional pixel units within a predetermined search range in the reference picture, thereby deriving motion information. Block similarity can be derived based on the difference in sample values ​​based on phase. For example, block similarity can be calculated based on the SAD between the current block (or template of the current block) and the reference block (or template of the reference block). In this case, motion information can be derived based on the reference block with the smallest SAD in the search space. The derived motion information can be signaled to the decoding device in various ways based on the interpretation mode.

[0130] The coding device can perform interpretation based on motion information for the current block (S820). The coding device can derive predicted samples for the current block based on motion information. The current block containing the predicted samples may be called a predicted block.

[0131] On the other hand, when deriving motion information for the current block, motion information candidates can be derived based on spatially adjacent blocks and temporally adjacent blocks, and motion information candidates for the current block can be selected based on the derived motion information candidates. In this case, the selected motion information candidate can be used as the motion information for the current block.

[0132] Figure 9 illustrates the spatially adjacent and temporally adjacent blocks of the current block.

[0133] Referring to Figure 9, a spatially adjacent block refers to an adjacent block located around the current block 900, which is the target of the interpretation prediction, and may include adjacent blocks located to the left of the current block 900, or adjacent blocks located above the current block 900. For example, a spatially adjacent block may include the adjacent block at the lower left corner of the current block 900, the adjacent block on the left side, the adjacent block at the upper right corner, the adjacent block above, and the adjacent block at the upper left corner. In Figure 9, a spatially adjacent block is indicated by "S".

[0134] In one embodiment, the encoding / decoding device can search for spatially adjacent blocks of the current block (for example, the left-lower corner adjacent block, the left-side adjacent block, the right-upper corner adjacent block, the upper adjacent block, and the left-upper corner adjacent block) in a specified order to detect available adjacent blocks, and derive motion information of the detected adjacent blocks as candidate spatial motion information.

[0135] A time-adjacent block is a block located on a different picture (i.e., a reference picture) from the current picture containing the current block 900, and refers to a block (collocated block; col block) that is in the same position as the current block 900 within the reference picture. Here, the reference picture may be earlier or later than the current picture in the Picture Order Count (POC). The reference picture used when deriving a time-adjacent block may be referred to as the same-position reference picture or col picture (collocated picture; col picture). Furthermore, a collocated block can refer to a block located within the col picture corresponding to the position of the current block 900, and may be referred to as a col block. For example, a time-adjacent block may include a col block located in the reference picture (i.e., col picture) corresponding to the lower-right corner sample position of the current block 900 (i.e., a col block containing the lower-right corner sample) and / or a col block located in the reference picture (i.e., col picture) corresponding to the center-lower-right sample position of the current block 900 (i.e., a col block containing the center-lower-right sample). In Figure 9, the time-adjacent block is indicated by "T".

[0136] In one embodiment, the encoding / decoding device can search for time-adjacent blocks of the current block (for example, a block containing a lower-right corner sample, a block containing a lower-right center sample) in a specified order to detect usable blocks, and derive motion information of the detected blocks as candidates for time-motion information. This technique of utilizing time-adjacent blocks may be called TMVP (Temporal Motion Vector Prediction). The candidates for time-motion information may also be called TMVP candidates.

[0137] On the other hand, depending on the interpretation mode, it is also possible to derive motion information at the subblock level and perform predictions. For example, in affine mode and TMVP mode, motion information can be derived at the subblock level. In particular, the method for deriving candidate time motion information at the subblock level can be called subblock-based TMVP (sbTMVP; subblock-based Temporal Motion Vector Prediction).

[0138] sbTMVP utilizes motion fields within a col picture to improve the motion vector prediction (MVP) and merge mode of coding units within the current picture, and the col picture in sbTMVP may be the same as the col picture used by TMVP. However, while TMVP performs motion prediction at the coding unit (CU) level, sbTMVP can perform motion prediction at the sub-block level or sub-coding unit (sub-CU) level. Furthermore, while TMVP derives temporal motion information from col blocks within a col picture (where the col block corresponds to the lower-right corner sample position of the current block or the lower-right center sample position of the current block), sbTMVP derives temporal motion information after applying a motion shift to the col picture. Here, the motion shift may include a process of obtaining a motion vector from one of the spatially adjacent blocks of the current block and being shifted by the aforementioned motion vector.

[0139] Figure 10 illustrates a spatially adjacent block that can be used to derive a candidate for time motion information (sbTMVP candidate) based on a subblock.

[0140] Referring to Figure 10, a spatially adjacent block may include at least one of the following adjacent blocks of the current block: the lower left corner adjacent block A0, the left adjacent block A1, the upper right corner adjacent block B0, and the upper adjacent block B1. Depending on the circumstances, a spatially adjacent block may further include other adjacent blocks other than those shown in Figure 10, or it may not include any specific adjacent block from those shown in Figure 10. Alternatively, a spatially adjacent block may include only a specific adjacent block, for example, only the left adjacent block A1 of the current block.

[0141] For example, an encoding / decoding device can search for spatially adjacent blocks according to a predetermined search order, detect the motion vector of the first available spatially adjacent block, and determine the block at the position indicated by the motion vector of the spatially adjacent block in the reference picture as the col block (i.e., collated reference block). Here, the motion vector of a spatially adjacent block may be called a temporal motion vector (temporal MV).

[0142] In this case, the availability of a spatially adjacent block can be determined by the reference picture information, prediction mode information, and location information of the spatially adjacent block. For example, if the reference picture of the spatially adjacent block is the same as the reference picture of the current block, the spatially adjacent block can be determined to be available. Alternatively, if the spatially adjacent block is coded in intra-prediction mode, or if the spatially adjacent block is located outside the current picture / tile, the spatially adjacent block can be determined to be unavailable.

[0143] Furthermore, the search order for spatially adjacent blocks can be defined in various ways; for example, it may be in the order A1, B1, B0, A0. Alternatively, only A1 can be searched to determine whether or not A1 is available.

[0144] Figure 11 is a schematic diagram illustrating the process of deriving candidate time-motion information (sbTMVP candidates) based on subblocks.

[0145] Referring to Figure 11, first, the encoding / decoding device can determine whether the spatially adjacent block of the current block (e.g., block A1) is available. For example, if the reference picture of the spatially adjacent block (e.g., block A1) uses a col picture, the spatially adjacent block (e.g., block A1) can be determined to be available, and the motion vector of the spatially adjacent block (e.g., block A1) can be derived. In this case, the motion vector of the spatially adjacent block (e.g., block A1) can be called tempMV, and this motion vector can be used for motion shifting. Alternatively, if the spatially adjacent block (e.g., block A1) is determined to be unavailable, tempMV (i.e., the motion vector of the spatially adjacent block) can be set to a zero vector. In other words, in this case, the motion shifting can be applied to the motion vector set to (0,0).

[0146] Next, the encoding / decoding device can apply a motion shift based on the motion vector of a spatially adjacent block (e.g., block A1). For example, the motion shift can be applied to the position indicated by the motion vector of the spatially adjacent block (e.g., block A1) (e.g., A1'). In other words, by applying a motion shift, the motion vector of the spatially adjacent block (e.g., block A1) can be added to the coordinates of the current block.

[0147] Next, the encoding / decoding device can derive motion-shifted col subblocks (collocated subblocks) on the col picture and obtain motion information (motion vector, reference index, etc.) for each col subblock. For example, the encoding / decoding device can derive each col subblock on the col picture corresponding to the motion-shifted position of each subblock within the current block (i.e., the position indicated by the motion vector of the spatially adjacent block (e.g., A1)). Then, the motion information of each col subblock can be used as motion information for each subblock relative to the current block (i.e., sbTMVP candidates).

[0148] Furthermore, scaling can be applied to the motion vectors of the col subblocks. This scaling can be performed based on the difference in temporal distance between the reference picture of the col block and the reference picture of the current block. Therefore, this scaling may be called time-motion scaling, and through this, the reference picture of the current block and the reference picture of the time-motion vector can be aligned. In this case, the encoding / decoding device can obtain the scaled motion vectors of the col subblocks as motion information for each subblock relative to the current block.

[0149] Furthermore, when deriving sbTMVP candidates, there may be cases where no motion information exists in the col subblock. In this case, base motion information (or default motion information) can be derived for col subblocks that lack motion information, and this base motion information can be used as the motion information for the subblock relative to the current block. Base motion information can be derived from the block located at the center of the col block (i.e., the col CU containing the col subblock). For example, motion information (e.g., a motion vector) can be derived from the block containing the sample located at the lower right end among the four samples located at the center of the col block, and this can be used as the base motion information.

[0150] As mentioned above, in the case of affine mode or sbTMVP mode, which derive movement information on a subblock basis, affine merge candidates and sbTMVP candidates can be derived, and a merge candidate list based on subblocks can be constructed based on these candidates. At this time, flag information indicating whether affine mode or sbTMVP mode is available or unavailable can be signaled. If sbTMVP mode is available based on the above flag information, the sbTMVP candidates derived as described above can be added as the first entry (firstly-ordered) in the merge candidate list based on subblocks. Then, affine merge candidates can be added as the next entry in the merge candidate list based on subblocks. Here, the maximum number of candidates in the merge candidate list based on subblocks may be five.

[0151] Furthermore, in sbTMVP mode, the size of the subblock can be fixed, for example, to an 8x8 size. Also, sbTMVP mode can only be applied to blocks where both the width and height are 8 or greater.

[0152] On the other hand, the current VVC standard allows for the derivation of time-motion information candidates (sbTMVP candidates) based on subblocks, as shown in Table 1 below.

[0153] [Table 1-1]

[0154] [Table 1-2]

[0155] [Table 1-3]

[0156] In deriving sbTMVP candidates according to the method shown in Table 1 above, default MV and subblock MV(s) may be considered. Here, default MV may be called subblock-based temporal merging base motion data or base motion data (base motion information). Referring to Table 1 above, default MV can correspond to ctrMV (or ctrMVLX) in Table 1. Subblock MV can correspond to mvSbCol (or mvLXSbcol) in Table 1.

[0157] For example, if a subblock or subblock MV is available through the derivation process of sbTMVP, the subblock MV is assigned to that subblock; or, if a subblock or subblock MV is unavailable, the default MV is used as the subblock MV for that subblock. Here, the default MV derives motion information from a position corresponding to the center pixel position of the corresponding block (i.e., col CU) on the col picture, and each subblock MV can derive motion information from the top-left position of the corresponding subblock (i.e., col subblock) on the col picture. In this case, the corresponding block (i.e., col CU) can be derived from a position shifted based on the motion vector (i.e., temporal MV) of the spatially adjacent block A1, as shown in Figure 11.

[0158] Figures 12 to 15 schematically illustrate the method for calculating the corresponding position for deriving the default MV and subblock MV based on the block size during the sbTMVP derivation process.

[0159] In Figures 12 to 15, pixels (samples) with dotted lines indicate the corresponding positions within each subblock for deriving each subblock MV, while pixels (samples) with solid lines indicate the corresponding positions within CU for deriving the default MV.

[0160] For example, referring to Figure 12, if the current block (i.e., the current CU) is 8x8 in size, the motion information of the subblock can be derived based on the upper left sample position within the 8x8 subblock, and the default motion information of the subblock can be derived based on the center sample position within the 8x8 current block (i.e., the current CU).

[0161] Alternatively, for example, referring to Figure 13, if the current block (i.e., current CU) is 16x8 in size, the motion information for each subblock can be derived based on the upper left sample position within each 8x8 subblock, and the default motion information for each subblock can be derived based on the center sample position within the 16x8 current block (i.e., current CU).

[0162] Alternatively, for example, referring to Figure 14, if the current block (i.e., current CU) is 8x16 in size, the motion information for each subblock can be derived based on the upper left sample position within each 8x8 subblock, and the default motion information for each subblock can be derived based on the center sample position within the 8x16 current block (i.e., current CU).

[0163] Alternatively, for example, referring to Figure 15, if the current block (i.e., current CU) is 16x16 in size, the motion information for each subblock can be derived based on the upper left sample position within each 8x8 subblock, and the default motion information for each subblock can be derived based on the center sample position within the 16x16 current block (i.e., current CU).

[0164] As can be seen from Figures 12 to 15 above, the motion information of the subblock is biased towards the upper left pixel position, resulting in a problem where the subblock MV is induced at a position far removed from the position from which the default MV, which represents the representative motion information of the current CU, is derived. As a simple example, in the case of an 8x8 block shown in Figure 12, one CU contains one subblock, but there is a contradiction in that the subblock MV and the default MV are represented by different motion information. Furthermore, since the method for calculating the corresponding position between the subblock and the current CU block is different (i.e., the corresponding position for deriving the subblock MV is the upper left sample position, and the corresponding position for deriving the default MV is the center sample position), an additional module may be required when implementing the hardware (H / W).

[0165] Therefore, in order to improve the above-mentioned problems, this document proposes a method for unifying the method of deriving the corresponding position of the CU for the default MV and the method of deriving the corresponding position of the subblock for each subblock MV in the process of deriving the sbTMVP candidate. According to one embodiment of this document, the effect of integration is achieved by using only one module that derives each corresponding position according to the block size from a hardware (H / W) perspective. For example, the method for calculating the corresponding position when the block size is 16x16 blocks and the method for calculating the corresponding position when the block size is 8x8 blocks can be implemented identically, thus achieving a simplification effect in terms of hardware implementation. Here, the 16x16 block can represent the CU, and the 8x8 block can represent each subblock.

[0166] In one embodiment, when deriving sbTMVP candidates, the center sample position can be used as the corresponding position for deriving subblock motion information and the corresponding position for deriving default motion information, and this can be embodied as shown in Table 2 below.

[0167] Table 2 below is a specification showing an example of a method for deriving subblock motion information and default motion information according to one embodiment of this document.

[0168] [Table 2-1]

[0169] [Table 2-2]

[0170] [Table 2-3]

[0171] Referring to Table 2 above, the position of the current block (i.e., the current CU) containing the subblock can be derived when deriving the sbTMVP candidate. The upper left sample position (xCtb, yCtb) of the coding tree block (or coding tree unit) containing the current block and the lower right center sample position (xCtr, yCtr) of the current block can be derived as shown in formulas (8-514) to (8-517) in Table 2 above. In this case, the positions (xCtb, yCtb) and (xCtr, yCtr) can be calculated based on the upper left sample position (xCb, yCb) of the current block, with respect to the upper left sample of the current picture.

[0172] Furthermore, a col block (i.e., col CU) on the col picture that corresponds to the current block (i.e., current CU) including subblocks can be derived. In this case, the position of the col block can be set to (xColCtrCb, yColCtrCb), and this position can indicate the position of the col block including the position (xCtr, yCtr) within the col picture, relative to the upper left edge sample of the col picture.

[0173] Furthermore, base motion data (i.e., default motion information) for sbTMVP can be derived. The base motion data may include a default MV (e.g., ctrMvLX). For example, a col block on a col picture can be derived. In this case, the position of the col block can be derived as (xColCb, yColCb), which may be the position obtained by applying a motion shift (e.g., tempMv) to the derived col block position (xColCtrCb, yColCtrCb). The motion shift can be performed by adding a motion vector (e.g., tempMv) derived from a spatially adjacent block of the current block (e.g., block A1) to the current col block position (xColCtrCb, yColCtrCb), as described above. Next, a default MV (e.g., ctrMvLX) can be derived based on the motion-shifted col block position (xColCb, yColCb). Here, the default MV (e.g., ctrMvLX) can represent the motion vector derived from the position corresponding to the center sample at the lower right edge of the col block.

[0174] Furthermore, the col subblocks on the col picture corresponding to the subblocks within the current block (referred to as current subblocks) can be derived. First, the position of each current subblock can be derived. The position of each subblock can be denoted as (xSb, ySb), and this position (xSb, ySb) indicates the position of the current subblock relative to the upper left edge sample of the current picture. For example, the position of the current subblock (xSb, ySb) can be calculated as shown in formulas (8-523) to (8-524) in Table 2 above, which indicates the position of the lower right edge center sample of the subblock. Next, the position of each col subblock on the col picture can be derived. The position of each col subblock can be denoted as (xColSb, yColSb), and this position (xColSb, yColSb) may be the position obtained by applying a motion shift (e.g., tempMv) to the position of the current subblock (xSb, ySb). Motion shifting can be performed, as described above, by adding a motion vector (e.g., tempMv) derived from the spatially adjacent block of the current block (e.g., block A1) to the current subblock's position (xSb, ySb). Next, based on the position (xColSb, yColSb) of each motion-shifted col subblock, motion information for the col subblock (e.g., motion vector mvLXSbCol, availableFlagLXSbCol, a flag indicating whether it is available or not) can be derived.

[0175] In this case, if there are unusable col subblocks in the col subblock (for example, if availableFlagLXSbCol is 0), base motion data (i.e., default motion information) can be used for the unusable col subblocks. For example, the motion vector (e.g., mvLXSbCol) for an unusable col subblock can use the default MV (e.g., ctrMvLX).

[0176] Figures 16 to 19 are illustrative diagrams illustrating a schematic method for integrating corresponding positions for deriving the default MV and subblock MV based on the block size during the sbTMVP derivation process.

[0177] In Figures 16 to 19, the pixels (samples) with dotted lines indicate the corresponding positions within each subblock for deriving each subblock MV, while the pixels (samples) with solid lines indicate the corresponding positions within the CU for deriving the default MV.

[0178] For example, referring to Figure 16, if the current block (i.e., current CU) is 8x8 in size, the motion information of the current subblock can be derived from the corresponding col subblock on the col picture based on the right-bottom center sample position within the 8x8 subblock. The default motion information of the current subblock can be derived from the corresponding col block (i.e., col CU) on the col picture based on the right-bottom center sample position within the 8x8 current block (i.e., current CU). In this case, as shown in Figure 16, the motion information of the current subblock and the default motion information can be derived from the same sample position (same corresponding position).

[0179] Alternatively, for example, referring to Figure 17, if the current block (i.e., current CU) is 16x8 in size, the motion information of the current subblock can be derived and used from the col subblock at the corresponding position on the col picture based on the right bottom center sample position within the 8x8 subblock. The default motion information of the current subblock can be derived and used from the col block (i.e., col CU) at the corresponding position on the col picture based on the right bottom center sample position within the 16x8 current block (i.e., current CU).

[0180] Alternatively, as shown in Figure 18, for example, if the current block (i.e., current CU) is 8x16 in size, the motion information of the current subblock can be derived from the corresponding position of the col subblock on the col picture based on the right-bottom center sample position within the 8x8 subblock. The default motion information of the current subblock can be derived from the corresponding position of the col block (i.e., col CU) on the col picture based on the right-bottom center sample position within the 8x16 current block (i.e., current CU).

[0181] Alternatively, as shown in Figure 19, for example, if the current block (i.e., current CU) is 16x16 in size or larger, the motion information of the current subblock can be derived from the corresponding position of the col subblock on the col picture based on the right bottom center sample position within the 8x8 subblock. The default motion information of the current subblock can be derived from the corresponding position of the col block (i.e., col CU) on the col picture based on the right bottom center sample position within the 16x16 (or larger) current block (i.e., current CU).

[0182] However, the embodiments described in this document above are merely examples, and the default motion information and the motion information of the current subblock can be derived not only from the center position (i.e., the lower right end sample position) but also from other sample positions. For example, the default motion information can be derived from the upper left end sample position of the current CU, and the motion information of the current subblock can be derived from the upper left end sample position of the subblock.

[0183] When implementing the embodiments described in this document in hardware, as described above, the same H / W module can be used to derive motion information (temporal motion), allowing for the configuration of pipelines as shown in Figures 20 and 21.

[0184] Figures 20 and 21 are schematic diagrams illustrating a pipeline configuration that allows for the integrated calculation of corresponding positions for deriving the default MV and subblock MV during the sbTMVP derivation process.

[0185] Referring to Figures 20 and 21, the corresponding position calculation module can calculate the corresponding positions for deriving the default MV and subblock MV. For example, as shown in Figures 20 and 21, if the block position (posX, posY) and block size (blkszX, blkszY) are input to the corresponding position calculation module, the center position of the input block (i.e., the lower right end sample position) can be output. If the current CU position and block size are input to the corresponding position calculation module, the center position of the col block on the col picture (i.e., the lower right end sample position), which is the corresponding position for deriving the default MV, can be output. Alternatively, if the current subblock position and block size are input to the corresponding position calculation module, the center position of the col subblock on the col picture (i.e., the lower right end sample position), which is the corresponding position for deriving the current subblock MV, can be output.

[0186] Thus, when the corresponding positions for deriving the default MV and subblock MV are output from the corresponding position calculation module, the motion vectors (i.e., temporal mv) derived from the above corresponding positions can be patched. Then, based on the patched motion vectors (i.e., temporal mv), time motion information (i.e., sbTMVP candidates) based on the subblock can be derived. For example, as shown in Figures 20 and 21, by implementing the hardware, sbTMVP candidates can be derived in parallel based on the clock cycle, or sequentially.

[0187] The following drawings were created to illustrate a specific example of this document. The names of specific devices, terms, and names (e.g., syntax / syntax element names) shown in the drawings are presented illustratively; therefore, the technical features of this document are not limited to the specific names used in the following drawings.

[0188] Figure 22 schematically shows an example of a video / image encoding method according to the embodiment described in this document.

[0189] The method disclosed in Figure 22 can be performed by the encoding device 200 disclosed in Figure 2. Specifically, steps (S2200) to (S2230) in Figure 22 can be performed by the prediction unit 220 (more specifically, the interpretation unit 221) disclosed in Figure 2, step (S2240) in Figure 22 can be performed by the residual processing unit 230 disclosed in Figure 2, and step (S2250) in Figure 22 can be performed by the entropy encoding unit 240 disclosed in Figure 2. Furthermore, the method disclosed in Figure 22 can be performed including the embodiments described above in this document. Therefore, in Figure 22, specific explanations of content that overlaps with the embodiments described above will be omitted or simplified.

[0190] Referring to Figure 22, the encoding device can derive the reference subblock on the collocated reference picture for the subblock within the current block (S2200).

[0191] Here, the collated reference picture refers to the reference picture used to derive the time motion information (i.e., sbTMVP) as described above, and can refer to the col picture mentioned earlier. The reference subblock can refer to the col subblock mentioned earlier.

[0192] In one embodiment, the encoding device can derive a reference subblock on a collated reference picture based on the position of a subblock within the current block. Here, the current block may be referred to as the current coding unit (CU) or current coding block (CB), and the subblocks contained within the current block may also be referred to as current coding subblocks.

[0193] For example, an encoding device can first determine the position of the current block, and then determine the position of subblocks within the current block. As explained with reference to Table 2 above, the position of the current block can be determined based on the upper left sample position (xCtb, yCtb) of the coding tree block and the lower right center sample position (xCtr, yCtr) of the current block. The positions of subblocks within the current block can be determined by (xSb, ySb), and these positions (xSb, ySb) indicate the lower right center sample position of the subblock. Here, the lower right center sample position (xSb, ySb) of the subblock can be calculated based on the upper left sample position and the subblock size, and can be calculated as shown in formulas (8-523) to (8-524) in Table 2 above.

[0194] The encoding device can then derive reference subblocks on the collated reference picture based on the right-bottom center sample position of each subblock within the current block. As explained with reference to Table 2 above, reference subblocks can be indicated by their position (xColSb, yColSb) on the collated reference picture, and the position (xColSb, yColSb) can be derived on the collated reference picture based on the right-bottom center sample position (xSb, ySb) of each subblock within the current block.

[0195] On the other hand, the upper left sample position used in this document may also be referred to as the upper left sample position, the upper left sample position, etc., and the lower right center sample position may also be referred to as the lower right center sample position, the center lower right sample position, the lower right center sample position, the center lower right sample position, etc.

[0196] Furthermore, motion shifting can be applied when deriving reference subblocks. The encoding device can perform motion shifting based on motion vectors derived from the spatially adjacent blocks of the current block. The spatially adjacent block of the current block may be the left-side adjacent block located to the left of the current block, for example, block A1 shown in Figures 10 and 11. In this case, if the left-side adjacent block (e.g., block A1) is available, a motion vector can be derived from the left-side adjacent block; or, if the left-side adjacent block is unavailable, a zero vector can be derived. Here, whether or not a spatially adjacent block is available can be determined by the reference picture information, prediction mode information, position information, etc., of the spatially adjacent block. For example, if the reference picture of the spatially adjacent block and the reference picture of the current block are the same, the spatially adjacent block can be determined to be available. Alternatively, if the spatially adjacent block is coded in intra-prediction mode, or if the spatially adjacent block is located outside the current picture / tile, the spatially adjacent block can be determined to be unavailable.

[0197] In other words, the encoding device can apply a motion shift (i.e., the motion vector of the spatially adjacent block (e.g., block A1)) to the right-side lower-end center sample position (xSb, ySb) of each subblock within the current block, and derive the reference subblock on the collated reference picture based on the motion-shifted position. At this time, the position of the reference subblock (xColSb, yColSb) can be represented by the position that has been motion-shifted at the right-side lower-end center sample position (xSb, ySb) of each subblock within the current block to the position indicated by the motion vector of the spatially adjacent block (e.g., block A1), and can be calculated as shown in formulas (8-525) to (8-526) in Table 2 above.

[0198] The encoding device can derive candidate time-motion information based on the subblock based on the reference subblock (S2210).

[0199] On the other hand, in this document, the subblock-based temporal motion information candidate refers to the aforementioned sbTMVp (subblock-based Temporal Motion Vector Prediction) candidate, and can be substituted for or used interchangeably with the subblock-based temporal motion vector predictor candidate. That is, as described above, when motion information is derived and prediction is performed on a subblock basis, an sbTMVP candidate can be derived, and motion prediction can be performed at the subblock level (or subcoding unit (sub-CU) level) based on the above sbTMVP candidate.

[0200] The encoding device can derive motion information for subblocks within the current block based on candidate time motion information based on subblocks (S2220).

[0201] Candidate time-motion information based on a subblock may include subblock unit motion vectors. In this case, subblock unit motion vectors may include motion vectors derived based on a reference subblock.

[0202] In one embodiment, the encoding device can derive a subblock unit motion vector for a reference subblock as motion information for a subblock within the current block. For example, the encoding device can derive a subblock unit motion vector based on whether the reference subblock is available or not. For available reference subblocks in a reference subblock, the encoding device can derive a subblock unit motion vector for the available reference subblock based on the motion vector of the available reference subblock. For unavailable reference subblocks in a reference subblock, the encoding device can use a base motion vector as the subblock unit motion vector for the unavailable reference subblock.

[0203] The base motion vector can correspond to the default motion vector mentioned above and can be derived on the collated reference picture based on the current block's position.

[0204] In one embodiment, the encoding device can determine the position of a reference coding block on a collated reference picture based on the right-side lower center sample position of the current block, and derive a base motion vector based on the position of the reference coding block. The reference coding block may be referred to as a col block located on a collated reference picture corresponding to the current block, including subblocks. As described with reference to Table 2 above, the position of the reference coding block can be denoted as (xColCtrCb, yColCtrCb), where (xColCtrCb, yColCtrCb) indicates the position of the reference coding block covering the position (xCtr, yCtr) within the collated reference picture, relative to the left-side upper sample of the collated reference picture. The position (xCtr, yCtr) can indicate the right-side lower center sample position of the current block.

[0205] Furthermore, in deriving the base motion vector, a motion shift can be applied to the reference coding block positions (xColCtrCb, yColCtrCb). As mentioned above, the motion shift can be performed by adding the motion vector derived from the spatially adjacent block of the current block (e.g., block A1) to the reference coding block positions (xColCtrCb, yColCtrCb) that cover the right-side lower center sample. The encoding device can derive the base motion vector based on the motion-shifted reference coding block positions (xColCb, yColCb). That is, the base motion vector may be a motion vector derived from a position motion-shifted on the collated reference picture based on the right-side lower center sample position of the current block.

[0206] On the other hand, whether the above reference subblock is usable can be determined based on whether it is located outside the collated reference picture or on its motion vector. For example, an unusable reference subblock may include a reference subblock located outside (or deviating from) the collated reference picture or a reference subblock for which the motion vector is unusable. For example, if the reference subblock is based on intra mode, IBC (Intra Block Copy) mode, or palette mode, the above reference subblock may be a subblock for which the motion vector is unusable. Alternatively, if the reference coding block covering a modified location derived based on the location of the reference subblock is based on intra mode, IBC mode, or palette mode, the above reference subblock may be a subblock for which the motion vector is unusable.

[0207] In one embodiment, the motion vector of an available reference subblock can be derived based on the motion vector of a block covering a modified location derived based on the upper-left sample position of the reference subblock. For example, as shown in Table 2 above, the modified location can be derived based on a formula such as ((xColSb>>3)<<3, (yColSb>>3)<<3), where xColSb and yColSb represent the x and y coordinates of the upper-left sample position of the reference subblock, respectively, and >> can represent an arithmetic right shift and << can represent an arithmetic left shift.

[0208] On the other hand, as mentioned above, when deriving candidate time motion information based on subblocks, the motion vector for the reference subblock is derived based on the position of the subblock within the current block, and the base motion vector is derived based on the position of the current block. For example, as explained in Figures 16 to 19, for an 8x8 size current block, the motion vector for the reference subblock and the base motion vector can be derived based on the right-side lower center sample position of the current block. For a current block larger than 8x8 size, the motion vector for the reference subblock is derived based on the right-side lower center sample position of each subblock within the current block, and the base motion vector can be derived based on the right-side lower center sample position of the current block.

[0209] The encoding device can generate predicted samples of the current block based on motion information for subblocks within the current block (S2230).

[0210] The encoding device can select the optimal motion information based on the Rate-Distortion (RD) cost and generate predictive samples based on this. For example, if motion information derived for the current block at the sub-block level (i.e., sbTMVP) is selected as the optimal motion information, the encoding device can generate predictive samples for the current block based on the motion information for the sub-blocks of the current block derived as described above.

[0211] The encoding device can derive residual samples based on predicted samples (S2240) and encode video information including information about the residual samples (S2250).

[0212] In other words, the encoding device can derive residual samples based on the original samples for the current block and the predicted samples for the current block. The encoding device can then generate information about the residual samples. Here, the information about the residual samples may include information such as the values ​​of the quantized transformation coefficients derived by performing transformation and quantization on the residual samples, positional information, transformation technique, transformation kernel, and quantization parameters.

[0213] The encoding device can encode information about the residual sample and output it as a bitstream, which can then be transmitted to the decoding device via a network or storage medium.

[0214] Figure 23 schematically shows an example of a video / image decoding method according to the embodiment described in this document.

[0215] The method disclosed in Figure 23 can be performed by the decoding device 300 disclosed in Figure 3. Specifically, steps (S2300) to (S2330) in Figure 23 can be performed by the prediction unit 330 (more specifically, the interpretation unit 332) disclosed in Figure 3, and step (S2340) in Figure 23 can be performed by the addition unit 340 disclosed in Figure 3. Furthermore, the method disclosed in Figure 23 can be performed including the embodiments described above in this document. Therefore, in Figure 23, specific explanations of content that overlaps with the embodiments described above will be omitted or simplified.

[0216] Referring to Figure 23, the decoding device can derive the reference subblock on the collocated reference picture for the subblock within the current block (S2300).

[0217] Here, the collated reference picture refers to the reference picture used to derive the time motion information (i.e., sbTMVP) as described above, and can refer to the col picture mentioned earlier. The reference subblock can refer to the col subblock mentioned earlier.

[0218] In one embodiment, the decoding device can derive a reference subblock on a collated reference picture based on the position of the subblock within the current block. Here, the current block may be referred to as the current coding unit (CU) or current coding block (CB), and the subblocks contained within the current block may also be referred to as current coding subblocks.

[0219] For example, a decoding device can first determine the position of the current block, and then determine the position of subblocks within the current block. As explained with reference to Table 2 above, the position of the current block can be determined based on the upper left sample position (xCtb, yCtb) of the coding tree block and the lower right center sample position (xCtr, yCtr) of the current block. The positions of subblocks within the current block can be determined by (xSb, ySb), and these positions (xSb, ySb) indicate the lower right center sample position of the subblock. Here, the lower right center sample position (xSb, ySb) of the subblock can be calculated based on the upper left sample position and the subblock size, and can be calculated as shown in formulas (8-523) to (8-524) in Table 2 above.

[0220] The decoding device can then derive the reference subblock on the collated reference picture based on the right-side lower-end center sample position of each subblock within the current block. As explained with reference to Table 2 above, the reference subblock can be indicated by its position (xColSb, yColSb) on the collated reference picture, and the position (xColSb, yColSb) can be derived on the collated reference picture based on the right-side lower-end center sample position (xSb, ySb) of each subblock within the current block.

[0221] On the other hand, the upper left sample position used in this document may also be referred to as the upper left sample position, the upper left sample position, etc., and the lower right center sample position may also be referred to as the lower right center sample position, the center lower right sample position, the lower right center sample position, the center lower right sample position, etc.

[0222] Furthermore, motion shifting can be applied when deriving reference subblocks. The decoding device can perform motion shifting based on motion vectors derived from the spatially adjacent blocks of the current block. The spatially adjacent block of the current block may be the left-side adjacent block located to the left of the current block, for example, block A1 shown in Figures 10 and 11. In this case, if the left-side adjacent block (e.g., block A1) is available, motion vectors can be derived from the left-side adjacent block; or, if the left-side adjacent block is unavailable, a zero vector can be derived. Here, whether or not a spatially adjacent block is available can be determined by the reference picture information, prediction mode information, position information, etc., of the spatially adjacent block. For example, if the reference picture of the spatially adjacent block and the reference picture of the current block are the same, the spatially adjacent block can be determined to be available. Alternatively, if the spatially adjacent block is coded in intra-prediction mode, or if the spatially adjacent block is located outside the current picture / tile, the spatially adjacent block can be determined to be unavailable.

[0223] In other words, the decoding device applies a motion shift (i.e., the motion vector of the spatially adjacent block (e.g., block A1)) to the right-side lower-end center sample position (xSb, ySb) of each subblock within the current block, and derives the reference subblock on the collated reference picture based on the motion-shifted position. At this time, the position of the reference subblock (xColSb, yColSb) can be represented by the position that has been motion-shifted at the right-side lower-end center sample position (xSb, ySb) of each subblock within the current block to the position indicated by the motion vector of the spatially adjacent block (e.g., block A1), and can be calculated as shown in formulas (8-525) to (8-526) in Table 2 above.

[0224] The decoding device can derive candidate time-motion information based on the subblock based on the reference subblock (S2310).

[0225] On the other hand, in this document, the subblock-based temporal motion information candidate refers to the aforementioned sbTMVp (Subblock-Based Temporal Motion Vector Prediction) candidate, and can be substituted for or used interchangeably with the subblock-based temporal motion vector predictor candidate. That is, as described above, when motion information is derived and prediction is performed on a subblock basis, an sbTMVP candidate can be derived, and motion prediction can be performed at the subblock level (or subcoding unit (sub-CU) level) based on the above sbTMVP candidate.

[0226] The decoding device can derive motion information for subblocks within the current block based on candidate time motion information based on subblocks (S2320).

[0227] Candidate time-motion information based on a subblock may include subblock unit motion vectors. In this case, subblock unit motion vectors may include motion vectors derived based on a reference subblock.

[0228] In one embodiment, the decoding device can derive a subblock unit motion vector for a reference subblock as motion information for a subblock within the current block. For example, the decoding device can derive a subblock unit motion vector based on whether the reference subblock is available or not. For available reference subblocks in a reference subblock, the decoding device can derive a subblock unit motion vector for the available reference subblock based on the motion vector of the available reference subblock. For unavailable reference subblocks in a reference subblock, the decoding device can use a base motion vector as the subblock unit motion vector for the unavailable reference subblock.

[0229] The base motion vector can correspond to the default motion vector mentioned above and can be derived on the collated reference picture based on the current block's position.

[0230] In one embodiment, the decoding device can locate the position of a reference coding block on a collated reference picture based on the right-side lower center sample position of the current block, and derive a base motion vector based on the position of the reference coding block. The reference coding block may refer to a col block located on a collated reference picture corresponding to the current block, including subblocks. As described with reference to Table 2 above, the position of the reference coding block can be denoted as (xColCtrCb, yColCtrCb), where (xColCtrCb, yColCtrCb) indicates the position of the reference coding block covering the position (xCtr, yCtr) within the collated reference picture, relative to the left-side upper sample of the collated reference picture. The position (xCtr, yCtr) can indicate the right-side lower center sample position of the current block.

[0231] Furthermore, in deriving the base motion vector, a motion shift can be applied to the reference coding block positions (xColCtrCb, yColCtrCb). As mentioned above, the motion shift can be performed by adding the motion vector derived from the spatially adjacent block of the current block (e.g., block A1) to the reference coding block positions (xColCtrCb, yColCtrCb) that cover the right-side lower center sample. The decoding device can derive the base motion vector based on the motion-shifted reference coding block positions (xColCb, yColCb). That is, the base motion vector may be a motion vector derived from a position motion-shifted on the collated reference picture based on the right-side lower center sample position of the current block.

[0232] On the other hand, whether the above reference subblock is usable can be determined based on whether it is located outside the collated reference picture or on its motion vector. For example, an unusable reference subblock may include a reference subblock that deviates from the outside of the collated reference picture, or a reference subblock for which the motion vector is unusable. For example, if the reference subblock is based on intra mode, IBC (Intra Block Copy) mode, or palette mode, the above reference subblock may be a subblock for which the motion vector is unusable. Alternatively, if the reference coding block covering a modified position derived based on the position of the reference subblock is based on intra mode, IBC mode, or palette mode, the above reference subblock may be a subblock for which the motion vector is unusable.

[0233] In one embodiment, the motion vector of an available reference subblock can be derived based on the motion vector of a block covering a modified location derived based on the upper-left sample position of the reference subblock. For example, as shown in Table 2 above, the modified location can be derived based on a formula such as ((xColSb>>3)<<3, (yColSb>>3)<<3), where xColSb and yColSb represent the x and y coordinates of the upper-left sample position of the reference subblock, respectively, and >> can represent an arithmetic right shift and << can represent an arithmetic left shift.

[0234] On the other hand, as mentioned above, when deriving candidate time motion information based on subblocks, the motion vector for the reference subblock is derived based on the position of the subblock within the current block, and the base motion vector is derived based on the position of the current block. For example, as explained in Figures 16 to 19, for an 8x8 size current block, the motion vector for the reference subblock and the base motion vector can be derived based on the right-side lower center sample position of the current block. For a current block larger than 8x8 size, the motion vector for the reference subblock is derived based on the right-side lower center sample position of each subblock within the current block, and the base motion vector can be derived based on the right-side lower center sample position of the current block.

[0235] The decoding device can generate predicted samples of the current block based on movement information for subblocks within the current block (S2330).

[0236] In one embodiment, in the case of a prediction mode in which the decoding device performs a prediction on the current block based on subblock unit motion information (i.e., sbTMVP mode), the decoding device can generate a prediction sample of the current block based on the motion information of the current block to its subblocks derived as described above.

[0237] The decoding device can generate a reconstructed sample based on the predicted sample (S2340).

[0238] In one embodiment, the decoding device can immediately use the predicted sample as the reconstructed sample in prediction mode, or it can generate a reconstructed sample by adding a residual sample to the predicted sample.

[0239] The decoding device can receive information about residuals relative to the current block, if residual samples exist for the current block. The residual information may include conversion coefficients for the residual samples. Based on the residual information, the decoding device can derive residual samples (or residual sample arrays) relative to the current block. Based on the predicted samples and residual samples, the decoding device can generate reconstructed samples, and based on the reconstructed samples, can derive reconstructed blocks or reconstructed pictures. Subsequently, as previously stated, the decoding device may apply deblock filtering and / or in-loop filtering procedures such as SAO procedures to the reconstructed picture to improve subjective / objective image quality as needed.

[0240] In the embodiments described above, the method is explained based on a flowchart in a series of steps or blocks. However, the embodiments in this document are not limited to the order of the steps, and some steps may occur in a different order or simultaneously than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of this document.

[0241] The method described in this document can be implemented in software form, and the encoding and / or decoding devices described in this document can be included in devices that perform video processing, such as TVs, computers, smartphones, set-top boxes, and display devices.

[0242] When the embodiments described in this document are implemented in software, the methods described above can be implemented by modules (processes, functions, etc.) that perform the functions described above. These modules are stored in memory and can be executed by a processor. The memory may be internal or external to the processor and can be connected to it by a variety of well-known means. The processor may include an ASIC (Application-Specific Integrated Circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (Read-Only Memory), RAM (Random Access Memory), flash memory, memory cards, storage media, and / or other storage devices. In other words, the embodiments described in this document can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored in a digital storage medium.

[0243] Furthermore, decoding and encoding devices to which this document applies may include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providers, OTT video (Over The Top video) devices, internet streaming service providers, 3D video devices, VR (Virtual Reality) devices, AR (Augmented Reality) devices, image-phone video devices, transportation terminals (e.g., vehicles (including autonomous vehicles), aircraft terminals, ship terminals, etc.), and medical video equipment, and may be used to process video signals or data signals. For example, OTT video (Over The Top video) devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorders), etc.

[0244] Furthermore, the processing methods to which this document applies can be produced in the form of programs executed on a computer and stored on a computer-readable storage medium. Multimedia data having the data structure described in this document can also be stored on a computer-readable storage medium. The computer-readable storage medium includes all types of storage devices and distributed storage devices that store data that can be read by a computer. Examples of computer-readable storage media include Blu-ray discs (BDs), Universal Serial Bus (USB), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable storage medium also includes media embodied in the form of carrier waves (e.g., transmission over the Internet). Additionally, bitstreams generated by encoding methods can be stored on computer-readable storage media or transmitted over wireless networks.

[0245] Furthermore, the embodiments described in this document can be embodied in computer program products using program code, and the program code can be executed on a computer according to the embodiments described in this document. The program code can be stored on a computer-readable carrier.

[0246] Figure 24 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.

[0247] Referring to Figure 24, the content streaming system applicable to the embodiments described in this document can be broadly classified to include an encoding server, a streaming server, a web server, a media storage device (storage), a user device, and a multimedia input device.

[0248] The above-mentioned encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then transmitting this bitstream to the above-mentioned streaming server. In other cases, if the multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the above-mentioned encoding server can be omitted.

[0249] The bitstream described above can be generated by an encoding method or bitstream generation method applicable to the embodiments of this document, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving it.

[0250] The streaming server transmits multimedia data to user devices based on user requests via the web server, and the web server acts as an intermediary to inform the user about available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.

[0251] The above-mentioned streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from the above-mentioned encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the above-mentioned streaming server can store the above-mentioned bitstream for a certain period of time.

[0252] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), navigation systems, slate PCs, tablet PCs, ultrabooks (ULTRABOOK®), wearable devices (such as smartwatches, smart glasses, HMDs (Head Mounted Displays)), digital TVs, desktop computers, and digital signatures (Signiji).

[0253] Each server within the above content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.

[0254] The claims described in this document can be combined in various ways. For example, the technical features of the method claims in this document can be combined and embodied in an apparatus, and the technical features of the apparatus claims in this document can be combined and embodied in a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims in this document can be combined and embodied in an apparatus, and the technical features of the method claims and the technical features of the apparatus claims in this document can be combined and embodied in a method.

Claims

1. In a video decoding method performed by a decoding device, Steps include obtaining residual information from the bitstream, The steps include: deriving the collated subblocks in the collated picture for the subblocks in the current block, The steps include: deriving candidate time-motion information based on the aforementioned collated subblocks, The steps include: deriving motion information for the sub-block within the current block based on candidate time motion information based on the sub-block; The steps include generating a predicted sample of the current block based on the movement information for the subblock within the current block, The steps include generating a residual sample of the current block based on the residual information, A step of generating a reconstructed sample based on the predicted sample and the residual sample, The step of applying deblocking filtering to the restored sample is included, Based on the availability of the collated subblock, the position of each collated subblock within the collated picture is derived based on the position of the subblock within the current block. The respective positions of the collated subblocks are derived from the collated picture by applying a motion shift to the center sample position of the lower right end of each subblock within the current block. The motion shift is performed based on motion vectors derived from spatially adjacent blocks of the current block. The adjacent block in the space is the adjacent block on the left. The candidate time motion information based on the aforementioned subblock includes a subblock unit motion vector, A method in which the base motion vector is used as the subblock unit motion vector for collated subblocks that are unusable among the collated subblocks.

2. In a video encoding method performed by an encoding device, The steps include: deriving the collated subblocks in the collated picture for the subblocks in the current block, The steps include: deriving candidate time-motion information based on the aforementioned collated subblocks, The steps include: deriving motion information for the sub-block within the current block based on candidate time motion information based on the sub-block; The steps include generating a predicted sample of the current block based on the movement information for the subblock within the current block, The steps include: deriving residual samples based on the aforementioned predicted samples; The step includes encoding video information that includes information about the residual sample, Based on the availability of the collated subblock, the position of each collated subblock within the collated picture is derived based on the position of the subblock within the current block. The respective positions of the collated subblocks are derived from the collated picture by applying a motion shift to the center sample position of the lower right end of each subblock within the current block. The motion shift is performed based on motion vectors derived from spatially adjacent blocks of the current block. The adjacent block in the space is the adjacent block on the left. The candidate time motion information based on the aforementioned subblock includes a subblock unit motion vector, A method in which the base motion vector is used as the subblock unit motion vector for collated subblocks that are unusable among the collated subblocks.

3. A step of obtaining a bitstream relating to video, wherein the bitstream is The steps include: deriving the collated subblocks in the collated picture for the subblocks in the current block, The steps include: deriving candidate time-motion information based on the aforementioned collated subblocks, The steps include: deriving motion information for the sub-block within the current block based on candidate time motion information based on the sub-block; The steps include generating a predicted sample of the current block based on the movement information for the subblock within the current block, The steps include: deriving residual samples based on the aforementioned predicted samples; A step of encoding video information including information about the residual sample and outputting the bitstream, and a step of generating based on, The step of transmitting data including the bitstream, Based on the availability of the collated subblock, the position of each collated subblock within the collated picture is derived based on the position of the subblock within the current block. The position of each of the collated subblocks is derived from the collated picture by applying a motion shift to the center sample position of the lower right end of each of the subblocks within the current block. The motion shift is performed based on motion vectors derived from spatially adjacent blocks of the current block. The adjacent block in the space is the adjacent block on the left. The candidate time motion information based on the aforementioned subblock includes a subblock unit motion vector, A method in which the base motion vector is used as the subblock unit motion vector for collated subblocks that are unusable among the collated subblocks.