Image or video compilation of a sub-block unit based on time motion vector prediction sub-candidates

By derive the motion vector using the reference sub-block inside the juxtaposed reference picture in the sub-block prediction, the problem of low image and video compilation efficiency in the prior art is solved, and more efficient compression and transmission are achieved.

CN114080809BActive Publication Date: 2025-06-24LG ELECTRONICS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080049404.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-13
Filing Date
2020-06-15
Publication Date
2025-06-24
Estimated Expiration
2040-06-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively compress and transmit high-resolution, high-quality images and videos, especially in image or video compilation techniques based on time motion vector prediction involving sub-block units, where there are challenges to improve compilation efficiency.

Method used

In sub-block based time motion vector prediction, the sub-block unit motion vector for the current block is derived based on the reference sub-block inside the co-arranged reference picture, and the base motion vector is used in the case where the reference sub-block is not available.

Benefits of technology

Improves image/video compression efficiency, reduces computational complexity, simplifies hardware implementation, and improves inter-frame prediction performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114080809B_ABST
    Figure CN114080809B_ABST
Patent Text Reader

Abstract

According to the disclosure of this document, the sub-block position for deriving the motion vector of a sub-block unit in sub-block based temporal motion vector prediction (sbTMVP) can be effectively calculated, thereby enabling an improvement in video / image compilation efficiency and obtaining a simplified effect of hardware implementation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to video or image compilation, and for example, relates to an image or video compilation technology for time motion vector prediction sub-candidates of sub-block units. Background Art

[0002] Recently, the demand for high-resolution and high-quality images and videos such as ultra-high definition (HUD) images and 4K or 8K or larger videos has been increasing in various fields. As image and video data become high-resolution and high-quality, the amount of information or the number of bits transmitted relatively increases compared to existing image and video data. Therefore, if media such as existing wired or wireless broadband lines are used to transmit image data or existing storage media are used to store image and video data, the transmission cost and storage cost increase.

[0003] In addition, recently, the interest and demand for immersive media such as virtual reality (VR), artificial reality (AR) content, or holograms have been increasing. The broadcasting of images and videos with different image characteristics from those of real images, such as game images, has been increasing.

[0004] Therefore, in order to effectively compress and transmit or store and play back information of high-resolution and high-quality images and videos with such various characteristics, efficient image and video compression technologies are required.

[0005] In addition, in order to improve image / video compilation efficiency, a sub-block-based time motion vector prediction technology has been discussed. For this purpose, a scheme for effectively performing the process of patching the motion vectors of sub-block units in sub-block-based time motion vector prediction is required. Summary of the Invention

[0006] Technical Problem

[0007] The purpose of this document is to provide a method and apparatus for improving video / image compilation efficiency.

[0008] Another purpose of this document is to provide a method and apparatus for effective inter-frame prediction.

[0009] Another purpose of this document is to provide a method and apparatus for improving prediction performance by deriving a sub-block-based time motion vector.

[0010] Another purpose of this document is to provide a method and apparatus for effectively deriving the corresponding position of a sub-block to derive a sub-block-based time motion vector.

[0011] Another object of this document is to provide a method and apparatus for unifying corresponding positions at the sub-compilation block level and corresponding positions at the compilation block level to derive sub-block-based temporal motion vectors.

[0012] Technical solution

[0013] According to an embodiment of this document, a sub-block unit motion vector for a current block can be derived in sub-block-based temporal motion vector prediction (sbTMVP) based on a reference sub-block inside a collocated reference picture.

[0014] According to an embodiment of this document, a reference sub-block inside a collocated reference picture for a sub-block in a current block can be derived based on the center sample position of each of the sub-blocks in the current block.

[0015] According to an embodiment of this document, a basic motion vector can be used in the sub-block unit motion vector for an unavailable reference sub-block among the reference sub-blocks.

[0016] According to an embodiment of this document, a basic motion vector can be derived from a collocated reference picture based on the center sample position of the current block.

[0017] According to an embodiment of this document, a video / image decoding method executed by a decoding device is provided. The video / image decoding method may include the methods disclosed in the embodiments of this document.

[0018] According to an embodiment of this document, a decoding device for performing video / image decoding is provided. The decoding device may execute the methods disclosed in the embodiments of this document.

[0019] According to an embodiment of this document, a video / image encoding method executed by an encoding device is provided. The video / image encoding method may include the methods disclosed in the embodiments of this document.

[0020] According to an embodiment of this document, an encoding device for performing video / image encoding is provided. The encoding device may execute the methods disclosed in the embodiments of this document.

[0021] According to an embodiment of this document, a computer-readable digital storage medium is provided, in which encoded video / image information is generated according to the video / image encoding method disclosed in at least one embodiment of this document.

[0022] According to an embodiment of this document, a computer-readable digital storage medium is provided, in which encoding information or encoded video / image information that causes a decoding device to execute the video / image decoding method disclosed in at least one embodiment of this document is stored.

[0023] Beneficial effects

[0024] This document can have various effects. For example, it can improve the overall image / video compression efficiency. In addition, through effective inter-frame prediction, the computational complexity can be reduced and the overall compilation efficiency can be improved. Moreover, the efficiency in terms of complexity and prediction performance can be improved because, in sub-block based temporal motion vector prediction (sbTMVP), the corresponding positions of the sub-blocks for deriving the sub-block based temporal motion vectors are effectively calculated. Further, since the method for calculating the corresponding positions at the sub-compilation block level and the corresponding positions at the compilation block level for deriving the sub-block based temporal motion vectors is unified, a simplification effect in terms of hardware implementation can be obtained.

[0025] The effects that can be obtained through the detailed embodiments of this document are not limited to the listed effects. For example, there can be various technical effects that those of ordinary skill in the relevant art can understand or derive from this document. Therefore, the detailed effects of this document are not limited to the effects clearly described in this document, and can include various effects that can be understood or derived from the technical features of this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 Schematically illustrates an example of a video / image compilation system to which the embodiments of this document can be applied.

[0027] Figure 2 Is a diagram schematically depicting the configuration of a video / image encoding device to which the embodiments of this document can be applied.

[0028] Figure 3 Is a diagram schematically depicting the configuration of a video / image decoding device to which the embodiments of this document can be applied.

[0029] Figure 4 Illustrates an example of a schematic video / image encoding method to which the embodiments of this document can be applied.

[0030] Figure 5 Illustrates an example of a schematic video / image decoding method to which the embodiments of this document can be applied.

[0031] Figure 6 Illustrates an example of a schematic inter-frame prediction based video / image encoding method to which the embodiments of this document can be applied.

[0032] Figure 7 Illustrates an example of a schematic inter-frame prediction based video / image decoding method to which the embodiments of this document can be applied.

[0033] Figure 8 Exemplarily illustrates the inter-frame prediction process.

[0034] Figure 9The exemplary figure illustrates the spatial neighboring blocks and temporal neighboring blocks of the current block.

[0035] Figure 10 The exemplary figure illustrates the spatial neighboring blocks that can be used to derive sub - block - based temporal motion vector prediction candidate (sbTMVP candidate).

[0036] Figure 11 is a figure for schematically describing the process of deriving sub - block - based temporal motion vector prediction candidate (sbTMVP candidate).

[0037] Figures 12 to 15 is a figure for schematically describing the method of calculating the corresponding positions for deriving the default MV and sub - block MV based on the block size during the sbTMVP derivation process.

[0038] Figures 16 to 19 is an example figure for schematically describing the method of unifying the method of deriving the corresponding positions of the default MV and sub - block MV based on the block size during the sbTMVP derivation process.

[0039] Figure 20 and Figure 21 is an example figure for schematically illustrating the configuration of the pipeline through which, during the sbTMVP derivation process, the corresponding positions for deriving the default MV and sub - block MV can be unified and calculated.

[0040] Figure 22 Schematically illustrates an example of a video / image encoding method according to one or more embodiments of this document.

[0041] Figure 23 Schematically illustrates an example of a video / image decoding method according to one or more embodiments of this document.

[0042] Figure 24 Illustrates an example of a content streaming system to which the embodiments disclosed in this document can be applied. Detailed implementation

[0043] This document can be modified in various ways and can have various embodiments, and specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are used to describe specific embodiments, rather than to limit the technical spirit of this document. Unless otherwise clearly indicated in the context, singular expressions include plural expressions. Terms such as "including" or "having" in this specification should be understood as indicating the presence of the features, numbers, steps, operations, elements, components, or combinations thereof described in this specification, without excluding the possibility of the presence or addition of one or more features, numbers, steps, operations, elements, components, or combinations thereof.

[0044] Meanwhile, for the convenience of descriptions related to different feature functions, the elements in the drawings described in this document are independently illustrated. This does not mean that each element is implemented as a separate piece of hardware or separate software. For example, at least two elements can be combined to form a single element, or a single element can be divided into multiple elements. Embodiments in which elements are combined and / or separated are also included within the scope of the rights of this document, unless it deviates from the essence of this document.

[0045] In this document, the term "A or B" can mean "only A", "only B", or "both A and B". In other words, in this document, the term "A or B" can be interpreted as indicating "A and / or B". For example, in this document, the term "A, B, or C" can mean "only A", "only B", "only C", or "any combination of A, B, and C".

[0046] The slashes " / " or commas used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B, or C".

[0047] In this document, "at least one of A and B" can mean "only A", "only B", or "both A and B". In addition, in this document, the expressions "at least one of A or B" or "at least one of A and / or B" can be interpreted as "at least one of A and B".

[0048] Furthermore, in this document, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". In addition, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".

[0049] In addition, the parentheses used in this document can mean "for example". Specifically, in the case of the expression "prediction (intra prediction)", it can indicate that "intra prediction" is presented as an example of "prediction". In other words, the term "prediction" in this document is not limited to "intra prediction", and it can indicate that "intra prediction" is presented as an example of "prediction". In addition, even in the case of the expression "prediction (i.e., intra prediction)", it can also indicate that "intra prediction" is presented as an example of "prediction".

[0050] This document relates to video / image compilation. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the Versatile Video Coding (VVC). In addition, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the Emerging Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the Audio Video Coding Standard 2 (AVS2), or the next-generation video / image coding standards (e.g., H.267 or H.268, etc.).

[0051] This document presents various embodiments of video / image compilation, and unless otherwise mentioned, these embodiments can be executed in combination with each other.

[0052] In this document, video can mean a collection of a series of images according to the passage of time. A picture generally means a unit representing an image for a specific time period, and a slice / tile is a unit that forms part of a picture in compilation. A slice / tile can include one or more Coding Tree Units (CTUs). A picture can be composed of one or more slices / tiles. A tile is a rectangular area of CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular area where the height of the CTUs is equal to the height of the picture and the width is specified by a syntax element in the picture parameter set. A tile row is a rectangular area where the height of the CTUs is specified by a syntax element in the picture parameter set and the width is equal to the width of the picture. Tile scanning is a specific sequential ordering of the CTUs that divide a picture: the CTUs can be sequentially ordered in raster scan within a tile, while the tiles in a picture can be sequentially ordered in raster scan of the tiles of the picture. A slice includes an integer number of complete tiles or an integer number of consecutive complete CTU rows within the tiles of a picture that can be exclusively contained in a single NAL unit.

[0053] Meanwhile, a picture can be divided into two or more sub-pictures. A sub-picture can be a rectangular area of one or more slices within a picture.

[0054] A pixel or pel can mean the smallest unit that constitutes a picture (or image). Additionally, "sample" can be used as a term corresponding to a pixel. A sample generally can represent a pixel or the value of a pixel, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Alternatively, a sample can mean a pixel value in the spatial domain, or when the pixel value is transformed into the frequency domain, it can mean a transform coefficient in the frequency domain.

[0055] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to the area. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the situation, terms such as unit and blocks, regions, etc. may be used interchangeably. Usually, an M×N block may include a set (or array) of samples (or sample array) consisting of M columns and N rows or a set (or array) of transform coefficients.

[0056] In addition, in this document, at least one of quantization / dequantization and / or transformation / inverse transformation may be omitted. When quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When transformation / inverse transformation is omitted, the transform coefficients may be called coefficients or residual coefficients, or for the sake of uniformity of expression, may still be called transform coefficients.

[0057] In this document, the quantized transform coefficients and the transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information may include information about the transform coefficients, and information about the transform coefficients may be signaled by a residual coding syntax. The transform coefficients may be derived based on the residual information (or information about the transform coefficients), and the scaled transform coefficients may be derived by inverse-transforming (scaling) the transform coefficients. The residual samples may be derived based on the inverse transformation of the scaled transform coefficients. This may also be applied / expressed in other parts of this document.

[0058] In this document, the technical features separately explained in one figure may be implemented separately or may be implemented simultaneously.

[0059] Hereinafter, the preferred embodiments of this document will be described more specifically with reference to the drawings. Hereinafter, in the drawings, the same reference numerals are used for the same elements, and repeated descriptions of the same elements may be omitted.

[0060] Figure 1 An example of a video / image coding system to which the embodiments of this document can be applied is schematically illustrated.

[0061] Reference Figure 1 , the video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transfer the encoded video / image information or data to the receiving device in the form of a file or a stream via a digital storage medium or a network.

[0062] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0063] The video source may obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate video / images. For example, virtual video / images may be generated by a computer or the like. In this case, the video / image capture process may be replaced by a process of generating relevant data.

[0064] The encoding device may encode the input video / images. The encoding device may perform a series of processes such as prediction, transformation, and quantization for compression and compilation efficiency. The encoded data (encoded video / image information) may be output in the form of a bitstream.

[0065] The transmitter may send the encoded video / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream through a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include elements for generating a media file in a predetermined file format and may include elements for sending through a broadcast / communication network. The receiver may receive / extract the bitstream and send the received / extracted bitstream to the decoding device.

[0066] The decoding device may decode the video / images by performing a series of processes such as dequantization, inverse transformation, prediction, etc., corresponding to the operations of the encoding device.

[0067] The renderer may render the decoded video / images. The rendered video / images may be displayed through a display.

[0068] Figure 2 is a diagram schematically depicting the configuration of a video / image encoding device to which this document may be applied. Hereinafter, the encoding device may include an image encoding device and / or a video encoding device.

[0069] Reference Figure 2, the encoding device 200 may include an image splitter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to an embodiment, the above-described image splitter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be constituted by one or more hardware components (e.g., an encoder chipset or a processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.

[0070] The image splitter 210 may split an input image (or picture or frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a largest coding unit (LCU), the coding unit may be recursively split according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, a coding unit may be divided into multiple coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure may be applied first, and then the binary tree structure and / or the ternary tree structure may be applied. Alternatively, the binary tree structure may be applied first. The coding process according to this document may be performed based on the final coding unit that is not further split. In this case, based on the coding efficiency according to the image characteristics, the largest coding unit may be directly used as the final coding unit. Alternatively, the coding unit may be recursively split into coding units with an even deeper depth as needed, such that the coding unit with the optimal size may be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be divided or split from the above-described final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal based on the transformation coefficients.

[0071] Depending on the context, terms such as unit and terms like block, region, etc. can be used interchangeably. In a general case, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. Samples can generally represent pixels or pixel values, and can represent only the pixels / pixel values of the luminance component, or only the pixels / pixel values of the chrominance component. Samples can be used as a term corresponding to the pixels or pels of a picture (or image).

[0072] In the encoding device 200, a prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 is subtracted from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) in the encoder 200 can be referred to as the subtractor 231. The predictor can perform prediction on a processing target block (hereinafter referred to as the "current block") and can generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As discussed subsequently in the description of each prediction mode, the predictor can generate various prediction-related information such as prediction mode information and send the generated information to the entropy encoder 240. Information about the prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0073] The intra-frame predictor 222 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the reference samples can be located near the current block or separated from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The non-directional modes can include, for example, the DC mode and the plane mode. Depending on the level of detail of the prediction direction, the directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes can be used depending on the settings. The intra-frame predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to neighboring blocks.

[0074] The inter - frame predictor 221 can derive a prediction block for a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information can be predicted on a block, sub - block, or sample basis based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same as or different from each other. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter - frame predictor 221 can configure a motion information candidate list based on neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter - frame prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter - frame predictor 221 can use the motion information of neighboring blocks as the motion information of the current block. In the skip mode, different from the merge mode, a residual signal cannot be transmitted. In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of neighboring blocks can be used as a motion vector prediction term, and the motion vector of the current block can be indicated by signaling a motion vector difference.

[0075] The predictor 220 can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra - frame prediction or inter - frame prediction to the prediction of a block, and can also apply intra - frame prediction and inter - frame prediction simultaneously. This can be referred to as combined inter - frame and intra - frame prediction (CIIP). Additionally, the predictor can be based on the intra - block copy (IBC) prediction mode or the palette mode in order to perform prediction on a block. The IBC prediction mode or the palette mode can be used for content image / video compilation such as games like screen content coding (SCC). Although IBC basically performs prediction in the current picture, the way it is performed is similar to inter - frame prediction in that it derives a reference block in the current picture. That is, IBC can use at least one of the inter - frame prediction techniques described in this document. The palette mode can be regarded as an example of intra - frame coding or intra - frame prediction. When the palette mode is applied, the sample values in the picture can be signaled based on information about the palette index and the palette table.

[0076] The prediction signal generated by a predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) can be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, GBT means a transform obtained from a curve graph when the relational information between pixels is represented by the curve graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Additionally, the transform process can be applied to square pixel blocks of the same size, or can be applied to blocks of variable size rather than square blocks.

[0077] Quantizer 233 may quantize the transform coefficients and send them to entropy encoder 240, and entropy encoder 240 may encode the quantized signal (information regarding the quantized transform coefficients) and output the encoded signal in a bitstream. The information regarding the quantized transform coefficients may be referred to as residual information. Quantizer 233 may rearrange the quantized transform coefficients of a block type into a one-dimensional vector form based on the coefficient scan order, and generate information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Entropy encoder 240 may perform various coding methods such as, for example, exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), etc. Entropy encoder 240 may encode the information required for video / image reconstruction, in addition to the quantized transform coefficients (e.g., values of syntax elements, etc.), together or separately. The encoded information (e.g., encoded video / image information) may be sent or stored in the form of a bitstream on a unit basis of a network abstraction layer (NAL). The video / image information may also include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), etc. Additionally, the video / image information may also include general constraint information. In this document, the information and / or syntax elements sent to / signaled to the decoding device from the encoding device may be included in the video / picture information. The video / image information may be encoded through the above encoding process and included in the bitstream. The bitstream may be transmitted through a network or stored in a digital storage medium. Here, the network may include a broadcast network, a communication network, and / or the like, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that sends the signal output from entropy encoder 240 or a memory (not shown) that stores it may be configured as an internal / external element of encoding device 200, or the transmitter may be included in entropy encoder 240.

[0078] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transformation to the quantized transform coefficients via the dequantizer 234 and the inverse transformer 235, a residual signal (residual block or residual samples) can be reconstructed. The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222, so that a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) can be generated. When there is no residual for the target block to be processed as in the case of applying the skip mode, the predicted block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of the next target block in the current picture and, as described later, can be used for inter-prediction of the next picture by filtering.

[0079] In addition, in the picture encoding and / or reconstruction process, luminance mapping and chrominance scaling (LMCS) can be applied.

[0080] The filter 260 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and the modified reconstructed picture can be stored in the memory 270, especially in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. As discussed later in the description of each filtering method, the filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240. The information about the filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0081] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-prediction unit 221. Accordingly, the encoding device can avoid prediction mismatch in the encoding device 100 and the decoding device when applying inter-prediction, and can also improve the encoding efficiency.

[0082] The DPB of the memory 270 can store the modified reconstructed picture for using it as a reference picture in the inter-prediction unit 221. The memory 270 can store the motion information of the blocks in the current picture from which the motion information has been derived (or encoded) and / or the motion information of the blocks in the reconstructed pictures. The stored motion information can be sent to the inter-prediction unit 221 to be used as the motion information of neighboring blocks or temporally neighboring blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-prediction unit 222.

[0083] Figure 3FIG. is a diagram schematically depicting the configuration of a video / image decoding apparatus to which this document can be applied. Hereinafter, the decoding apparatus may include an image decoding apparatus and / or a video decoding apparatus.

[0084] Reference Figure 3 , the video decoding apparatus 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be constituted by one or more hardware components (e.g., a decoder chipset or a processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.

[0085] When receiving a bitstream including video / image information, the decoding apparatus 300 may reconstruct an image corresponding to the processing of the video / image information that has been processed in the Figure 2 encoding apparatus accordingly. For example, the decoding apparatus 300 may derive units / blocks based on information related to block segmentation obtained from the bitstream. The decoding apparatus 300 may perform decoding by using the processing units applied in the encoding apparatus. Thus, the decoding processing units may be, for example, coding units, which may be segmented from a coding tree unit or a largest coding unit along a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding units. The reconstructed image signal decoded and output by the decoding apparatus 300 may be reproduced by a reproducing apparatus.

[0086] The decoding apparatus 300 may receive, in the form of a bitstream, from Figure 2The signal output by the encoding device, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can also include information about various parameter sets such as adaptive parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), video parameter set (VPS), etc. Additionally, the video / image information can further include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the general constraint information. In this document, the information and / or syntax elements transmitted / received subsequently can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information in the bitstream based on coding methods such as exponential Golomb coding, CAVLC, or CABAC, and can output the values of the syntax elements required for image reconstruction and the quantization values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, use the decoding target syntax element information and the decoding information of the neighboring and decoding target blocks or the information of the symbols / bins decoded in the previous step to determine the context model, predict the bin generation probability according to the determined context model, and perform arithmetic decoding on the bins to generate symbols corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model using the information of the symbols / bins decoded by the context model for the next symbol / bin after determining the context model. Among the information decoded in the entropy decoder 310, the information about prediction can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantized transform coefficients) and associated parameter information for which entropy decoding has been performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive the residual signal (residual block, residual sample, residual sample array). Additionally, the information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not shown) that receives the signal output by the encoding device can also configure the decoding device 300 as an internal / external component, and the receiver can be a component of the entropy decoder 310. Additionally, the decoding device according to this document can be referred to as a video / image / picture decoding device, and the decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0087] The dequantizer 321 can output transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into the form of a two-dimensional block. In this case, the rearrangement can be performed based on the order of coefficient scanning that has been performed in the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.

[0088] The inverse transformer 322 obtains a residual signal (residual block, residual sample array) by performing an inverse transform on the transform coefficients.

[0089] The predictor can perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on the prediction information output from the entropy decoder 310 and can determine a specific intra / inter prediction mode.

[0090] The predictor 320 can generate a prediction signal based on various prediction methods. For example, the predictor can apply not only intra prediction or inter prediction to the prediction of a block but also apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). Additionally, the predictor can be based on the intra block copy (IBC) prediction mode or the palette mode in order to perform prediction on the block. The IBC prediction mode or the palette mode can be used for content image / video coding such as games like screen content coding (SCC). Although IBC basically performs prediction in the current picture, the way it performs is similar to inter prediction in that it derives a reference block in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.

[0091] The intra predictor 331 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the reference samples can be located near the current block or separated from the current block. In intra prediction, the prediction mode can include multiple non - directional modes and multiple directional modes. The intra predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to neighboring blocks.

[0092] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information can be predicted on a block, sub-block, or sample basis based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter-frame prediction for the current block.

[0093] The adder 340 adds the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (inter-frame predictor 332 or intra-frame predictor 331), so that a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) can be generated. When there is no residual for the block to be processed, such as in the case of applying the skip mode, the prediction block can be used as the reconstructed block.

[0094] The adder 340 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, can be output through filtering as described below, or can be used for inter-frame prediction of the next picture.

[0095] In addition, luminance mapping and chrominance scaling (LMCS) can be applied to the picture decoding process.

[0096] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 360, specifically, in the DPB of the memory 360. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0097] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter - predictor 332. The memory 360 can store the motion information of the blocks from which the motion information in the current picture is derived (or decoded) and / or the motion information of the blocks in the already - reconstructed pictures. The stored motion information can be sent to the inter - predictor 332 to be used as the motion information of spatially - adjacent blocks or temporally - adjacent blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and transfer the reconstructed samples to the intra - predictor 331.

[0098] In the present disclosure, the embodiments described in the filter 260, the inter - predictor 221, and the intra - predictor 222 of the encoding device 200 can be the same as or respectively corresponding to the filter 350, the inter - predictor 332, and the intra - predictor 331 of the decoding device 300 and be applied. This can also be applied to the unit 332 and the intra - predictor 331.

[0099] As described above, when performing video encoding, prediction is performed to improve the compression efficiency. A prediction block including prediction samples for the current block, that is, the target encoding block, can be generated through prediction. In this case, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived in the same way in both the encoding device and the decoding device. The encoding device can improve the image encoding efficiency by signaling to the decoding device information about the residual between the original block rather than the original sample values of the original block and the prediction block (residual information). The decoding device can derive a residual block including residual samples based on the residual information, can generate a reconstructed block including reconstructed samples by adding the residual block and the prediction block, and can generate a reconstructed picture including the reconstructed block.

[0100] The residual information can be generated through a transformation and quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, can derive transform coefficients by performing a transformation process on the residual samples (residual sample array) included in the residual block, can derive quantized transform coefficients by performing a quantization process on the transform coefficients, and can signal the relevant residual information to the decoding device (through the bitstream). In this case, the residual information can include information such as value information, position information, transformation scheme, transformation kernel, and quantization parameters of the quantized transform coefficients. The decoding device can perform a de - quantization / inverse - transformation process based on the residual information and can derive the residual samples (or residual block). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. In addition, the encoding device can derive a residual block by performing de - quantization / inverse - transformation on the quantized transform coefficients for use as a reference for inter - prediction of subsequent pictures and can generate a reconstructed picture.

[0101] Figure 4 An example of a schematic video / image encoding method to which the embodiments of this document are applicable is illustrated.

[0102] Figure 4 The method disclosed in Figure 2 can be executed by the above-mentioned encoding apparatus 200. Specifically, S400 can be executed by the inter-frame predictor 221 or the intra-frame predictor 222 of the encoding apparatus 200, and S410, S420, S430, and S440 can be executed by the subtractor 231, the transformer 232, the quantizer 233, and the entropy encoder 240 of the encoding apparatus 200.

[0103] Referring to Figure 4 , the encoding apparatus can derive a prediction sample by predicting the current block (S400). The encoding apparatus can determine whether to perform inter-frame prediction or intra-frame prediction on the current block, and can determine a specific inter-frame prediction mode or a specific intra-frame prediction mode based on the RD cost. According to the determined mode, the encoding apparatus can derive the prediction sample of the current block.

[0104] The encoding apparatus can derive a residual sample by comparing the initial sample and the prediction sample of the current block (S410).

[0105] The encoding apparatus can derive transform coefficients through a transformation process for the residual sample (S420), and quantize the derived transform coefficients to derive quantized transform coefficients (S430).

[0106] The encoding apparatus can encode the image information including prediction information and residual information, and output the encoded image information in the form of a bitstream (S440). The prediction information is information related to the prediction process, and can include prediction mode information and motion information (e.g., when inter-frame prediction is applied). The residual information can include information about the quantized transform coefficients. The residual information can be entropy-coded.

[0107] The output bitstream can be transmitted to the decoding apparatus through a storage medium or a network.

[0108] Figure 5 An example of a schematic video / image decoding method to which the embodiments of the present disclosure are applicable is shown.

[0109] Figure 5 The method disclosed in Figure 3The decoding apparatus 300 performs it. Specifically, S500 may be performed by the inter - frame predictor 332 or the intra - frame predictor 331 of the decoding apparatus 300. The process of deriving the values of relevant syntax elements by decoding the prediction information included in the bitstream in S500 may be performed by the entropy decoder 310 of the decoding apparatus 300. S510, S520, S530, and S540 may be performed by the entropy decoder 310, the de - quantizer 321, the inverse transformer 322, and the adder 340 of the decoding apparatus 300, respectively.

[0110] Reference Figure 5 , the decoding apparatus may perform operations corresponding to those performed by the encoding apparatus. The decoding apparatus may perform inter - frame prediction or intra - frame prediction on the current block based on the received prediction information, and derive predicted samples (S500).

[0111] The decoding apparatus may derive quantization transform coefficients for the current block based on the received residual information (S510). The decoding apparatus may derive quantization transform coefficients from the residual information through entropy decoding.

[0112] The decoding apparatus may de - quantize the quantized transform coefficients to derive transform coefficients (S520).

[0113] The decoding apparatus derives residual samples through the inverse transformation process of the transform coefficients (S530).

[0114] The decoding apparatus may generate reconstructed samples for the current block based on the predicted samples and the residual samples, and generate a reconstructed picture based on them (S540). As described above, the in - loop filtering process may be further applied to the reconstructed picture subsequently.

[0115] Meanwhile, as described above, when performing prediction on the current block, intra - frame prediction or inter - frame prediction may be applied. Hereinafter, the case of applying inter - frame prediction to the current block will be described.

[0116] The predictor (more specifically, the inter - frame predictor) of an encoding / decoding device can derive a prediction sample by performing inter - frame prediction on a block - by - block basis. Inter - frame prediction can represent a prediction derived by a method that depends on data elements (e.g., sample values or motion information) of pictures other than the current picture. When applying inter - frame prediction to a current block, a prediction block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture indicated by a reference picture index. In this case, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information of the current block can be predicted on a block, sub - block, or sample basis based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter - frame prediction type (L0 prediction, L1 prediction, Bi - prediction, etc.) information. In the case of applying inter - frame prediction, neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same as or different from each other. Temporal neighboring blocks can be referred to by names such as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, a motion information candidate list can be configured based on neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) can be signaled in order to derive the motion vector and / or reference picture index of the current block. Inter - frame prediction can be performed based on various prediction modes, and for example, in the case of the skip mode and the merge mode, the motion information of the current block can be the same as that of the selected neighboring block. In the case of the skip mode, different from the merge mode, a residual signal may not be sent. In the case of the motion vector prediction (MVP) mode, the motion vector of the selected neighboring block can be used as a motion vector predictor and a motion vector difference can be signaled. In this case, the motion vector of the current block can be derived by using the sum of the motion vector predictor and the motion vector difference.

[0117] The motion information may further include L0 motion information and / or L1 motion information according to an inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). The L0-direction motion vector may be referred to as the L0 motion vector or MVL0, and the L1-direction motion vector may be referred to as the L1 motion vector or MVL1. The prediction based on the L0 motion vector may be referred to as L0 prediction, the prediction based on the L1 motion vector may be referred to as L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector may be referred to as bi-prediction. Herein, the L0 motion vector may indicate a motion vector associated with the reference picture list L0, and the L1 motion vector may indicate a motion vector associated with the reference picture list L1. The reference picture list L0 may include pictures before the current picture in the output order, and the reference picture list L1 may include pictures after the current picture in the output order as reference pictures. The previous pictures may be referred to as forward (reference) pictures and the subsequent pictures may be referred to as backward (reference) pictures. The reference picture list L0 may further include pictures after the current picture in the output order as reference pictures. In this case, the previous pictures may be indexed first in the reference picture list L0, and then the subsequent pictures may be indexed. The reference picture list L1 may further include pictures before the current picture in the output order as reference pictures. In this case, the subsequent pictures may be indexed first in the reference picture list L1, and then the previous pictures may be indexed. Herein, the output order may correspond to the picture order count (POC) order.

[0118] In addition, various inter-frame prediction modes may be used to predict a current block in a picture. For example, various modes such as a merge mode, a skip mode, a motion vector prediction (MVP) mode, an affine mode, a sub-block merge mode, a merge with MVD (MMVD) mode, and a history motion vector prediction (HMVP) mode may be used. A decoder-side motion vector refinement (DMVR) mode, an adaptive motion vector resolution (AMVR) mode, a bi-prediction with CU-level weight (BCW), a bidirectional optical flow (BDOF), etc. may be further used as additional modes. The affine mode may also be referred to as an affine motion prediction mode. The MVP mode may also be referred to as an advanced motion vector prediction (AMVP) mode. In this document, some modes and / or motion information candidates derived from some modes may also be included in one of the motion information-related candidates of other modes. For example, an HMVP candidate may be added to the merge candidates of the merge / skip mode, or may also be added to the mvp candidates of the MVP mode. If the HMVP candidate is used as a motion information candidate of the merge mode or the skip mode, the HMVP candidate may be referred to as an HMVP merge candidate.

[0119] Prediction mode information indicating an inter - frame prediction mode of a current block can be signaled from an encoding device to a decoding device. In this case, the prediction mode information can be included in a bitstream and received by the decoding device. The prediction mode information can include index information indicating one of a plurality of candidate modes. Alternatively, the inter - frame prediction mode can be indicated by hierarchical signaling of flag information. In this case, the prediction mode information can include one or more flags. For example, it can be signaled whether to apply a skip mode by signaling a skip flag, and when the skip mode is not applied, it can be signaled whether to apply a merge mode by signaling a merge flag, and a flag for additional discrimination can be further signaled when indicating to apply an MVP mode or when the merge mode is not applied. The affine mode can be signaled as an independent mode or as a dependent mode with respect to the merge mode or the MVP mode. For example, the affine mode can include an affine merge mode and an affine MVP mode.

[0120] In addition, when inter - frame prediction is applied to a current block, motion information of the current block can be used. The encoding device can derive optimal motion information for the current block through a motion estimation process. For example, the encoding device can search for a similar reference block with high correlation in a predetermined search range in a reference picture in units of fractional pixels using an original block in an original picture for the current block, and derive motion information based on the searched reference block. The similarity of blocks can be derived based on the difference in sample values based on phase. For example, the similarity of a block can be calculated based on the sum of absolute differences (SAD) between the current block (or a template of the current block) and the reference block (or a template of the reference block). In this case, the motion information can be derived based on the reference block having the minimum SAD in the search area. The derived motion information can be signaled to the decoding device according to various methods based on the inter - frame prediction mode.

[0121] A prediction block for a current block can be derived based on motion information derived according to an inter-frame prediction mode. The prediction block may include prediction samples (prediction sample array) of the current block. When the motion vector (MV) of the current block indicates a fractional sample unit, an interpolation process may be performed, and through the interpolation process, prediction samples of the current block may be derived based on reference samples of the fractional sample unit in a reference picture. When affine inter-frame prediction is applied to the current block, prediction samples may be generated based on sample / sub-block unit MVs. When dual prediction is applied, prediction samples derived by weighted summation or weighted averaging of prediction samples derived based on L0 prediction (i.e., prediction using reference pictures in reference picture list L0 and MVL0) and (depending on the phase) prediction samples derived based on L1 prediction (i.e., prediction using reference pictures in reference picture list L1 and MVL1) may be used as prediction samples of the current block. When dual prediction is applied, if the reference picture used for L0 prediction and the reference picture used for L1 prediction are in different temporal directions based on the current picture (i.e., if the prediction corresponds to dual prediction and bi-directional prediction), this may be referred to as true dual prediction.

[0122] Reconstructed samples and a reconstructed picture may be generated based on the derived prediction samples, and thereafter, a process such as in-loop filtering may be performed as described above.

[0123] Figure 6 An example of a schematic inter-frame prediction-based video / image encoding method to which embodiments of this document are applicable is illustrated.

[0124] Figure 6 The method disclosed in Figure 2 may be performed by the above-described encoding apparatus 200. Specifically, S600 may be performed by the inter-frame predictor 221 of the encoding apparatus 200, S610 may be performed by the subtractor 231 of the encoding apparatus 200, and S620 may be performed by the entropy encoder 240 of the encoding apparatus 200.

[0125] Refer to Figure 6, the encoding device can perform inter prediction on the current block (S600). The encoding device can derive an inter prediction mode and motion information of the current block, and generate a prediction sample of the current block. Here, the inter prediction mode determination process, the motion information derivation process, and the prediction sample generation process can be performed simultaneously, and any one of the processes can be performed earlier than the other processes. For example, the inter prediction unit of the encoding device can include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit, and the prediction mode determination unit can determine the prediction mode of the current block, the motion information derivation unit can derive the motion information of the current block, and the prediction sample derivation unit can derive the prediction sample of the current block. For example, the inter prediction unit of the encoding device can search for a block similar to the current block in a predetermined area (search area) of a reference picture through motion estimation, and derive a reference block with the smallest difference from the current block or equal to or less than a predetermined criterion. A reference picture index indicating the reference picture where the reference block is located can be derived based on this, and a motion vector can be derived based on the position difference between the reference block and the current block. The encoding device can determine a mode applied to the current block among various prediction modes. The encoding device can compare the RD costs of various prediction modes and determine the best prediction mode of the current block.

[0126] For example, when the skip mode or the merge mode is applied to the current block, the encoding device can configure a merge candidate list described below, and derive a reference block with the smallest difference from the current block or equal to or less than a predetermined criterion among the reference blocks indicated by the merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block can be selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the current block can be derived by using the motion information of the selected merge candidate.

[0127] As another example, when the (A)MVP mode is applied to the current block, the encoding device can configure an (A)MVP candidate list and use the motion vector of the selected mvp candidate among the motion vector predictors (mvps) included in the (A)MVP candidate list as the mvp of the current block. In this case, for example, the motion vector of the reference block derived through motion estimation can be used as the motion vector of the current block, and the mvp candidate with the smallest difference from the motion vector of the current block among the mvp candidates can become the selected mvp candidate. A motion vector difference (MVD), which is the difference obtained by subtracting the mvp from the motion vector of the current block, can be derived. In this case, information about the MVD can be signaled to the decoding device. In addition, when the (A)MVP mode is applied, the value of the reference picture index can be configured as reference picture index information and signaled to the decoding device separately.

[0128] The encoding device may derive a residual sample based on a prediction sample (S610). The encoding device may derive a residual sample by comparing an initial sample and a prediction sample of a current block.

[0129] The encoding device may encode image information including prediction information and residual information (S620). The encoding device may output the encoded image information in the form of a bitstream. Here, the prediction information may include information about prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and information about motion information, as information related to the prediction process. The information about motion information may include candidate selection information (e.g., merge index, mvp flag, or mvp index), which is information for deriving a motion vector. In addition, the information about motion information may include information about MVD and / or reference picture index information. In addition, the information about motion information may include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information about the residual sample. The residual information may include information about quantization transform coefficients for the residual sample.

[0130] The output bitstream may be stored in a (digital) storage medium and transmitted to a decoding device or transmitted to the decoding device via a network.

[0131] Meanwhile, as described above, the encoding device may generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on a reference sample and a residual sample. This is to derive the same prediction result as the prediction result executed by the decoding device, and thus, the compilation efficiency can be improved. Therefore, the encoding device may store the reconstructed picture (or reconstructed samples or reconstructed blocks) in a memory and use the reconstructed picture as a reference picture. As described above, an in-loop filtering process may be further applied to the reconstructed picture.

[0132] Figure 7 An example of a schematic inter-frame prediction-based video / image decoding method to which the embodiments of this document are applicable is shown.

[0133] Figure 7 The method disclosed in may be executed by the decoding device 300 described above Figure 3 Specifically, S700 may be executed by the inter-frame predictor 332 of the decoding device 300. Through the entropy decoder 310 of the decoding device 300, in S700, a process of decoding the prediction information included in the bitstream to derive the values of relevant syntax elements is executed. S710 and S720 may be executed by the inter-frame predictor 332 of the decoding device 300, S730 may be executed by the residual processor 320 of the decoding device 300, and S740 may be executed by the adder 340 of the decoding device 300.

[0134] Refer to Figure 7, the decoding device can perform operations corresponding to those performed by the encoding device. The decoding device can perform prediction on a current block based on the received prediction information and derive predicted samples.

[0135] Specifically, the decoding device can determine a prediction mode of the current block based on the received prediction information (S700). The decoding device can determine which inter prediction mode to apply to the current block based on the prediction mode information in the prediction information.

[0136] For example, based on a merge flag, it can be determined whether the merge mode or the (A)MVP mode is applied to the current block. Alternatively, one of various inter prediction mode candidates can be selected based on a mode index. The inter prediction mode candidates can include a skip mode, a merge mode, and / or the (A)MVP mode, or can include various inter prediction modes described above.

[0137] The decoding device derives motion information of the current block based on the determined inter prediction mode (S710). For example, when the skip mode or the merge mode is applied to the current block, the decoding device can configure a merge candidate list and select one merge candidate among the merge candidates included in the merge candidate list. Here, the selection can be performed based on selection information (merge index). The motion information of the current block can be derived by using the motion information of the selected merge candidate. The motion information of the selected merge candidate can be used as the motion information of the current block.

[0138] As another example, when the (A)MVP mode is applied to the current block, the decoding device can configure an (A)MVP candidate list, and use the motion vector of the selected mvp candidate among the motion vector predictors (mvps) included in the (A)MVP candidate list as the mvp of the current block. Here, the selection can be performed based on selection information (mvp flag or mvp index). In this case, the MVD of the current block can be derived based on information about the MVD, and the motion vector of the current block can be derived based on the mvp and the MVD of the current block. In addition, the reference picture index of the current block can be derived based on the reference picture index information. The picture indicated by the reference picture index in the reference picture list of the current block can be derived as the reference picture referred to for the inter prediction of the current block.

[0139] Meanwhile, the motion information of the current block can be derived without configuring a candidate list. In this case, the motion information of the current block can be derived according to the process disclosed in the prediction mode. In this case, the configuration of the candidate list can be omitted.

[0140] The decoding device may generate prediction samples for the current block based on the motion information of the current block (S720). In this case, the reference picture may be derived based on the reference picture index of the current block, and the prediction samples of the current block may be derived by using the samples of the reference block indicated by the motion vector of the current block on the reference picture. In this case, in some cases, a prediction sample filtering process may be further performed for all or some of the prediction samples of the current block.

[0141] For example, the inter prediction unit of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit may determine the prediction mode of the current block based on the received prediction mode information. The motion information derivation unit may derive the motion information (motion vector and / or reference picture index) of the current block based on the information regarding the received motion information. And the prediction sample derivation unit may derive the prediction samples of the current block.

[0142] The decoding device generates residual samples for the current block based on the received residual information (S730). The decoding device may generate reconstructed samples for the current block based on the prediction samples and the residual samples, and generate a reconstructed picture based on the generated reconstructed samples (S740). Thereafter, as described above, an in-loop filtering process may be further applied to the reconstructed picture.

[0143] Figure 8 An inter prediction process is exemplarily illustrated. Figure 8 The inter prediction process disclosed in Figure 6 and Figure 7 may be applied to the inter prediction process illustrated in

[0144] Reference Figure 8 , as described above, the inter prediction process may include an inter prediction mode determination step, a motion information derivation step according to the determined prediction mode, and a prediction process (prediction sample generation) step based on the derived motion information. As described above, the inter prediction process may be performed by the encoding device and the decoding device. In this document, the compiling device may include the encoding device and / or the decoding device.

[0145] The encoding device can determine an inter - prediction mode for the current block (S800). Various inter - prediction modes can be used to predict the current block in a picture. For example, various modes such as the merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub - block merge mode, merge with MVD (MMVD) mode, and history motion vector prediction (HMVP) mode can be used. Decoder - side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, bi - prediction with CU - level weights (BCW), bi - directional optical flow (BDOF), etc. can be further used as additional modes. The affine mode can also be referred to as the affine motion prediction mode. The MVP mode can also be referred to as the advanced motion vector prediction (AMVP) mode. In this document, some modes and / or candidates of motion information derived from some modes can also be included in one of the motion - information - related candidates of other modes. For example, the HMVP candidate can be added to the merge candidates of the merge / skip mode, or can also be added to the mvp candidates of the MVP mode. If the HMVP candidate is used as a motion - information candidate for the merge mode or skip mode, the HMVP candidate can be referred to as the HMVP merge candidate.

[0146] The prediction - mode information indicating the inter - prediction mode of the current block can be signaled from the encoding device to the decoding device. In this case, the prediction - mode information can be included in the bitstream and received by the decoding device. The prediction - mode information can include index information indicating one of multiple candidate modes. Alternatively, the inter - prediction mode can be indicated by hierarchical signaling of flag information. In this case, the prediction - mode information can include one or more flags. For example, whether the skip mode is applied can be indicated by signaling a skip flag, whether the merge mode is applied can be indicated by signaling a merge flag when the skip mode is not applied, and a flag for additional discrimination can be further signaled when indicating the application of the MVP mode or when the merge mode is not applied. The affine mode can be signaled as an independent mode or as a subordinate mode with respect to the merge mode or MVP mode. For example, the affine mode can include an affine merge mode and an affine MVP mode.

[0147] The encoding device can derive motion information for the current block (S810). The motion - information derivation can be based on the inter - prediction mode.

[0148] The encoding device may perform inter prediction using the motion information of the current block. The encoding device may derive the optimal motion information for the current block through a motion estimation process. For example, the encoding device may search for similar reference blocks with high correlation in a predetermined search range in the reference picture with fractional pixels as units for the original block in the original picture for the current block, and derive the motion information through the searched reference blocks. The similarity of the block may be derived based on the difference of the sample values based on the phase. For example, the similarity of the block may be calculated based on the sum of absolute differences (SAD) between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, the motion information may be derived based on the reference block with the minimum SAD in the search area. The derived motion information may be signaled to the decoding device according to various methods based on the inter prediction mode.

[0149] The encoding device may perform inter prediction (S820) based on the motion information for the current block. The encoding device may derive the (one or more) predicted samples for the current block based on the motion information. The current block including the predicted samples may be referred to as a prediction block.

[0150] Meanwhile, when deriving the motion information of the current block, (one or more) motion information candidates may be derived based on (one or more) spatially neighboring blocks and (one or more) temporally neighboring blocks, and a motion information candidate for the current block may be selected based on the derived (one or more) motion information candidates. At this time, the selected motion information candidate may be used as the motion information of the current block.

[0151] Figure 9 The exemplary diagram shows the spatially neighboring blocks and temporally neighboring blocks of the current block.

[0152] Reference Figure 9 , the spatially neighboring blocks refer to the neighboring blocks located around the current block 900 (which is the target for performing inter prediction currently), and may include the neighboring blocks located around the left side of the current block 900 or the neighboring blocks located around the upper side of the current block 900. For example, the spatially neighboring blocks may include the lower left neighboring block, left neighboring block, upper right neighboring block, upper neighboring block, and upper left neighboring block of the current block 900. Figure 9 The spatially neighboring blocks are illustrated as “S”.

[0153] According to an exemplary embodiment, the encoding device / decoding device may detect available neighboring blocks by searching the spatially neighboring blocks (e.g., lower left neighboring block, left neighboring block, upper right neighboring block, upper neighboring block, and upper left neighboring block) of the current block in a predetermined order, and derive the motion information of the detected neighboring blocks as spatially motion information candidates.

[0154] A temporally neighboring block is a block located on a picture different from the current picture including the current block 900 (i.e., a reference picture), and refers to a juxtaposed block of the current block 900 in the reference picture. Here, the reference picture can be before or after the current picture in terms of the picture order count (POC). In addition, the reference picture used to derive the temporally neighboring block can be referred to as a juxtaposed reference picture or a col picture (juxtaposed picture). Further, the juxtaposed block can refer to a block located at a position corresponding to the position of the current block 900 in the col picture, and is called a col block. For example, as Figure 9 shown, the temporally neighboring block can include a col block (i.e., a col block including the bottom-right sample) located at a position corresponding to the position of the bottom-right sample of the current block 900 in the reference picture (i.e., the col picture) and / or a col block (i.e., a col block including the bottom-right center sample) located at a position corresponding to the position of the bottom-right center sample of the current block 900 in the reference picture (i.e., the col picture). Figure 9 The temporally neighboring block is illustrated as "T".

[0155] According to an exemplary embodiment, the encoding device / decoding device can detect available blocks by searching for temporally neighboring blocks of the current block (e.g., a col block including the bottom-right sample and a col block including the bottom-right center sample) in a predetermined order, and derive the motion information of the detected blocks as temporally motion information candidates. As described above, the technique using temporally neighboring blocks can be referred to as temporal motion vector prediction (TMVP). In addition, the temporally motion information candidates can be referred to as TMVP candidates.

[0156] Meanwhile, prediction can also be performed by deriving motion information in units of sub-blocks according to the inter-frame prediction mode. For example, in the affine mode or the TMVP mode, motion information can be derived in units of sub-blocks. Specifically, the method for deriving temporally motion information candidates in units of sub-blocks can be referred to as sub-block based temporal motion vector prediction (sbTMVP) candidates.

[0157] sbTMVP is a method that uses the sports field within the col picture to improve the motion vector prediction (MVP) and merge mode of the coding unit within the current picture. The col picture of sbTMVP can be the same as the col picture used by TMVP. However, in TMVP, motion prediction is performed at the coding unit (CU) level. In contrast, in sbTMVP, motion prediction can be performed at the sub-block level or sub-coding unit (sub-CU) level. Additionally, in TMVP, the temporal motion information is derived from the col block within the col picture (in this case, the col block is the col block corresponding to the bottom-right sample position of the current block or the center bottom-right sample position of the current block). In sbTMVP, after applying a motion shift from the col picture, the temporal motion information is derived. In this case, the motion shift can include the process of obtaining a motion vector from one of the spatial neighboring blocks of the current block and shifting by that motion vector.

[0158] Figure 10 The exemplary diagram shows the spatial neighboring blocks that can be used to derive the sub-block-based temporal motion information candidate (sbTMVP candidate).

[0159] Reference Figure 10 , the spatial neighboring blocks can include at least one of the bottom-left neighboring block A0, left neighboring block A1, top-right neighboring block B0, and top neighboring block B1 of the current block. In some cases, the spatial neighboring blocks can further include another neighboring block in addition to Figure 10 the neighboring blocks shown, or may not include Figure 10 a specific neighboring block among the neighboring blocks shown. Additionally, the spatial neighboring blocks can include only specific neighboring blocks, for example, can include only the left neighboring block A1 of the current block.

[0160] For example, the encoding device / decoding device can first detect the motion vectors of the available spatial neighboring blocks while searching the spatial neighboring blocks in a predetermined search order, and can determine the block at the position indicated by the motion vectors of the spatial neighboring blocks in the reference picture as the col block (i.e., the collocated reference block). In this case, the motion vectors of the spatial neighboring blocks can be represented as temporal motion vectors (temporal MVs).

[0161] In this case, it can be determined whether the spatial neighboring block is available based on the reference picture information, prediction mode information, position information, etc. of the spatial neighboring block. For example, if the reference picture of the spatial neighboring block is the same as the reference picture of the current block, it can be determined that the corresponding spatial neighboring block is available. Alternatively, if the spatial neighboring block is coded in the intra prediction mode or the spatial neighboring block is outside the current picture / tile, it can be determined that the corresponding spatial neighboring block is unavailable.

[0162] In addition, the search order of spatially neighboring blocks can be defined differently, and can be in the order of, for example, A1, B1, B0, and A0. Alternatively, it can be determined whether A1 is available by searching only A1.

[0163] Figure 11 is a diagram for schematically describing the process of deriving a sub-block based temporal motion information candidate (sbTMVP candidate).

[0164] Refer to Figure 11 , first, the encoding / decoding device can determine whether a spatially neighboring block (e.g., block A1) of the current block is available. For example, if the reference picture of the spatially neighboring block (e.g., block A1) uses the col picture, it can be determined that the spatially neighboring block (e.g., block A1) is available, and the motion vector of the spatially neighboring block (e.g., block A1) can be derived. In this case, the motion vector of the spatially neighboring block (e.g., block A1) can be represented as a temporal MV (tempMV), and the motion vector can be used in motion shifting. Alternatively, if it is determined that the spatially neighboring block (e.g., block A1) is not available, the temporal MV (i.e., the motion vector of the spatially neighboring block) can be set to a zero vector. In other words, in this case, a motion vector set to (0, 0) can be applied to motion shifting.

[0165] Next, the encoding / decoding device can apply motion shifting based on the motion vector of the spatially neighboring block (e.g., block A1). For example, the motion shifting can be shifted (e.g., A1') to the position indicated by the motion vector of the spatially neighboring block (e.g., block A1). That is, by applying motion shifting, the motion vector of the spatially neighboring block (e.g., block A1) can be added to the coordinates of the current block.

[0166] Next, the encoding / decoding device can derive the juxtaposed sub-blocks (col sub-blocks) of the motion shifting on the col picture, and can obtain the motion information (motion vector, reference index, etc.) of each col sub-block. For example, the encoding / decoding device can derive each col sub-block on the col picture corresponding to the motion shifting position (i.e., the position indicated by the motion vector of the spatially neighboring block (e.g., A1)) at each sub-block position within the current block. In addition, the motion information of each col sub-block can be used as the motion information (i.e., sbTMVP candidate) of each sub-block of the current block.

[0167] In addition, scaling can be applied to the motion vectors of the col sub-blocks. The scaling can be performed based on the temporal distance difference between the reference picture of the col block and the reference picture of the current block. Thus, the scaling can be expressed as temporal motion scaling, and thus the reference pictures of the current block and the reference picture of the temporal motion vector can be arranged. In this case, the encoding / decoding device can obtain the scaled motion vectors of the col sub-blocks as the motion information of each sub-block of the current block.

[0168] In addition, when deriving sbTMVP candidates, motion information may not exist in the col sub-blocks. In this case, for the col sub-blocks where motion information does not exist, basic motion information (or default motion information) can be derived. The basic motion information can be used as the motion information for the sub-blocks of the current block. The basic motion information can be derived from the block located at the center of the col block (i.e., the col CU including the col sub-blocks). For example, motion information (e.g., motion vector) can be derived from the block including the bottom-right sample among the four samples located at the center of the col block, and can be used as the basic motion information.

[0169] As described above, in the case of the affine mode or sbTMVP mode where motion information is derived in sub-block units, affine merge candidates and sbTMVP candidates can be derived, and a merge candidate list based on sub-blocks can be configured based on these candidates. In this case, flag information indicating whether the affine mode or sbTMVP mode is enabled or disabled can be signaled. If the sbTMVP mode is enabled based on the flag information, the sbTMVP candidates derived as described above can be added to the first sorting of the merge candidate list based on sub-blocks. In addition, the affine merge candidates can be added to the next entry of the merge candidate list based on sub-blocks. In this case, the maximum number of candidates in the merge candidate list based on sub-blocks can be 5.

[0170] In addition, in the case of the sbTMVP mode, the sub-block size can be fixed and can be fixed to, for example, 8×8 size. In addition, the sbTMVP mode can be applied only to blocks having both a width and a height equal to or greater than 8.

[0171] Meanwhile, in the current VVC standard, as shown in Table 1, sub-block-based temporal motion information candidates (sbTMVP candidates) can be derived.

[0172] [Table 1]

[0173]

[0174]

[0175]

[0176]

[0177] When deriving sbTMVP candidates according to the method illustrated in Table 1, the default MV and sub-block MVs can be considered. In this case, the default MV can be referred to as sub-block-based temporal merged basic motion data or basic motion vectors (basic motion information). Referring to Table 1, the default MV can correspond to the ctrMV (or ctrMVLX) in Table 1. The sub-block MV can correspond to the mvSbCol (or mvLXSbcol) in Table 1.

[0178] For example, if a sub-block or sub-block MV is available according to the sbTMVP derivation process, the sub-block MV can be assigned to the corresponding sub-block, or if the sub-block or sub-block MV is not available, the default MV can be used as the corresponding sub-block MV for the corresponding sub-block. In this case, the default MV can derive motion information from a position corresponding to the center pixel position of the corresponding block (i.e., col CU) on the col picture, and each sub-block MV can derive motion information from the upper left position of the corresponding sub-block (i.e., col sub-block) on the col picture. In this case, the corresponding block (i.e., col CU) can be derived from the motion shift position based on the motion vector of the spatially adjacent block A1 (i.e., the temporal MV), as described above in Figure 11 as described.

[0179] Figures 12 to 15 is a diagram for schematically describing a method for calculating corresponding positions for deriving the default MV and sub-block MVs based on the block size during the sbTMVP derivation process.

[0180] Figures 12 to 15 The pixels (samples) outlined by the dashed line in

[0181] indicate the corresponding positions of each sub-block for deriving each sub-block MV, and the pixels (samples) outlined by the solid line illustrate the corresponding positions of the CU for deriving the default MV. Figure 12 For example, referring to

[0182] , if the current block (i.e., the current CU) has a size of 8×8, the motion information of the sub-block can be derived based on the upper left sample position within the sub-block having a size of 8×8, and the default motion information of the sub-block can be derived based on the center sample position within the current block (i.e., the current CU) having a size of 8×8. Figure 13 Alternatively, for example, referring to

[0183] Alternatively, for example, referring to Figure 14 , if the current block (i.e., the current CU) has a size of 8×16, the motion information of each sub-block can be derived based on the top-left sample position within each sub-block having a size of 8×8, and the default motion information of each sub-block can be derived based on the center sample position within the current block (i.e., the current CU) having a size of 8×16.

[0184] Alternatively, for example, referring to Figure 15 , if the current block (i.e., the current CU) has a size of 16×16, the motion information of each sub-block can be derived based on the top-left sample position within each sub-block having a size of 8×8, and the default motion information of each sub-block can be derived based on the center sample position within the current block (i.e., the current CU) having a size of 16×16.

[0185] From Figures 12 to 15 it can be seen that since the motion information of the sub-blocks tends to the top-left pixel position, there is the following problem: the sub-block MVs are derived at positions far from the position where the default MV indicating the representative motion information of the current CU is derived. As an example, in the case of the 8×8 block shown in Figure 12 , one CU includes one sub-block, but there is a contradiction where the sub-block MV and the default MV are represented as different motion information. In addition, since the methods for calculating the corresponding positions of the sub-block and the current CU block are different (i.e., the corresponding position for deriving the MV of the sub-block is the top-left sample position, while the corresponding position for deriving the default MV is the center sample position), additional modules may be required in hardware (H / W) implementation.

[0186] Therefore, to improve the problem, this document proposes a solution for unifying the method of deriving the corresponding position of the CU for the default MV and the method of deriving the corresponding position of the sub-block for each sub-block MV during the process of deriving sbTMVP candidates. According to the embodiments of this document, there is a unifying effect, that is, from the perspective of hardware (H / W), only one module for deriving each corresponding position based on the block size can be used. For example, since the method of calculating the corresponding position if the block size is 16×16 and the method of calculating the corresponding position if the block size is 8×8 can be implemented in the same way, there is a simplification effect in terms of hardware implementation. In this case, the 16×16 block can represent the CU, and the 8×8 block can represent each sub-block.

[0187] As an embodiment, when deriving sbTMVP candidates, the center sample position can be used as the corresponding position for deriving the motion information of the sub-block and the corresponding position for deriving the default motion information, and it can be implemented as shown in Table 2 below.

[0188] Table 2 below is an illustration of an example of a method for deriving motion information and default motion information of a derived sub-block according to an embodiment of this document.

[0189] [Table 2]

[0190]

[0191]

[0192]

[0193]

[0194] Referring to Table 2, when deriving sbTMVP candidates, the position of the current block (i.e., the current CU) including the sub-block can be derived. The upper-left sample position (xCtb, yCtb) of the coding tree block (or coding tree unit) including the current block and the lower-right center sample position (xCtr, yCtr) of the current block can be derived as in equations (8-514) to (8-517) in Table 2. In this case, the positions (xCtb, yCtb) and (xCtr, yCtr) can be calculated based on the upper-left sample position (xCb, yCb) of the current block relative to the upper-left sample of the current picture.

[0195] In addition, the col block (i.e., col CU) on the col picture corresponding to the positioning of the current block (i.e., the current CU) including the sub-block can be derived. In this case, the position of the col block can be set to (xColCtrCb, yColCtrCb). The position can represent the position of the col block, which includes the position (xCtr, yCtr) relative to the upper-left sample of the col picture within the col picture.

[0196] In addition, the basic motion data (i.e., default motion information) for sbTMVP can be exported. The basic motion data may include a default MV (e.g., CtrMvLX). For example, col blocks on a col picture can be exported. In this case, the position of a col block can be exported as (xColCb, yColCb). This position can be the position where a motion shift (e.g., tempMv) has been applied to the exported col block position (xColCtrCb, yColCtrCb). As described above, the motion shift can be performed by adding the motion vector (e.g., tempMv) derived from a spatially neighboring block of the current block (e.g., A1 block) to the current col block position (xColCtrCb, yColCtrCb). Next, a default MV (e.g., ctrMvLX) can be derived based on the position (xColCb, yColCb) of the motion-shifted col block. In this case, the default MV (e.g., ctrMvLX) can represent the motion vector derived from the position corresponding to the lower-right center sample of the col block.

[0197] In addition, col sub-blocks on a col picture corresponding to sub-blocks (represented as current sub-blocks) in the current block can be exported. First, the position of each current sub-block can be exported. The position of each in the sub-block can be represented as (xSb, ySb). The position (xSb, ySb) can represent the position of the current sub-block based on the top-left sample of the current picture. For example, the position (xSb, ySb) of the current sub-block can be calculated as in equations (8-523) to (8-524) of Table 2, which can represent the lower-right center sample position of the sub-block. Next, the position of each col sub-block in the col sub-blocks on the col picture can be exported. The position of each col sub-block can be represented as (xColSb, yColSb). The position (xColSb, yColSb) can be the position where a motion shift (e.g., tempMv) is applied to the position (xSb, ySb) of the current sub-block. As described above, the motion shift can be performed by adding the motion vector (e.g., tempMv) derived from a spatially neighboring block of the current block (e.g., A1 block) to the current sub-block position (xSb, ySb). Next, the motion information (e.g., motion vector mvLXSbCol, flag availableFlagLXSbCol indicating availability) of the col sub-blocks can be derived based on the position (xColSb, yColSb) of each of the motion-shifted col sub-blocks.

[0198] In this case, if the col sub-blocks are unavailable in the col sub-blocks (e.g., when availableFlagLXSbCol is 0), the basic motion data (i.e., the default motion information) can be used for the unavailable col sub-blocks. For example, the default MV (e.g., ctrMvLX) can be used as the motion vector (e.g., mvLXSbCol) for the unavailable col sub-blocks.

[0199] Figures 16 to 19 is an example diagram for schematically describing a method for uniformly deriving the corresponding positions of the default MV and the sub-block MV based on the block size during the sbTMVP derivation process.

[0200] Figures 16 to 19 The pixels (samples) outlined by the dashed line in the figure indicate the corresponding positions within each sub-block for deriving each sub-block MV, and the pixels (samples) outlined by the solid line illustrate the corresponding positions of the CU for deriving the default MV.

[0201] For example, referring to Figure 16 , if the current block (i.e., the current CU) has a size of 8×8, the motion information can be derived from the col sub-block at the corresponding position on the col picture based on the lower-right center sample position within the sub-block with a size of 8×8, and can be used as the motion information of the current sub-block. The motion information can be derived from the col block (i.e., the col CU) at the corresponding position on the col picture based on the lower-right center sample position within the current block (i.e., the current CU) with a size of 8×8, and can be used as the default motion information of the current sub-block. In this case, as Figure 16 shown, the motion information and the default motion information of the current sub-block can be derived from the same sample position (the same corresponding position).

[0202] Alternatively, for example, referring to Figure 17 , if the current block (i.e., the current CU) has a size of 16×8, the motion information can be derived from the col sub-block at the corresponding position on the col picture based on the lower-right center sample position within the sub-block with a size of 8×8, and can be used as the motion information of the current sub-block. The motion information can be derived from the col block (i.e., the col CU) at the corresponding position on the col picture based on the lower-right center sample position within the current block (i.e., the current CU) with a size of 16×8, and can be used as the default motion information of the current sub-block.

[0203] Alternatively, for example, referring to Figure 18, if the current block (i.e., the current CU) has a size of 8×16, the motion information can be derived from the col sub-block at the corresponding position on the col picture based on the lower-right center sample position within the sub-block having a size of 8×8, and can be used as the motion information of the current sub-block. The motion information can be derived from the col block (i.e., col CU) at the corresponding position on the col picture based on the lower-right center sample position within the current block (i.e., the current CU) having a size of 8×16, and can be used as the default motion information of the current sub-block.

[0204] Alternatively, for example, refer to Figure 19 , if the current block (i.e., the current CU) has a size equal to or greater than 16×16, the motion information can be derived from the col sub-block at the corresponding position on the col picture based on the lower-right center sample position within the sub-block having a size of 8×8, and can be used as the motion information of the current sub-block. The motion information can be derived from the col block (i.e., col CU) at the corresponding position on the col picture based on the lower-right center sample position within the current block (i.e., the current CU) having a size of 16×16 (or a size of 16×16 or larger), and can be used as the default motion information of the current sub-block.

[0205] However, the foregoing embodiments of this document are merely examples, and the default motion information and the motion information of the current sub-block can be derived based on another sample position (i.e., the lower-right sample position) other than the center position. For example, the default motion information can be derived based on the upper-left sample position of the current CU, and the motion information of the current sub-block can be derived based on the upper-left sample position of the sub-block.

[0206] If the embodiments of this document are implemented in hardware, pipelines such as Figure 20 and Figure 21 can be configured because the same H / W module can be used to derive the motion information (temporal motion).

[0207] Figure 20 and Figure 21 are example diagrams schematically illustrating the configuration of the pipeline through which the corresponding positions for deriving the default MV and the sub-block MV can be unified and calculated during the sbTMVP derivation process.

[0208] Refer to Figure 20 and Figure 21 , the corresponding position calculation module can calculate the corresponding positions for deriving the default MV and the sub-block MV. For example, as Figure 20 and Figure 21As shown, when the position (posX, posY) and size (blkszX, blkszY) of a block are input to the corresponding position calculation module, the center position of the input block (i.e., the lower-right sample position) can be output. When the position and size of the current CU are input to the corresponding position calculation module, the center position of the col block on the col picture (i.e., the lower-right sample position), that is, the corresponding position for deriving the default MV, can be output. Alternatively, when the position and size of the current sub-block are input to the corresponding position calculation module, the corresponding position of the col sub-block on the col picture (i.e., the lower-right sample position) can be output, that is, the corresponding position for deriving the MV of the current sub-block.

[0209] As described above, when the corresponding positions for deriving the default MV and sub-block MV are output from the corresponding position calculation module, the motion vectors derived from the corresponding positions (i.e., temporal mv) can be patched. In addition, the sub-block based temporal motion information (i.e., sbTMVP candidate) can be derived based on the patched motion vectors (i.e., temporal mv). For example, as in Figure 20 and 21 , depending on the H / W implementation, the sbTMVP candidates can be derived in parallel based on clock cycles, or the sbTMVP candidates can be derived sequentially.

[0210] The following drawings are prepared to describe detailed examples of this document. The names of the detailed devices or the detailed terms or names written in the drawings (e.g., the name of the syntax / syntax name) are exemplary, and thus the technical features of this document are not limited to the detailed names used in the following drawings.

[0211] Figure 22 Schematically illustrates an example of a video / image encoding method according to an embodiment of this document.

[0212] Figure 22 The method disclosed in Figure 2 can be executed by the encoding device 200 disclosed in Figure 22 . Specifically, Figure 2 the steps S2200 to S2230 in Figure 22 can be executed by the predictor 220 (more specifically, the inter-frame predictor 221) disclosed in Figure 2 . Figure 22 The step S2240 in Figure 2 can be executed by the residual processor 230 disclosed in Figure 22 . The step S2250 in Figure 22 can be executed by the entropy encoder 240 disclosed in

[0213] Reference Figure 22 The encoding device can derive a reference sub-block within a collocated reference picture for a sub-block in the current block (S2200).

[0214] In this case, as described above, the collocated reference picture refers to the reference picture used to derive the temporal motion information (i.e., sbTMVP), and can represent the aforementioned col picture. The reference sub-block can represent the aforementioned col sub-block.

[0215] As an example, the encoding device can derive a reference sub-block within the collocated reference picture based on the position of the sub-block in the current block. In this case, the current block can be represented as the current coding unit (CU) or the current coding block (CB). The sub-blocks included in the current block can also be represented as the current coding sub-blocks.

[0216] For example, the encoding device can first specify the position of the current block, and then specify the position of the sub-block in the current block. As described in reference table 2, the position of the current block can be represented based on the top-left sample position (xCtb, yCtb) of the coding tree block and the bottom-right center sample position (xCtr, yCtr) of the current block. Each of the positions of the sub-blocks in the current block can be represented as (xSb, ySb). The position (xSb, ySb) can represent the bottom-right center sample position of the sub-block. In this case, the bottom-right center sample position (xSb, ySb) of the sub-block can be calculated based on the top-left sample position of the sub-block and the sub-block size, and can be calculated similar to equations (8-523) to (8-524) in table 2.

[0217] In addition, the encoding device can derive a reference sub-block within the collocated reference picture based on the bottom-right center sample position of each sub-block in the current block. As described in reference table 2, the reference sub-block can be represented as the position (xColSb, yColSb) within the collocated reference picture. The position (xColSb, yColSb) can be derived from the collocated reference picture based on the bottom-right center sample position (xSb, ySb) of each of the sub-blocks in the current block.

[0218] Meanwhile, the top-left sample position used in this document can also be represented as the upper-left sample position, the left-above sample position, etc. The bottom-right center sample position can also be represented as the bottom-right center sample position, the center below-right sample position, the lower-right center sample position, the center bottom-right sample position, etc.

[0219] In addition, when exporting a reference sub-block, motion shifting can be applied. The encoding device can perform motion shifting based on a motion vector derived from a spatial neighboring block of the current block. The spatial neighboring block of the current block can be a left neighboring block located on the left side of the current block, and can represent, for example, Figure 10 and Figure 11 the A1 block shown. In this case, if the left neighboring block (e.g., the A1 block) is available, a motion vector can be derived from the left neighboring block, or if the left neighboring block is not available, a zero vector can be derived. In this case, it can be determined whether the spatial neighboring block is available based on the reference picture information, prediction mode information, position information, etc. of the spatial neighboring block. For example, if the reference picture of the spatial neighboring block is the same as that of the current block, the corresponding spatial neighboring block can be determined to be available. Alternatively, if the spatial neighboring block is coded in an intra prediction mode or the spatial neighboring block is outside the current picture / tile, the corresponding spatial neighboring block can be determined to be unavailable.

[0220] That is, the encoding device can apply motion shifting (i.e., the motion vector of the spatial neighboring block (e.g., the A1 block)) to the lower right center sample position (xSb, ySb) of each sub-block in the current block, and can derive a reference sub-block within the collocated reference picture based on the motion shifting position. In this case, the position (xColSb, yColSb) of the reference sub-block can be expressed as the position obtained by motion shifting from the lower right center sample position (xSb, ySb) of each sub-block in the current block to the position indicated by the motion vector of the spatial neighboring block (e.g., the A1 block), and can be calculated similar to equations (8-525) to (8-526) in Table 2.

[0221] The encoding device can derive sub-block based temporal motion information candidates (S2210) based on the reference sub-block.

[0222] Meanwhile, in this document, the sub-block based temporal motion information candidates represent the aforementioned sub-block based temporal motion vector prediction (sbTMVP) candidates, and are replaced or used interchangeably by sub-block based temporal motion vector prediction sub-candidates. That is, as described above, if prediction is performed by deriving motion information in sub-block units, sbTMVP candidates can be derived. Motion prediction can be performed at the sub-block level (or sub-coding unit (sub-CU) level) based on the sbTMVP candidates.

[0223] The encoding device can derive motion information for the sub-blocks in the current block (S2220) based on the sub-block based temporal motion information candidates.

[0224] The sub-block based temporal motion information candidate may include a sub-block unit motion vector. In this case, the sub-block unit motion vector may include a motion vector derived based on a reference sub-block.

[0225] As an example, the encoding device may derive the sub-block unit motion vector for the reference sub-block as the motion information of the sub-blocks in the current block. For example, the encoding device may determine whether to derive the sub-block unit motion vector based on the availability of the reference sub-block. Regarding the available reference sub-blocks among the reference sub-blocks, the encoding device may derive the sub-block unit motion vector of the available reference sub-blocks based on the motion vectors of the available reference sub-blocks. Regarding the unavailable reference sub-blocks among the reference sub-blocks, the encoding device may use the base motion vector as the sub-block unit motion vector for the unavailable reference sub-blocks.

[0226] The base motion vector may correspond to the aforementioned default motion vector and may be derived from the collocated reference picture based on the position of the current block.

[0227] As an example, the encoding device may specify the position of the reference compilation block inside the collocated reference picture based on the lower right center sample position of the current block, and may derive the base motion vector based on the position of the reference compilation block. The reference compilation block may represent a col block inside the collocated reference picture corresponding to the current block including the sub-blocks. As described in reference table 2, the position of the reference compilation block may be represented as (xColCtrCb, yColCtrCb), and the position (xColCtrCb, yColCtrCb) may represent the position of the reference compilation block that covers the position (xCtr, yCtr) inside the collocated reference picture based on the upper left sample of the collocated reference picture. The position (xCtr, yCtr) may represent the lower right center sample position of the current block.

[0228] In addition, when deriving the base motion vector, a motion shift may be applied to the position (xColCtrCb, yColCtrCb) of the reference compilation block. As described above, the motion shift may be performed by adding the motion vector derived from the spatially neighboring block of the current block (e.g., A1 block) to the position (xColCtrCb, yColCtrCb) of the reference compilation block that covers the lower right center sample. The encoding device may derive the base motion vector based on the motion shifted position (xColCb, yColCb) of the reference compilation block. That is, the base motion vector may be a motion vector derived from the motion shifted position inside the collocated reference picture based on the lower right center sample position of the current block.

[0229] Meanwhile, it is possible to determine whether a reference sub-block is available based on whether the reference sub-block is outside the collocated reference picture or the motion vector. For example, an unavailable reference sub-block may include a reference sub-block detached from outside the collocated reference picture or a reference sub-block whose motion vector is unavailable. For example, if the reference sub-block is based on an intra mode, an intra block copy (IBC) mode, or a palette mode, the reference sub-block may be a sub-block whose motion vector is unavailable. Alternatively, if the reference compiled block covering the modified position derived based on the position of the reference sub-block is based on an intra mode, an IBC mode, or a palette mode, the reference sub-block may be a sub-block whose motion vector is unavailable.

[0230] In this case, as an example, the motion vector of the available reference sub-block can be derived based on the motion vector of the block covering the modified position derived based on the upper left sample position of the reference sub-block. For example, as shown in Table 2, the modified position can be derived based on an equation such as ((xColSb >> 3) << 3, (yColSb >> 3) << 3). In this case, xColSb and yColSb may respectively indicate the x coordinate and y coordinate of the upper left sample position of the reference sub-block, >> may indicate arithmetic right shift, and << may indicate arithmetic left shift.

[0231] Meanwhile, as described above, when deriving candidates for the temporal motion information based on sub-blocks, it can be seen that the motion vector of the reference sub-block is derived based on the position of the sub-block in the current block, and the basic motion vector is derived based on the position of the current block. For example, as referred to Figures 16 to 19 above, for a current block having a size of 8×8, the motion vector of the reference sub-block and the basic motion vector can be derived based on the lower right center sample position of the current block. For a current block having a size larger than 8×8, the motion vector of the reference sub-block can be derived based on the lower right center sample position of each sub-block in the current block, and the basic motion vector can be derived based on the lower right center sample position of the current block.

[0232] The encoding device can generate prediction samples for the current block based on the motion information of the sub-blocks in the current block (S2230).

[0233] The encoding device can select the best motion information based on the rate distortion (RD) cost, and can generate prediction samples based on the best motion information. For example, if the motion information derived in sub-block units with respect to the current block (i.e., sbTMVP) is selected as the best motion information, the encoding device can generate prediction samples for the current block based on the motion information of the sub-blocks of the current block derived as described above.

[0234] The encoding device can derive residual samples based on the prediction samples (S2240), and can encode the image information including information about the residual samples (S2250).

[0235] That is, the encoding device can derive residual samples based on the initial samples of the current block and the predicted samples of the current block. In addition, the encoding device can generate information about the residual samples. In this case, the information about the residual samples can include information such as the value information, position information, transform scheme, transform kernel, and quantization parameters of the quantized transform coefficients derived by performing a transform and quantization on the residual samples.

[0236] The encoding device can encode the information about the residual samples, output the information as a bitstream, and send the bitstream to the decoding device via a network or a storage medium.

[0237] Figure 23 An example of a video / image decoding method according to an embodiment of this document is schematically shown.

[0238] Figure 23 The method disclosed in Figure 3 can be executed by the decoding device 300 disclosed in Figure 23 Specifically, Figure 3 the steps S2300 to S2330 in Figure 23 can be executed by the predictor 330 (more specifically, the inter-frame predictor 332) disclosed in Figure 3 Figure 23 The step S2340 in can be executed by the adder 340 disclosed in Figure 23 In addition, the method disclosed in

[0239] including the embodiments described in this document can be executed. Therefore, the detailed description of the content redundant with the embodiments described with reference to Figure 23 is omitted or simply given.

[0240] Referring to

[0241]

[0242] ​For example, the decoding device may first specify the position of the current block and then specify the position of the sub-blocks within the current block. As described in Reference Table 2, the position of the current block can be represented based on the top-left sample position (xCtb, yCtb) of the compiled tree block and the bottom-right center sample position (xCtr, yCtr) of the current block. Each of the positions of the sub-blocks within the current block can be represented as (xSb, ySb). The position (xSb, ySb) can represent the bottom-right center sample position of the sub-block. In this case, the bottom-right center sample position (xSb, ySb) of the sub-block can be calculated based on the top-left sample position of the sub-block and the sub-block size, and can be calculated similar to equations (8-523) to (8-524) in Table 2.

[0243] In addition, the decoding device can derive a reference sub-block within the collocated reference picture based on the bottom-right center sample position of each sub-block within the current block. As described in Reference Table 2, the reference sub-block can be represented as a position (xColSb, yColSb) within the collocated reference picture. The position (xColSb, yColSb) can be derived from the collocated reference picture based on the bottom-right center sample position (xSb, ySb) of each of the sub-blocks within the current block.

[0244] Meanwhile, the top-left sample position used in this document can also be represented as the upper-left sample position, the left-above sample position, etc. The bottom-right center sample position can also be represented as the bottom-right center sample position, the center below-right sample position, the lower-right center sample position, the center bottom-right sample position, etc.

[0245] In addition, when deriving the reference sub-block, motion shift can be applied. The decoding device can perform motion shift based on the motion vector derived from the spatially neighboring block of the current block. The spatially neighboring block of the current block can be the left neighboring block located on the left side of the current block, and can be represented, for example Figure 10 and Figure 11Block A1 as shown. In this case, if a left neighboring block (e.g., Block A1) is available, a motion vector can be derived from the left neighboring block, or if the left neighboring block is not available, a zero vector can be derived. In this case, it is possible to determine whether a spatial neighboring block is available based on reference picture information, prediction mode information, position information, etc. of the spatial neighboring block. For example, if the reference picture of the spatial neighboring block is the same as that of the current block, the corresponding spatial neighboring block can be determined to be available. Alternatively, if the spatial neighboring block is coded in an intra prediction mode or the spatial neighboring block is outside the current picture / tile, the corresponding spatial neighboring block can be determined to be unavailable.

[0246] That is to say, the decoding device can apply a motion shift (i.e., the motion vector of a spatial neighboring block (e.g., Block A1)) to the lower right center sample position (xSb, ySb) of each sub-block in the current block, and can derive a reference sub-block within the collocated reference picture based on the motion shift position. In this case, the position (xColSb, yColSb) of the reference sub-block can be expressed as the position obtained by motion-shifting from the lower right center sample position (xSb, ySb) of each sub-block in the current block to the position indicated by the motion vector of the spatial neighboring block (e.g., Block A1), and can be calculated similar to equations (8-525) to (8-526) in Table 2.

[0247] The decoding device can derive sub-block-based temporal motion information candidates based on the reference sub-block (S2310).

[0248] Meanwhile, in this document, the sub-block-based temporal motion information candidates represent the aforementioned sub-block-based temporal motion vector prediction (sbTMVP) candidates, and are replaced by or interchangeably used with sub-block-based temporal motion vector prediction sub-candidates. That is, as described above, if prediction is performed by deriving motion information in sub-block units, sbTMVP candidates can be derived. Motion prediction can be performed at the sub-block level (or sub-coding unit (sub-CU) level) based on the sbTMVP candidates.

[0249] The decoding device can derive motion information for the sub-blocks in the current block based on the sub-block-based temporal motion information candidates (S2320).

[0250] The sub-block-based temporal motion information candidates can include sub-block unit motion vectors. In this case, the sub-block unit motion vectors can include motion vectors derived based on the reference sub-block.

[0251] As an example, the decoding device may derive the sub-block unit motion vector for the reference sub-block as the motion information of the sub-blocks in the current block. For example, the decoding device may derive the sub-block unit motion vector based on whether the reference sub-block is available. Regarding the available reference sub-blocks among the reference sub-blocks, the decoding device may derive the sub-block unit motion vector of the available reference sub-blocks based on the motion vectors of the available reference sub-blocks. Regarding the unavailable reference sub-blocks among the reference sub-blocks, the decoding device may use the basic motion vector as the sub-block unit motion vector for the unavailable reference sub-blocks.

[0252] The basic motion vector may correspond to the aforementioned default motion vector and may be derived from the collocated reference picture based on the position of the current block.

[0253] As an example, the decoding device may specify the position of the reference compilation block inside the collocated reference picture based on the lower right center sample position of the current block, and may derive the basic motion vector based on the position of the reference compilation block. The reference compilation block may represent the col block inside the collocated reference picture corresponding to the current block including the sub-blocks. As described in Reference Table 2, the position of the reference compilation block may be represented as (xColCtrCb, yColCtrCb), and the position (xColCtrCb, yColCtrCb) may represent the position of the reference compilation block that covers the position (xCtr, yCtr) inside the collocated reference picture based on the upper left sample of the collocated reference picture. The position (xCtr, yCtr) may represent the lower right center sample position of the current block.

[0254] In addition, when deriving the basic motion vector, a motion shift may be applied to the position (xColCtrCb, yColCtrCb) of the reference compilation block. As described above, the motion shift may be performed by adding the motion vector derived from the spatially neighboring block (e.g., A1 block) of the current block to the position (xColCtrCb, yColCtrCb) of the reference compilation block that covers the lower right center sample. The decoding device may derive the basic motion vector based on the motion-shifted position (xColCb, yColCb) of the reference compilation block. That is, the basic motion vector may be the motion vector derived from the motion-shifted position inside the collocated reference picture based on the lower right center sample position of the current block.

[0255] Meanwhile, it is possible to determine whether a reference sub-block is available based on whether the reference sub-block is located outside the collocated reference picture or the motion vector. For example, an unavailable reference sub-block may include a reference sub-block that has detached from outside the collocated reference picture or a reference sub-block whose motion vector is unavailable. For example, if the reference sub-block is based on an intra mode, an intra block copy (IBC) mode, or a palette mode, the reference sub-block may be a sub-block whose motion vector is unavailable. Alternatively, if a reference compilation block that covers a modified position derived based on the position of the reference sub-block is based on an intra mode, an IBC mode, or a palette mode, the reference sub-block may be a sub-block whose motion vector is unavailable.

[0256] In this case, as an example, the motion vector of an available reference sub-block can be derived based on the motion vector of a block that covers a modified position derived based on the top-left sample position of the reference sub-block. For example, as shown in Table 2, the modified position can be derived based on an equation such as ((xColSb>>3)<<3,(yColSb>>3)<<3). In this case, xColSb and yColSb may respectively indicate the x coordinate and y coordinate of the top-left sample position of the reference sub-block, >> may indicate arithmetic right shift, and << may indicate arithmetic left shift.

[0257] Meanwhile, as described above, when deriving candidates for sub-block-based temporal motion information, it can be seen that the motion vector of a reference sub-block is derived based on the position of a sub-block in a current block, and the basic motion vector is derived based on the position of the current block. For example, as described Figures 16 to 19 above, for a current block having a size of 8×8, the motion vector of the reference sub-block and the basic motion vector can be derived based on the bottom-right center sample position of the current block. For a current block having a size larger than 8×8, the motion vector of the reference sub-block can be derived based on the bottom-right center sample position of each sub-block in the current block, and the basic motion vector can be derived based on the bottom-right center sample position of the current block.

[0258] The decoding device may generate prediction samples for the current block based on the motion information of the sub-blocks in the current block (S2330).

[0259] As an example, for a prediction mode (i.e., the sbTMVP mode) that performs prediction on the current block based on sub-block unit motion information, the decoding device may generate prediction samples for the current block based on the motion information of the sub-blocks of the current block derived as above.

[0260] The decoding device may generate reconstructed samples based on the prediction samples (S2340).

[0261] As an example, the decoding device may use the prediction samples as the reconstructed samples depending on the prediction mode, or may generate the reconstructed samples by adding the residual samples to the prediction samples.

[0262] If there are residual samples for the current block, the decoding device may receive information regarding the residual of the current block. The information regarding the residual may include transform coefficients of the residual samples. The decoding device may derive the residual samples (or an array of residual samples) for the residual information of the current block. The decoding device may generate reconstructed samples based on the predicted samples and the residual samples, and may derive a reconstructed block or a reconstructed picture based on the reconstructed samples. Thereafter, as described above, the decoding device may apply in-loop filtering processes such as deblocking filtering and / or SAO process to the reconstructed picture in order to improve subjective / objective picture quality if necessary.

[0263] In the above-mentioned embodiments, although these methods have been described based on flowcharts in the form of a series of steps or units, the embodiments of this document are not limited to the order of these steps, and some of these steps may be executed in an order different from the order of other steps or may be executed simultaneously with other steps. In addition, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and without affecting the scope of the rights of this document, these steps may include additional steps or one or more steps in the flowchart may be deleted.

[0264] The above-mentioned methods according to this document may be implemented in software form, and the encoding device and / or decoding device according to this document may be included in a device for performing image processing such as a TV, a computer, a smart phone, a set-top box, or a display device.

[0265] In this document, when the embodiments are implemented in software form, the above-mentioned methods may be implemented as modules (programs, functions, etc.) for performing the above-mentioned functions. The modules may be stored in a memory and executed by a processor. The memory may be arranged inside or outside the processor and is connected to the processor by various well-known means. The processor may include an application specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document may be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units illustrated in the drawings may be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, the information (e.g., information regarding instructions) or algorithms for such implementation may be stored in a digital storage medium.

[0266] In addition, the decoding device and encoding device applying this document may be included in a multimedia broadcast transmission and reception device, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a video camera, a video-on-demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, a video phone video device, a vehicle terminal (e.g., a vehicle (including an autonomous vehicle) terminal, an aircraft terminal, and a ship terminal), and a medical video device, and may be used to process video signals or data signals. For example, an over-the-top (OTT) video device may include a game console, a Blueray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, and a digital video recorder (DVR).

[0267] In addition, the processing method applying this document may be generated in the form of a program executable by a computer and may be stored in a computer-readable recording medium. Multimedia data having a data structure according to this document may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices storing computer-readable data. The computer-readable recording medium may include, for example, a Blueray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmitted through the Internet). In addition, a bitstream generated using an encoding method may be stored in a computer-readable recording medium or may be transmitted through wired and wireless communication networks.

[0268] In addition, the embodiments of this document may be implemented as a computer program product using program code. The program code may be executed by a computer according to the embodiments of this document. The program code may be stored on a carrier wave readable by a computer.

[0269] Figure 24 The figure shows an example of a content streaming system to which the embodiments disclosed in this document may be applied.

[0270] Reference Figure 24 , the content streaming media system applying the embodiments of the present invention may basically include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0271] The encoding server compresses the content input from multimedia input devices such as smart phones, cameras, video cameras, etc. into digital data to generate a bitstream, and transmits the bitstream to the streaming server. As another example, when a multimedia input device such as a smart phone, camera, video camera, etc. directly generates a bitstream, the encoding server can be omitted.

[0272] The bitstream can be generated by applying the encoding method or the bitstream generation method of the embodiments of this document, and the streaming server can temporarily store the bitstream in the process of sending or receiving the bitstream.

[0273] The streaming server sends the multimedia data to the user device through the web server based on the user's request, and the web server serves as a medium for informing the user of the service. When the user requests a desired service from the web server, the web server forwards it to the streaming server, and the streaming server sends the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between the devices in the content streaming system.

[0274] The streaming server can receive the content from the media storage device and / or the encoding server. For example, when receiving the content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream within a predetermined time.

[0275] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, touchscreen PCs, tablet PCs, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc.

[0276] Each server in the content streaming system can operate as a distributed server, in which case the data received from each server can be distributed.

[0277] The claims described herein can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and implemented as a device, and the technical features of the device claims in this specification can be combined and implemented as a method. In addition, the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented as a device, and the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented as a method.

Claims

1. An image decoding method performed by a decoding device, the method comprising: Deriving a reference sub-block within a collocated reference picture for a sub-block in a current block; Deriving sub-block-based temporal motion information candidates based on the reference sub-block; Deriving motion information for the sub-block in the current block based on the sub-block-based temporal motion information candidates; Generating a prediction sample for the current block based on the motion information for the sub-block in the current block; And Generating a reconstructed sample based on the prediction sample, wherein the reference sub-block is derived from the collocated reference picture by applying a motion shift to the lower-right center sample position of each of the sub-blocks in the current block, wherein the motion shift is performed based on a motion vector derived from a spatially neighboring block of the current block, wherein the sub-block-based temporal motion information candidates include sub-block unit motion vectors, and wherein a base motion vector is used as the sub-block unit motion vector for an unavailable reference sub-block among the reference sub-blocks.

2. The image decoding method according to claim 1, Among them, Deriving the base motion vector from the collocated reference picture based on the lower-right center sample position of the current block.

3. The image decoding method according to claim 1, Among them, Deriving a sub-block unit motion vector for an available reference sub-block among the reference sub-blocks based on the motion vector of the available reference sub-block, wherein the motion vector of the available reference sub-block is derived based on the motion vector of a block covering a modified position, and the modified position is derived based on the upper-left sample position of the reference sub-block, wherein the modified position is derived based on the following formula, and ((xColSb>>3)<<3,(yColSb>>3)<<3) wherein xColSb and yColSb respectively indicate the x coordinate and y coordinate of the upper-left sample position of the reference sub-block, >> indicates arithmetic right shift, and << indicates arithmetic left shift.

4. The image decoding method according to claim 1, Among them, The unavailable reference sub-block includes a reference sub-block located outside the collocated reference picture or a reference sub-block whose motion vector is unavailable.

5. The image decoding method according to claim 1, Among them, For a current block having a size of 8x8, deriving a motion vector for the reference sub-block and the base motion vector based on the lower-right center sample position of the current block.

6. The image decoding method according to claim 1, Among them, For a current block having a size larger than 8x8, deriving a motion vector for the reference sub-block based on the lower-right center sample position of each of the sub-blocks in the current block, and deriving the base motion vector based on the lower-right center sample position of the current block.

7. The image decoding method according to claim 1, Among them, Calculating the lower-right center sample position of each of the sub-blocks in the current block based on the upper-left sample position of each of the sub-blocks in the current block and the size of each of the sub-blocks in the current block.

8. An image encoding method performed by an encoding device, comprising: Derive a reference sub-block inside the juxtaposed reference picture for sub-blocks in the current block; Based on the reference sub-block, derive a sub-block-based temporal motion information candidate; Based on the sub-block-based temporal motion information candidate, derive motion information for the sub-blocks in the current block; Based on the motion information for the sub-blocks in the current block, generate a prediction sample for the current block; Derive a residual sample based on the prediction sample; And Encode image information including information about the residual sample, wherein the reference sub-block is derived from the juxtaposed reference picture by applying a motion shift to the lower-right center sample position of each of the sub-blocks in the current block, wherein the motion shift is performed based on a motion vector derived from spatially neighboring blocks of the current block, wherein the sub-block-based temporal motion information candidate includes a sub-block unit motion vector, and wherein a basic motion vector is used as the sub-block unit motion vector for an unavailable reference sub-block among the reference sub-blocks.

9. The image encoding method according to claim 8, Among them, Derive the basic motion vector from the juxtaposed reference picture based on the lower-right center sample position of the current block.

10. The image encoding method according to claim 8, Among them, Derive a sub-block unit motion vector for an available reference sub-block among the reference sub-blocks based on the motion vector of the available reference sub-block, wherein the motion vector of the available reference sub-block is derived based on the motion vector of a block covering a modified position, and the modified position is derived based on the upper-left sample position of the reference sub-block, wherein the modified position is derived based on the following formula, and ((xColSb>>3)<<3,(yColSb>>3)<<3) where xColSb and yColSb respectively indicate the x coordinate and y coordinate of the upper-left sample position of the reference sub-block, >> indicates arithmetic right shift, and << indicates arithmetic left shift.

11. The image encoding method according to claim 8, Among them, The unavailable reference sub-block includes a reference sub-block located outside the juxtaposed reference picture or a reference sub-block whose motion vector is unavailable.

12. The image encoding method according to claim 8, Among them, For a current block having a size of 8x8, derive a motion vector for the reference sub-block and the basic motion vector based on the lower-right center sample position of the current block.

13. The image encoding method according to claim 8, Among them, For the current block having a size larger than 8x8, derive a motion vector for the reference sub-block based on the lower-right center sample position of each of the sub-blocks in the current block, and derive the basic motion vector based on the lower-right center sample position of the current block.

14. The image encoding method according to claim 8, Among them, Calculate the lower-right center sample position of each of the sub-blocks in the current block based on the upper-left sample position of each of the sub-blocks in the current block and the size of each of the sub-blocks in the current block.

15. A method for transmitting data of image information, the method comprising: Obtain a bitstream of picture information including information on residual samples, wherein the bitstream is generated based on the following operations: derive a reference sub-block within a collocated reference picture for a sub-block in a current block; based on the reference sub-block, derive a sub-block-based temporal motion information candidate; based on the sub-block-based temporal motion information candidate, derive motion information for the sub-block in the current block; based on the motion information for the sub-block in the current block, generate a prediction sample for the current block; based on the prediction sample, derive a residual sample; encode the picture information including information on the residual sample; and transmit data of the bitstream including the picture information, the picture information including information on the residual sample, wherein the reference sub-block is derived from the collocated reference picture by applying a motion shift to the lower-right center sample position of each of the sub-blocks in the current block, wherein the motion shift is performed based on a motion vector derived from a spatially neighboring block of the current block, wherein the sub-block-based temporal motion information candidate includes a sub-block unit motion vector, and wherein a base motion vector is used as the sub-block unit motion vector for an unavailable reference sub-block among the reference sub-blocks.

Citation Information

Patent Citations

  • Video coding device and video decoding device

    WO2019078169A1

  • KR20190053238A