MDMVR-based video coding method and apparatus
The MDMVR-based video coding method addresses the need for efficient compression of high-resolution and immersive media by refining motion vectors and improving inter-coding efficiency, thus reducing data transmission and storage costs.
Patent Information
- Application Number
- JP2024519056
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-10-01
- Filing Date
- 2022-09-30
- Publication Date
- 2025-10-15
- Estimated Expiration
- 2042-09-30
AI Technical Summary
The increasing demand for high-resolution, high-quality video and immersive media requires a highly efficient video compression technology to manage the increased data transmission and storage costs associated with higher resolution and quality, as well as the diverse visual characteristics of content like VR and AR.
A multi-layer Decoder-side Motion Vector Refinement (MDMVR) based video coding method and apparatus that determines the use of MDMVR for current blocks, derives refined motion vectors, and generates reconstructed samples to improve inter-coding efficiency without increasing processing complexity.
This approach enhances overall image/video compression efficiency, supports decoder-side motion vector refinement using multiple layers, and improves the inter-coding structure performance without increasing processing complexity.
Smart Images

Figure 0007755059000001 
Figure 0007755059000002 
Figure 0007755059000003
Abstract
Description
[Technical Field]
[0001] This document relates to video or image coding techniques, for example, multi-layer decoder-side motion vector refinement (MDMVR) based video coding techniques. [Background technology]
[0002] Recently, the demand for high-resolution, high-quality video / images, such as 4K or 8K or higher UHD (Ultra High Definition) video / images, is increasing in various fields. As the video / image data has higher resolution and quality, the amount of information or bits to be transmitted increases relatively compared to existing video / image data, which increases the transmission and storage costs when transmitting video / image data using existing media such as wired or wireless broadband lines or storing video / image data using existing storage media.
[0003] In addition, interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content and holograms has been increasing recently, and the broadcast of videos / images with different visual characteristics from real images, such as game images, is increasing.
[0004] Therefore, a highly efficient video compression technology is required to effectively compress, transmit, store, and play back high-resolution, high-quality video / image information having the above-mentioned various characteristics. Summary of the Invention [Problem to be solved by the invention]
[0005] The technical problem of this document is to provide a method and apparatus for improving video / image coding efficiency.
[0006] Another technical problem of this document is to provide a DMVR-based video coding method and apparatus that uses multiple layers to improve inter-coding efficiency. [Means for solving the problem]
[0007] According to one embodiment of the present document, there is provided a video decoding method executed by a decoding device, the method including: determining whether to use multi-layer Decoder-side Motion Vector Refinement (MDMVR) for a current block; deriving a refined motion vector for the current block based on whether the MDMVR is used for the current block; deriving prediction samples for the current block based on the refined motion vector; and generating reconstructed samples for the current block based on the prediction samples.
[0008] According to another embodiment of the present document, there is provided a video encoding method performed by an encoding apparatus, the method including: determining whether to use multi-layer Decoder-side Motion Vector Refinement (MDMVR) for a current block; deriving a refined motion vector for the current block based on whether the MDMVR is used for the current block; deriving prediction samples for the current block based on the refined motion vector; deriving residual samples based on the prediction samples; and encoding video information including information on the residual samples to generate a bitstream.
[0009] According to another embodiment of the present document, a computer-readable digital storage medium is provided that stores encoded video / image information and / or a bitstream generated by the video / image encoding method disclosed in at least one of the embodiments of the present document.
[0010] According to another embodiment of the present document, there is provided a method for transmitting data including a bitstream of video information, the method including: obtaining the bitstream of video information, determining whether a multi-layer Decoder-side Motion Vector Refinement (MDMVR) is used for a current block, deriving a refined motion vector for the current block based on whether the MDMVR is used for the current block, deriving prediction samples for the current block based on the refined motion vector, deriving residual samples based on the prediction samples, and encoding video information including information on the residual samples; and transmitting the data including the bitstream. [Effects of the Invention]
[0011] This document may have various advantages. For example, this document may improve overall image / video compression efficiency. This document may also enable efficient decoder-side motion vector refinement-based image coding using multiple layers. This document may also provide a decoder-side motion vector refinement method using multiple layers that improves the performance of an inter-coding structure and does not increase processing complexity. This document may also improve processing efficiency without increasing processing complexity by providing various methods for determining whether to use decoder-side motion vector refinement using multiple layers.
[0012] The effects that can be obtained through the specific embodiments of this document are not limited to the effects listed above. For example, there may be various technical effects that a person having ordinary skill in the related art can understand or derive from this document. Therefore, the specific effects of this document are not limited to those explicitly described in this document, but may include various effects that can be understood or derive from the technical features of this document. [Brief explanation of the drawings]
[0013] [Figure 1] 1 illustrates schematically an example of a video / image coding system to which embodiments of the present document can be applied. [Figure 2] 1 is a diagram illustrating the configuration of a video / image encoding device to which the embodiments of this document can be applied. [Figure 3] 1 is a diagram illustrating the configuration of a video / image decoding device to which an embodiment of the present document can be applied. [Figure 4] 1 shows an exemplary hierarchical structure for coded video / images. [Figure 5] FIG. 1 is a diagram illustrating an embodiment of a process for performing Decoder-side Motion Vector Refinement (DMVR). [Figure 6] 1 is a diagram illustrating an embodiment of a process for performing Decoder-side Motion Vector Refinement (DMVR) using SAD (sum of absolute differences). FIG. [Figure 7] 1 illustrates an exemplary MDMVR structure according to one embodiment of the present document. [Figure 8] 10 illustrates an exemplary MDMVR structure according to another embodiment of the present document. [Figure 9]1 illustrates an example of a video / image encoding method and associated components according to embodiment(s) of the present document; [Figure 10] 1 illustrates an example of a video / image encoding method and associated components according to embodiment(s) of the present document; [Figure 11] 1 illustrates an example of a video / image decoding method and associated components according to embodiment(s) of the present document; [Figure 12] 1 illustrates an example of a video / image decoding method and associated components according to embodiment(s) of the present document; [Figure 13] 1 illustrates an example of a content streaming system to which the embodiments disclosed herein can be applied. DETAILED DESCRIPTION OF THE INVENTION
[0014] This document may be modified in various ways and may have various embodiments. Specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit the embodiments of this document to the specific embodiments. The terms used in this document are used merely to describe specific embodiments and are not intended to limit the technical ideas of this document. Singular expressions include plural expressions unless the context clearly dictates otherwise. In this document, terms such as "comprise" or "have" specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the document, and should be understood not to preclude the presence or possibility of addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0015] Meanwhile, each component in the drawings described in this document is illustrated independently for the convenience of explaining the different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also within the scope of this document as long as they do not deviate from the essence of this document.
[0016] In this document, "A or B" can mean "A only," "B only," or "both A and B." In other words, in this document, "A or B" can be interpreted as "A and / or B." For example, in this document, "A, B or C" can mean "A only," "B only," "C only," or "any combination of A, B, and C."
[0017] The slash ( / ) and comma used in this document can mean "and / or." For example, "A / B" can mean "A and / or B." This allows "A / B" to mean "A only," "B only," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0018] In this document, "at least one of A and B" can mean "A only," "B only," or "both A and B." Also, in this document, the expressions "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B."
[0019] Additionally, in this document, "at least one of A, B, and C" can mean "A only," "B only," "C only," or "any combination of A, B, and C." Additionally, "at least one of A, B, or C" and "at least one of A, B, and / or C" can mean "at least one of A, B, and C."
[0020] Furthermore, parentheses used in this document may mean "for example." Specifically, when "prediction (intra prediction)" is used, "intra prediction" is proposed as an example of "prediction." In other words, "prediction" in this document is not limited to "intra prediction," and "intra prediction" is proposed as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is used, "intra prediction" is proposed as an example of "prediction."
[0021] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document may be applied to methods disclosed in the versatile video coding (VVC) standard. The methods / embodiments disclosed in this document may also be applied to methods disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation audio video coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).
[0022] This document presents various embodiments relating to video / image coding, which, unless otherwise stated, may also be implemented in combination with one another.
[0023] In this document, video may refer to a collection of a series of images over time. A picture generally refers to a unit that represents an image at a specific time period, and a subpicture / slice / tile is a unit that constitutes part of a picture in coding. A subpicture / slice / tile may include one or more coding tree units (CTUs). A picture may consist of one or more subpictures / slices / tiles. A picture may consist of one or more tile groups. A tile group may include one or more tiles. A brick may refer to a rectangular area of a row of CTUs within a tile within a picture. A tile may be partitioned into multiple bricks, and each brick may consist of one or more rows of CTUs within the tile. A tile that is not partitioned into multiple bricks may also be called a brick. A brick scan can indicate a specific sequential ordering of CTUs that partition a picture, where the CTUs can be aligned within a brick by a CTU raster scan, the bricks within a tile can be aligned consecutively by a raster scan of the bricks in the tile, and the tiles within a picture can be aligned consecutively by a raster scan of the tiles in the picture. A subpicture can also indicate a rectangular region of one or more slices within a picture. That is, a subpicture can include one or more slices that collectively cover a rectangular region of a picture. A tile is a specific tile column and a rectangular region of a CTU within the specific tile column. The tile column is a rectangular region of a CTU, where the rectangular region has a height equal to the height of the picture, and the width can be specified by a syntax element in a picture parameter set. The tile row is a rectangular region of a CTU, where the rectangular region has a width equal to the height of the picture, and the width is specified by a syntax element in a picture parameter set.A tile scan may indicate a specific sequential ordering of CTUs that partition a picture, and the CTUs may be consecutively aligned in a raster scan of CTUs within a tile, and tiles within a picture may be consecutively aligned in a raster scan of the tiles of the picture. A slice may include an integer number of bricks of a picture, and the integer number of bricks may be included in one NAL unit. A slice may consist of multiple complete tiles or may be a continuous sequence of complete bricks of one tile. In this document, the terms tile group and slice may be used interchangeably. For example, in this document, tile group / tile group header may be referred to as slice / slice header.
[0024] A pixel or a pel may refer to the smallest unit that constitutes one picture (or image). A "sample" may also be used as a term corresponding to a pixel. A sample may generally refer to a pixel or a pixel value, or may refer to only a pixel / pixel value of a luma component, or may refer to only a pixel / pixel value of a chroma component.
[0025] A unit may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area. In a general case, an M×N block may include samples (or a sample array) consisting of M columns and N rows, or a set (or an array) of transform coefficients.
[0026] In this document, technical features individually described in one drawing may be embodied individually or simultaneously.
[0027] The following drawings are created to explain a specific example of the present document. The names of specific devices and names of specific signals / messages / fields shown in the drawings are provided for illustrative purposes only, and the technical features of the present document are not limited to the specific names used in the following drawings.
[0028] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Hereinafter, the same reference numerals will be used to refer to the same components in the drawings, and duplicated descriptions of the same components may be omitted.
[0029] FIG. 1 illustrates schematically an example of a video / image coding system in which embodiments of the present document may be applied.
[0030] As shown in Figure 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in a file or streaming format via a digital recording medium or a network.
[0031] The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may be a separate device or an external component.
[0032] A video source can acquire video / images through a video / image capture, synthesis, or generation process. A video source can include a video / image capture device and / or a video / image generation device. A video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer, etc., in which case the video / image capture process can be replaced by a process in which the associated data is generated.
[0033] An encoding device can encode input video / images. The encoding device can perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0034] The transmitter can transmit the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving device via a digital recording medium or a network in the form of a file or streaming. The digital recording medium can include various recording media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter can include elements for generating a media file in a predetermined file format and elements for transmission via a broadcasting / communication network. The receiver can receive / extract the bitstream and transmit it to a decoding device.
[0035] The decoding device can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, prediction, etc., which correspond to the operations of the encoding device.
[0036] The renderer can render the decoded video / image, and the rendered video / image can be displayed via a display unit.
[0037] 2 is a diagram illustrating a configuration of a video / image encoding device to which an embodiment of this document can be applied. Hereinafter, the encoding device may include an image encoding device and / or a video encoding device.
[0038] As shown in FIG. 2, the encoding device 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter predictor 221 and an intra predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image dividing unit 210, the predicting unit 220, the residual processing unit 230, the entropy encoding unit 240, the adding unit 250, and the filtering unit 260 may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. Also, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital recording medium. The hardware components may further include the memory 270 as an internal / external component.
[0039] The image division unit 210 may divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quad-tree, binary-tree, ternary-tree (QTBTTT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and then the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied first. The coding procedure according to this document may be performed based on the final coding unit that is not further divided. In this case, the largest coding unit may be immediately used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as needed, and the coding unit of the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be divided or partitioned from the final coding unit.The prediction unit is a unit of sample prediction, and the transform unit is a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0040] The term "unit" can be used interchangeably with terms such as "block" or "area." In general, an MxN block can refer to a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally refer to a pixel or pixel value, and can refer to only a pixel / pixel value of the luma component, or only a pixel / pixel value of the chroma component. A sample can also be used as a term corresponding to one pixel or pel of a picture (or image).
[0041] The encoding apparatus 200 may generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input image signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, a unit in the encoder 200 that subtracts the prediction signal (predicted block, prediction sample array) from the input image signal (original block, original sample array) may be referred to as a subtraction unit 231. The prediction unit may perform prediction on a current block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit may generate various information related to prediction, such as prediction mode information, and transmit the information to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The prediction information can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0042] The intra prediction unit 222 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away, depending on the prediction mode. In intra prediction, prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, DC mode and planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0043] The inter prediction unit 221 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (such as L0 prediction, L1 prediction, or Bi prediction). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), or the like, and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 221 may configure a motion information candidate list based on neighboring blocks and generate information indicating which candidates are used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of a skip mode or a merge mode, the inter predictor 221 may use motion information of neighboring blocks as motion information of the current block. In the case of the skip mode, unlike in the merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the neighboring block is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0044] The predictor 220 may generate a prediction signal based on various prediction methods, which will be described later. For example, the predictor may apply intra prediction or inter prediction for prediction of a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The predictor may also use an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be performed similarly to inter prediction in deriving a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described herein. The palette mode may be seen as an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture may be signaled based on information about a palette table and a palette index.
[0045] The prediction signal generated by the prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) may be used to generate a reconstructed signal or a residual signal. The transform unit 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph representing inter-pixel relationship information. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size or non-square blocks of variable size.
[0046] The quantizer 233 quantizes the transform coefficients and transmits them to the entropy encoder 240. The entropy encoder 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 240 may perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. In addition to the quantized transform coefficients, the entropy encoder 240 may also encode information required for video / image restoration (e.g., values of syntax elements, etc.) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of network abstraction layer (NAL) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. Information and / or syntax elements transmitted / signaled from an encoding device to a decoding device in this document may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream.The bitstream may be transmitted via a network or stored on a digital recording medium. Here, the network may include a broadcasting network and / or a communication network, and the digital recording medium may include various recording media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoding unit 240 and / or a storage unit (not shown) for storing the signal may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoding unit 240.
[0047] The quantized transform coefficients output from the quantization unit 233 may be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) may be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transform unit 235. The adder 250 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the current block, such as when skip mode is applied, a predicted block may be used as the reconstructed block. The adder 250 may be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, or may be used for inter prediction of the next picture after filtering, as described below.
[0048] Meanwhile, luma mapping with chroma scaling (LMCS) can be applied during picture encoding and / or reconstruction.
[0049] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 260 may generate various information related to filtering and transmit it to the entropy encoding unit 240, as will be described later in connection with each filtering method. The filtering information may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0050] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding apparatus can avoid prediction mismatch between the encoding apparatus 200 and the decoding apparatus 300 and can also improve encoding efficiency.
[0051] The memory 270DPB may store modified reconstructed pictures for use as reference pictures in the inter predictor 221. The memory 270 may store motion information of blocks from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in already reconstructed pictures. The stored motion information may be transmitted to the inter predictor 221 to be used as motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 270 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 222.
[0052] 3 is a diagram illustrating the configuration of a video / image decoding device to which an embodiment of this document can be applied. Hereinafter, the decoding device may include an image decoding device and / or a video decoding device.
[0053] Referring to FIG. 3, the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra predictor 331 and an inter predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. Depending on the embodiment, the entropy decoding unit 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be implemented as a single hardware component (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) and may be implemented as a digital storage medium. The hardware components may further include a memory 360 as an internal / external component.
[0054] When a bitstream including video / image information is input, the decoding device 300 can reconstruct an image corresponding to the process by which the video / image information was processed by the encoding device of FIG. 2. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using a processing unit applied by the encoding device. Accordingly, the processing unit for decoding is, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximal coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a playback device.
[0055] The decoding device 300 may receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal may be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 may parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The decoding device may further decode pictures based on the information on the parameter sets and / or the general constraint information. Signaling / received information and / or syntax elements, which will be described later in this document, may be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration, quantized values of transform coefficients related to residuals, etc. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using information on the syntax element to be decoded and decoded information on neighboring and current blocks, or information on symbols / bins decoded in previous steps, predicts the occurrence probability of the bins according to the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. After determining the context model, the CABAC entropy decoding method may update the context model using information on the decoded symbols / bins for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 310, information related to prediction is provided to a prediction unit (inter prediction unit 332 and intra prediction unit 331), and residual values entropy decoded by the entropy decoding unit 310, i.e., quantized transform coefficients and related parameter information, may be input to a residual processing unit 320. The residual processing unit 320 may derive a residual signal (residual block, residual sample, residual sample array). In addition, among the information decoded by the entropy decoding unit 310, information related to filtering may be provided to a filtering unit 350. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. Meanwhile, the decoding device according to this document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, the inverse transform unit 322, the addition unit 340, the filtering unit 350, the memory 360, the inter prediction unit 332, and the intra prediction unit 331.
[0056] The inverse quantization unit 321 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 321 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit 321 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.
[0057] The inverse transform unit 322 performs inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0058] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 310, and may determine a specific intra / inter prediction mode.
[0059] The predictor 320 may generate a prediction signal based on various prediction methods, which will be described later. For example, the predictor may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The predictor may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be performed similarly to inter prediction in deriving a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode may be seen as an example of intra coding or intra prediction. When the palette mode is applied, information regarding a palette table and a palette index may be included in the video / image information and signaled.
[0060] The intra prediction unit 331 may predict a current block by referring to samples in a current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block depending on the prediction mode. In intra prediction, prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 may also determine a prediction mode to be applied to the current block using prediction modes applied to neighboring blocks.
[0061] The inter prediction unit 332 may derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted from the inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter prediction unit 332 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.
[0062] The adder 340 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the current block, such as when a skip mode is applied, the predicted block may be used as a reconstructed block.
[0063] The adder 340 may be referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of a next block to be processed in the current picture, may be output after filtering as described below, or may be used for inter prediction of a next picture.
[0064] Meanwhile, LMCS (luma mapping with chroma scaling) can be applied during the picture decoding process.
[0065] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may transmit the modified reconstructed picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0066] The (modified) reconstructed picture stored in the DPB of the memory 360 may be used as a reference picture in the inter predictor 332. The memory 360 may store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 260 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 360 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 331.
[0067] In this document, the embodiments described for the filtering unit 260, inter prediction unit 221, and intra prediction unit 222 of the encoding device 200 can also be applied identically or correspondingly to the filtering unit 350, inter prediction unit 332, and intra prediction unit 331 of the decoding device 300, respectively.
[0068] FIG. 4 shows an exemplary hierarchical structure for coded video / pictures.
[0069] Referring to Figure 4, coded video / images can be divided into a VCL (video coding layer) that handles the video / image decoding process and itself, a lower system that transmits and stores coded information, and a NAL (network abstraction layer) that exists between the VCL and the lower system and is responsible for network adaptation functions.
[0070] For example, in VCL, VCL data including compressed image data (slice data) can be generated, or a parameter set including a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), a Video Parameter Set (VPS), or an SEI (Supplemental Enhancement Information) message that is additionally required for the video decoding process can be generated.
[0071] For example, in NAL, NAL units can be generated by adding header information (NAL unit header) to RBSP (Raw Byte Sequence Payload) generated by VCL. In this case, the RBSP can refer to slice data, parameter sets, SEI messages, etc. generated by VCL. The NAL unit header can include NAL unit type information specified by the RBSP data included in the corresponding NAL unit.
[0072] For example, as shown in Figure 4, NAL units can be classified into VCL NAL units and non-VCL NAL units according to the RBSP generated by the VCL. A VCL NAL unit can refer to a NAL unit that contains information about an image (slice data), and a non-VCL NAL unit can refer to a NAL unit that contains information required for video decoding (parameter set or SEI message).
[0073] The VCL NAL units and non-VCL NAL units can be transmitted over a network by attaching header information according to the subsystem's data standard. For example, the NAL units can be converted into a predetermined standard data format such as H.266 / VVC file format, real-time transport protocol (RTP), transport stream (TS), etc., and can be transmitted over various networks.
[0074] Also, as mentioned above, the NAL unit type of an NAL unit can be specified by the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored and signaled in the NAL unit header.
[0075] For example, NAL units can be classified into VCL NAL unit types and non-VCL NAL unit types depending on whether they contain information about a video (slice data). In addition, VCL NAL unit types can be classified according to the characteristics and type of a picture included in the VCL NAL unit, and non-VCL NAL unit types can be classified according to the type of parameter set.
[0076] The following are examples of NAL unit types specified by the types of parameter sets contained in the Non-VCL NAL unit types:
[0077] -APS (Adaptation Parameter Set) NAL unit: Type for NAL units containing APS
[0078] -DPS (Decoding Parameter Set) NAL unit: Type for NAL units containing DPS
[0079] -VPS (Video Parameter Set) NAL unit: Type for NAL units containing VPS
[0080] -SPS (Sequence Parameter Set) NAL unit: Type for NAL units containing SPS
[0081] -PPS (Picture Parameter Set) NAL unit: Type for NAL units containing PPS
[0082] -PH (Picture header) NAL unit: Type for NAL units containing PH
[0083] The above-mentioned NAL unit type may have syntax information for the NAL unit type, and the syntax information may be stored and signaled in a NAL unit header. For example, the syntax information may be nal_unit_type, and the NAL unit type may be specified as a nal_unit_type value.
[0084] Meanwhile, one picture may include multiple slices, and each slice may include a slice header and slice data. In this case, one picture header may be added for multiple slices (a set of slice headers and slice data). A picture header (picture header syntax) may include information / parameters commonly applicable to pictures. A slice header (slice header syntax) may include information / parameters commonly applicable to slices. An APS (APS syntax) or PPS (PPS syntax) may include information / parameters commonly applicable to one or more slices or pictures. An SPS (SPS syntax) may include information / parameters commonly applicable to one or more sequences. A VPS (VPS syntax) may include information / parameters commonly applicable to multiple layers. A DPS (DPS syntax) may include information / parameters commonly applicable to the entire picture. A DPS may include information / parameters related to the concatenation of a coded video sequence (CVS).
[0085] In this document, image / video information encoded from an encoding device to a decoding device and signaled in the form of a bitstream may include not only intra-picture partitioning-related information, intra / inter prediction information, inter-layer prediction-related information, residual information, in-loop filtering information, etc., but also information included in the slice header, information included in the picture header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS. In addition, the image / video information may further include information in a NAL unit header.
[0086] Meanwhile, as described above, prediction is performed to improve compression efficiency when performing video coding. Accordingly, a predicted block including predicted samples for a current block, which is a block to be coded, can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived in the same way by an encoding device and a decoding device, and the encoding device signals information (residual information) regarding the residual between an original block and a predicted block, rather than the original sample values of the original block, to the decoding device, thereby improving video coding efficiency. The decoding device derives a residual block including residual samples based on the residual information, and generates a reconstructed block including reconstructed samples by adding the residual block and the predicted block, thereby generating a reconstructed picture including the reconstructed block.
[0087] The residual information may be generated through a transform and quantization procedure. For example, an encoding device may derive a residual block between an original block and a predicted block, perform a transform procedure on residual samples (residual sample array) included in the residual block to derive transform coefficients, and perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, and then signal the related residual information (via a bitstream) to a decoding device. Here, the residual information may include information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding device may derive residual samples (or residual blocks) by performing an inverse quantization / inverse transform procedure based on the residual information. The decoding device may generate a reconstructed picture based on the predicted block and the residual block. The encoding device may also derive a residual block by inverse quantizing / inverse transforming the quantized transform coefficients for reference for inter-prediction of a future picture, and generate a reconstructed picture based on the residual block.
[0088] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. If the quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. If the transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for consistency of expression. Furthermore, whether the transform / inverse transform is omitted may be signaled based on transform_skip_flag.
[0089] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficient(s), and the information about the transform coefficient(s) may be signaled via residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficient(s), and scaled transform coefficients may be derived through an inverse transform (scaling) of the transform coefficients. Residual samples may be derived based on an inverse transform (transform) of the scaled transform coefficients. This may be similarly applied / expressed in other parts of this document.
[0090] Meanwhile, as described above, intra prediction or inter prediction can be applied to perform prediction on the current block. Hereinafter, a case where inter prediction is applied to the current block will be described.
[0091] A prediction unit (more specifically, an inter prediction unit) of an encoding / decoding device may perform inter prediction on a block-by-block basis to derive prediction samples. Inter prediction may refer to prediction derived in a manner dependent on data elements (e.g., sample values, motion information, etc.) of picture(s) other than the current picture. When inter prediction is applied to a current block, a predicted block (prediction sample array) for the current block may be derived based on a reference block (reference sample array) identified by a motion vector in a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, motion information of the current block may be predicted in block, sub-block, or sample units based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction type information (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). When inter prediction is applied, neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block may be the same or different. The temporally neighboring block may be called a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporally neighboring block may be called a collocated picture (colPic). For example, a motion information candidate list may be constructed based on the neighboring blocks of the current block, and flag or index information indicating which candidate is selected (used) to derive a motion vector and / or a reference picture index for the current block may be signaled.Inter prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block is the same as the motion information of the selected neighboring block. In skip mode, unlike merge mode, a residual signal is not transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected neighboring block is used as a motion vector predictor, and a motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.
[0092] The motion information may include L0 motion information and / or L1 motion information depending on the inter-prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be referred to as an L0 motion vector or MVL0, and a motion vector in the L1 direction may be referred to as an L1 motion vector or MVL1. Prediction based on an L0 motion vector may be referred to as L0 prediction, prediction based on an L1 motion vector may be referred to as L1 prediction, and prediction based on both an L0 motion vector and an L1 motion vector may be referred to as bi-prediction. Here, the L0 motion vector may indicate a motion vector associated with a reference picture list L0 (L0), and the L1 motion vector may indicate a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include pictures that are earlier in output order than the current picture as reference pictures, and the reference picture list L1 may include pictures that are later in output order than the current picture. The earlier picture may be referred to as a forward (reference) picture, and the later picture may be referred to as a backward (reference) picture. The reference picture list L0 may further include, as reference pictures, pictures that are later in output order than the current picture. In this case, the previous picture may be indexed first in the reference picture list L0, and the later picture may be indexed next. The reference picture list L1 may further include, as reference pictures, pictures that are earlier in output order than the current picture. In this case, the later picture may be indexed first in the reference picture list L1, and the previous picture may be indexed next. Here, the output order may correspond to the picture order count (POC) order.
[0093] In addition, various inter prediction modes can be used when applying inter prediction to the current block. For example, various modes such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, and historical motion vector prediction (HMVP) mode can be used. Decoder side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, and bi-directional optical flow (BDOF) can also be used as additional modes. The affine mode can also be referred to as affine motion prediction mode. The MVP mode can also be referred to as advanced motion vector prediction mode. In this document, some modes and / or motion information candidates derived by some modes can be included as one of the motion information-related candidates of other modes.
[0094] Prediction mode information indicating the inter prediction mode of the current block may be signaled from the encoding apparatus to the decoding apparatus. In this case, the prediction mode information may be included in a bitstream and received by the decoding apparatus. The prediction mode information may include index information indicating one of multiple candidate modes. Alternatively, the inter prediction mode may be indicated through hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether the skip mode is applied, a merge flag may be signaled to indicate whether the merge mode is applied when the skip mode is not applied, or an MVP mode may be applied when the merge mode is not applied, or a flag for additional classification may be further signaled. The affine mode may be signaled as an independent mode or as a mode dependent on the merge mode or MVP mode. For example, the affine mode may include affine merge mode and affine MVP mode.
[0095] Furthermore, motion information of the current block may be used when applying inter prediction to the current block. The encoding apparatus may derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding apparatus may search for a similar reference block with high correlation using an original block in an original picture for the current block in fractional pixel units within a predetermined search range in the reference picture, thereby deriving motion information. Block similarity may be derived based on a difference in phase-based sample values. For example, block similarity may be calculated based on the sum of absolute differences (SAD) between the current block (or a template of the current block) and a reference block (or a template of the reference block). In this case, motion information may be derived based on the reference block with the smallest SAD within the search range. The derived motion information may be signaled to the decoding apparatus in various ways based on the inter prediction mode.
[0096] As described above, a predicted block for the current block may be derived based on motion information derived according to the inter prediction mode. The predicted block may include predicted samples (prediction sample array) of the current block. If the motion vector (MV) of the current block points to a fractional sample unit, an interpolation procedure may be performed, through which predicted samples of the current block may be derived based on reference samples in fractional sample units within the reference picture. If affine inter prediction is applied to the current block, predicted samples may be generated based on sample / sub-block unit MVs. If bi-prediction is applied, predicted samples derived through a weighted sum or weighted average (by phase) of predicted samples derived based on L0 prediction (i.e., prediction using a reference picture in the reference picture list L0 and MVL0) and predicted samples derived based on L1 prediction (i.e., prediction using a reference picture in the reference picture list L1 and MVL1) may be used as predicted samples of the current block. When bi-prediction is applied, if the reference picture used for L0 prediction and the reference picture used for L1 prediction are located in different temporal directions relative to the current picture (i.e., if it is bi-predictive and corresponds to bidirectional prediction), this can be called true bi-prediction.
[0097] As described above, reconstructed samples and reconstructed pictures can be generated based on the derived predicted samples, and then procedures such as in-loop filtering can be performed.
[0098] On the other hand, skip mode and / or merge mode have limitations in motion prediction because they predict the motion of a current block based on the motion vector of a neighboring block without MVD (Motion Vector Difference). To overcome the limitations of skip mode and / or merge mode, the motion vector can be refined by applying Decoder-side Motion Vector Refinement (DMVR), Bi-directional optical flow (BDOF) mode, etc. The DMVR and BDOF modes can be used when true bi-prediction is applied to the current block.
[0099] FIG. 5 is a diagram illustrating an embodiment of a process for performing decoder-side motion vector refinement (DMVR).
[0100] DMVR is a method of performing motion prediction by refining motion information of neighboring blocks on the decoder side. When DMVR is applied, the decoder can derive refined motion information through cost comparison based on a template generated using motion information of neighboring blocks in merge / skip mode. In this case, the accuracy of motion prediction can be increased without additional signaling information, thereby improving compression performance.
[0101] For convenience of explanation, this document will be described mainly in terms of a decoding device, but the DMVR according to the embodiments of this document can be implemented in the same manner in an encoding device.
[0102] 5, the decoding apparatus may derive prediction blocks (i.e., reference blocks) identified by initial motion vectors (or motion information) (e.g., MV0 and MV1) in the list0 and list1 directions, and generate a template (or bilateral template) by weighting (e.g., averaging) the derived prediction blocks (step 1). Here, the initial motion vectors (MV0 and MV1) may indicate motion vectors derived using motion information of neighboring blocks in merge / skip mode.
[0103] The decoding apparatus can then derive motion vectors (e.g., MV0′ and MV1′) that minimize the difference between the template and the sample region of the reference picture through a template matching operation (step 2). Here, the sample region indicates a surrounding region of the initial prediction block in the reference picture, and the sample region may also be referred to as a surrounding region, reference region, search region, search range, search space, etc. The template matching operation may include calculating a cost measurement value between the template and the sample region of the reference picture. For example, the sum of absolute differences (SAD) may be used for the cost measurement. As an example, a normalized SAD may be used as the cost function. In this case, the matching cost may be given as SAD(T-mean(T), 2*P[x]-2*mean(P[x])), where T indicates the template and P[x] indicates a block within the search region. The motion vector that calculates the minimum template cost for each of the two reference pictures may be considered as an updated motion vector (replacing the initial motion vector). As shown in Figure 5, the decoding device can generate a final bi-directional prediction result (i.e., a final bi-directional prediction block) using the updated motion vectors MV0' and MV1'. In one embodiment, multi-iteration for deriving updated (or new) motion vectors can be used to obtain the final bi-directional prediction result.
[0104] In one embodiment, the decoding device may invoke the DMVR process to improve the accuracy of initial motion compensation prediction (i.e., motion compensation prediction via conventional merge / skip mode). For example, the decoding device may perform the DMVR process when the prediction mode of the current block is merge mode or skip mode and bidirectional bi-prediction, in which reference pictures in both directions are in opposite directions relative to the current picture in display order, is applied to the current block.
[0105] FIG. 6 is a diagram illustrating an embodiment of a process for performing decoder-side motion vector refinement (DMVR) using sum of absolute differences (SAD).
[0106] As described above, a decoding apparatus can measure the matching cost using SAD when performing DMVR. As an example, FIG. 6 illustrates a method for refining a motion vector by calculating the mean sum of absolute differences (MRSAD) between prediction samples in two reference pictures without generating a template. That is, the method of FIG. 6 illustrates an example of bilateral matching using MRSAD.
[0107] 6, the decoding apparatus may derive neighboring pixels of a pixel (sample) indicated by a motion vector (MV0) in the list0 (L0) direction on the L0 reference picture, and neighboring pixels of a pixel (sample) indicated by a motion vector (MV1) in the list1 (L1) direction on the L1 reference picture. The decoding apparatus may then measure the matching cost by calculating the MRSAD between an L0 predicted block (i.e., an L0 reference block) identified by a motion vector indicating the neighboring pixels derived on the L0 reference picture and an L1 predicted block (i.e., an L1 reference block) identified by a motion vector indicating the neighboring pixels derived on the L1 reference picture. In this case, the decoding apparatus may select a search point with the minimum cost (i.e., a search area with the minimum SAD between the L0 predicted block and the L1 predicted block) as a refined motion vector pair. That is, the refined motion vector pair may include a refined L0 motion vector pointing to a pixel location (L0 prediction block) with the smallest cost in the L0 reference picture and a refined L1 motion vector pointing to a pixel location (L1 prediction block) with the smallest cost in the L1 reference picture.
[0108] In one embodiment, after a search region of a reference picture is set for calculating the matching cost, unidirectional prediction may be performed using a regular 8-tap DCTIF interpolation filter. Also, in one example, 16-bit precision may be used for the MRSAD calculation, and clipping and / or rounding operations may not be applied before the MRSAD calculation in consideration of an internal buffer.
[0109] As described above, when true bi-prediction is applied to the current block, BDOF can be used to refine the bi-prediction signal. When bi-prediction is applied to the current block, BDOF (Bi-directional optical flow) can be used to calculate improved motion information and generate prediction samples based on the information. For example, BDOF can be applied at a 4x4 sub-block level. That is, BDOF can be performed in units of 4x4 sub-blocks within the current block. Alternatively, BDOF can be applied only to the luma component. Alternatively, BDOF can be applied only to the chroma component, or to both the luma component and the chroma component.
[0110] As the name suggests, BDOF mode is based on the optical flow concept, which assumes that the motion of objects is smooth. For each 4x4 sub-block, motion refinement (v) is performed by minimizing the difference value between the L0 and L1 predicted samples. x , v y ) can be calculated, and motion refinement can be used to adjust the bi-predictive sample values in the 4x4 sub-blocks.
[0111] The aforementioned DMVR and BDOF are techniques that refine motion information to perform prediction when true bi-prediction is applied (in this case, true bi-prediction refers to the case where motion prediction / compensation is performed using a reference picture in a different direction based on the picture of the current block), and can be seen as refinement techniques with a similar concept in that they assume that the movement of an object within a picture occurs at a constant speed and in a constant direction.
[0112] Meanwhile, the following describes structures and features that can be used in determining / parsing inter-mode(s) and / or inter-prediction(s) at a decoder to improve the performance of inter-coding structures. The described method(s) are based on Versatile Video Coding (VVC), but can also be applied to other past or future video coding technologies.
[0113] In this document, a multi-pass DMVR technique can be applied to improve inter-coding performance. Multi-pass DMVR (i.e., MDMVR) is a technique for additionally improving (and simplifying) DMVR technology in next-generation video codecs. In the first pass, bilateral matching (BM) is applied to the coding block, in the second pass, BM is applied to each 16x16 sub-block within the coding block, and in the third pass, BDOF is applied to refine the MV of each 8x8 sub-block. Here, the refined MV can be stored for spatial and temporal motion vector prediction.
[0114] More specifically, in the first pass of multi-pass DMVR, a refined MV can be derived by applying bilateral matching (BM) to a coding block. Similar to Decoder-Side Motion Vector Refinement (DMVR), a refined MV using bi-prediction can be searched around two initial MVs (MV0 and MV1) in reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) can be derived around the initial MVs based on the minimum bilateral matching cost between two reference blocks in L0 and L1.
[0115] The BM can perform a local search to derive the integer sample precision intDeltaMV. The local search can be repeated over the horizontal search range [-sHor, sHor] and the vertical search range [-sVer, sVer] by applying a 3x3 square search pattern, where the values of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.
[0116] The bidirectional matching cost can be calculated as bilCost = mvDistanceCost + sadCost. If the block size cbW * cbH is greater than 64, the MRSAD cost function can be applied to remove the DC distortion effect between reference blocks. If the bilCost at the center point of the 3x3 search pattern has the minimum cost, the intDeltaMV local search can be terminated. Otherwise, the minimum cost search can continue until the current minimum cost search point becomes the new center point of the 3x3 search pattern and the edge of the search range is reached.
[0117] The existing fractional sample refinement can be additionally applied to derive the final deltaMV. After the first pass, the refined MV can be derived as follows:
[0118] MV0_pass1=MV0+deltaMV
[0119] MV1_pass1=MV1-deltaMV
[0120] In the second pass of MDMVR, refined MVs can be derived by applying BM to 16x16 grid sub-blocks. For each sub-block, refined MVs can be searched around the two MVs (MV0_pass1, MV1_pass1) obtained in the first pass in the reference picture lists L0 and L1. Refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) can be derived based on the minimum bidirectional matching cost between the two reference sub-blocks in L0 and L1.
[0121] For each sub-block, the BM can perform a full search to derive the integer sample precision intDeltaMV. The full search has a search range of [-sHor, sHor] horizontally and [-sVer, sVer] vertically, where the values of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.
[0122] The bidirectional matching cost can be calculated by applying a cost factor to the SATD cost between two reference subblocks, as follows: bilCost = satdCost * costFactor. The search area (2 * sHor + 1) * (2 * sVer + 1) can be divided into up to five diamond-shaped search areas. Each search area is assigned a costFactor determined by the distance (intDeltaMV) between each search point and the starting MV. Each diamond area can be processed in order starting from the center of the search area. In each area, search points can be processed in raster scan order starting from the top left corner of the area to the bottom right corner. If the minimum bilCost within the current search area is less than a threshold, such as sbW * sbH, the int-pel full search is terminated. If not, the int-pel full search can continue with the next search area until all search points have been examined.
[0123] In the existing VVC, DMVR fractional sample refinement can be additionally applied to derive the final deltaMV(sbIdx2). In the second pass, the refined MV can be derived as follows:
[0124] MV0_pass2(sbIdx2)=MV0_pass1+deltaMV(sbIdx2)
[0125] MV1_pass2(sbIdx2)=MV1_pass1-deltaMV(sbIdx2)
[0126] In the third pass of MDMVR, refined MVs can be derived by applying BDOF to the 8x8 grid sub-blocks. For each 8x8 sub-block, BDOF refinement can be applied to derive scaled Vx and Vy without clipping, starting from the refined MV of the parent sub-block in the second pass. The derived bioMv(Vx, Vy) can be clipped between -32 and 32, rounded to 1 / 16 sample precision.
[0127] The improved MVs in the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) can be derived as follows:
[0128] MV0_pass3(sbIdx3)=MV0_pass2(sbIdx2)+bioMv
[0129] MV1_pass3(sbIdx3)=MV0_pass2(sbIdx2)-bioMv
[0130] Meanwhile, this document proposes a method for applying the above-mentioned decoder-side motion vector derivation (DMVD) (i.e., DMVR and / or multi-pass DMVR) to improve the performance of the inter-coding structure without increasing the processing complexity. To this end, the following aspects can be taken into consideration. That is, the proposed method can include the following embodiments, and the proposed embodiments can be applied individually or in combination.
[0131] 1. As an example, DMVR can be applied at multiple levels, passes, or steps to derive a final refined motion vector. For purposes of this disclosure, multi-layer DMVR can be referred to as MDMVR. Multiple layers can also be referred to as multiple passes, levels, or steps.
[0132] 2. Also, by way of example, the use of MDMVR can be determined by taking into consideration several factors, including:
[0133] Each CU can have a different number of layers, for example, the current PU can have two layers, and the next PU needs three layers.
[0134] b. The MVD information can be used to determine the level or layer number of DMVR that should or can be applied to a block.
[0135] i) For example, in the first case (i.e., use for on / off control of DMVR), if the MVD in merge mode exceeds a predetermined threshold, it is possible not to apply DMVR. For example, the threshold can be selected in 1 / 4 pixel (pel), 1 / 2 pixel (pel), 1 pixel (pel), 4 pixels (pel), and / or other suitable MVD units.
[0136] ii) For example, if any one of the threshold(s) for a particular block size is not met, DMVR is not applied.
[0137] iii) Alternatively, it may be possible to use such a threshold to decide whether to apply a single layer, two layers, or multiple layers to the DMVR. For example, if the MVD is within a certain threshold, it may be sufficient to perform the DMVR at a 16x16 level rather than an 8x8 level.
[0138] 3. Also, as an example, when implementing MDMVR, each layer can have the same search pattern or different search patterns.
[0139] 4. Also, as an example, search points may generally vary depending on the search pattern. Therefore, it may be possible and advantageous to initially use a large search that correlates with many search points, and then reduce the search space in subsequent layers to consider smaller search patterns and fewer search points.
[0140] 5. Additionally, the accuracy of the initial search pattern is integer-based, and the accuracy of the additional layers is 1 / 2-pel or 1 / 4-pel.
[0141] 6. Also, by way of example, contemplated search patterns may include, but are not limited to, squares, diamonds, crosses, rectangles, and / or other suitable shapes for capturing basic block movements.
[0142] For example, if the motion vector in the x direction is larger than the motion vector in the y direction, a rectangular search pattern is useful because it can be more adaptable in capturing the basic motion of the block.
[0143] b. Alternatively, a diamond / cross pattern may be more suitable for blocks with a lot of vertical movement.
[0144] 7. Also, as an example, the size of the search area may differ for each block depending on the size of the block.
[0145] For example, if larger blocks are used, it is possible to use 7x7 / 8x8 or larger or other suitable square / diamond search patterns. This can be used with variable DMVR granularity.
[0146] b. Or, if the blocks are smaller, the search area can be smaller than 5x5.
[0147] c. It is also possible to determine the size of the search area based on basic motion information, the motion characteristics of available neighboring blocks. For example, if the block MVD is greater than a threshold T (which can be predetermined), a 7x7 search area can be used.
[0148] 8. Also, as an example, reference samples can be pre-fetched and stored in memory while awaiting processing. For example, when samples are pre-fetched, they can be reference sample padded.
[0149] 9. In VVC, DMVR uses SAD as a means to estimate the refined MV using an iterative process, but several other distortion metrics can be used to estimate the distortion.
[0150] For example, the L0 norm can be used, where the initial and intermediate motion vectors can be used to determine whether there is a motion change in the x or y direction, which can then be used to evaluate whether early termination can be performed.
[0151] b. Alternatively, other forms of norms, such as the Euclidean Norm (L2), can be used to indicate which point has the largest displacement and therefore the outlier.
[0152] c. Alternatively, for example, MR-SAD (Mean Removed-Sum of Absolute Difference) can be used. Various variations of MR-SAD can also be used. For example, the MRSAD of all alternate rows / columns or the cumulative average of the previous block's MRSAD for each block can be used.
[0153] 10. Also, as an example, additional weighting factors can be added to attenuate / amplify distortion measurements. This consideration can be taken into account if distortion at a particular search point is more prevalent than at other search points.
[0154] For example, if an initial search point should be prioritized over other points within the search range, a weighting value can be applied to the initial error / distortion metric so that the initial value(s) generates the minimum distortion cost.
[0155] b. As another example, it is possible to take into account available motion information of neighboring blocks to determine the weighting values to use.
[0156] 11. Also, by way of example, it is possible to facilitate early termination within a single layer or within each layer of the MDMVR.
[0157] For example, the MDMVR can be terminated early if the distance between the initial starting MV and the MV at a point between iterations is less than a threshold T, after which the search can be terminated.
[0158] b. Alternatively, all termination conditions can be verified between layers. For example, a CU can signal to use three layers for DMVR. However, if a termination condition is met after the first layer, the DMVR can be terminated.
[0159] c. Alternatively, other early termination methodologies, such as sample-based difference or SAD-based difference, can also be used, either individually or in combination.
[0160] 12. Also, as an example, the use of single layer DMVR or MDMVR can be signaled with multiple parameter sets.
[0161] For example, the use of MDMVR can be fixed for the entire sequence by signaling a single flag in the SPS along with the associated GCI flag.
[0162] b. Sequences can be converted between DMVR and MDMVR using PPS (Picture Parameter Set) / PH (Picture header) / SH (Slice header) / CU (Coding unit) and / or other appropriate headers.
[0163] i) For example, the use of an existing DMVR or MDMVR can be signaled in the PPS via additional control present in the PH or lower level.
[0164] ii) For example, it is possible to switch between DMVR and MDMVR at the CU level, i.e., each CU can be switched independently.
[0165] c. When a flag or a PTL flag, i.e., a GCI restriction flag, is signaled in the SPS, it can be considered that it can be fixed length coded.
[0166] d. When syntax elements for MDMVR are signaled at the PPS or lower level, it is appropriate to use context coding to signal the necessary detailed information.
[0167] i) For example, information that can be signaled can include, but is not limited to, the number of layers, the granularity of MDMVR application, whether early termination is explicitly used, etc.
[0168] ii) Also, for example, the number of context models, initialization values can be determined taking into account relevant aspects of block statistics.
[0169] Meanwhile, the use of multi-layer DMVR (i.e., MDMVR) as described above can be considered advantageous in improving video quality through compression efficiency. Accordingly, this document proposes a method for efficiently signaling information related to the use of MDMVR. In this regard, examples of structures that can be implemented in a decoder can be shown in Figures 7 and 8.
[0170] FIG. 7 illustrates an exemplary MDMVR structure according to one embodiment of this document.
[0171] In the example of Figure 7, a flag / index can be used to determine whether an MDMVR or an existing DMVR is used (S700). Generally, such a flag (single bin or multiple fixed length coded bins) can be used to indicate whether a DMVR is used (e.g., index 0), whether a single-layer DMVR is used (e.g., index 1), or whether both a single layer and an MDMVR are used (e.g., index 2). For example, if an existing DMVR is used, an existing DMVR processing process is performed (S710). If an MDMVR is used, additional control information can be signaled in a slice header or a picture header (S720 to S740).
[0172] More specifically, the decoding device may acquire information related to whether MDMVR is available (e.g., MDMVR enable flag or index information) and determine whether DMVR is used or not based on the information (S700). The information related to whether MDMVR is available indicates whether MDMVR is enabled and may be signaled in a higher level (e.g., SPS) syntax.
[0173] For example, the decoding device can acquire flag information related to whether MDMVR can be used, and if the value of the flag information is 0, determine that DMVR is used, and if the value of the flag information is 1, determine that MDMVR is used.
[0174] Alternatively, for example, the decoding device may acquire index information related to whether MDMVR is available, and determine whether to use DMVR or MDMVR based on the value of the index information. For example, if the value of the index information is 0, it may determine that DMVR is to be used, and if the value of the index information is 1, it may determine that MDMVR is to be used. Alternatively, as described above, if the value of the index information is 0, it may determine that DMVR is to be used, if the value of the index information is 1, it may determine that single-layer DMVR is to be used, and if the value of the index information is 2, it may determine that multi-layer DMVR (i.e., MDMVR) is to be used.
[0175] If the decoding device determines that the DMVR is to be used based on the information related to the availability of the DMVR, the decoding device can execute the DMVR (S710).
[0176] If the decoding device determines that the MDMVR is to be used based on the information related to the availability of the MDMVR, the decoding device can obtain additional control information (S720).
[0177] The additional control information is MDMVR-related control information at a lower level (e.g., picture header, slice header, etc.) and may indicate, for example, whether an MDMVR-related syntax element is present in the picture header syntax. For example, a value of 0 in the additional control information may indicate that the additional control information (e.g., MDMVR-related syntax element) is not present in the picture header syntax, and a value of 1 in the additional control information may indicate that the additional control information (e.g., MDMVR-related syntax element) is present in the picture header syntax.
[0178] If the value of the additional control information is 0, the decoding device can determine that the MDMVR-related additional control information does not exist in the picture header syntax and can obtain the MDMVR-related information in the slice header (S730).
[0179] The MDMVR-related information in the slice header may indicate whether an MDMVR-related syntax element is present in the slice header. For example, if the value of the MDMVR-related information is 1, it indicates that an MDMVR-related syntax element is present in the slice header, and the MDMVR-related syntax element can be subsequently signaled / parsed from the slice header.
[0180] If the value of the additional control information is 1, the decoding device can determine that the MDMVR-related syntax element exists in the picture header syntax and can obtain the MDMVR-related syntax element from the picture header (S740).
[0181] Thereafter, the decoding device can perform MDMVR based on the MDMVR-related information signaled from the picture header or slice header.
[0182] FIG. 8 exemplarily illustrates an MDMVR structure according to another embodiment of the present document.
[0183] In the example of FIG. 8, MDMVR information can be signaled at the PPS or CU level. DMVR can be performed at the PU level. Here, when MDMVR is enabled, an additional syntax element can be parsed at the PPS to indicate whether the control element exists at the PPS or CU level. Because multiple frames typically refer to one PPS, parsing control information at the PPS can be less flexible than when elements are signaled at the CU. On the other hand, signaling at the PPS level can reduce signaling overhead.
[0184] 8, the decoding device may obtain information related to whether the MDMVR is available (e.g., an MDMVR enable flag or index information) and determine whether the DMVR is used or not based on the information (S800). The information related to whether the MDMVR is available indicates whether the MDMVR is enabled and may be signaled in a higher level (e.g., SPS) syntax.
[0185] For example, the decoding device can acquire flag information related to whether MDMVR can be used, and if the value of the flag information is 0, determine that DMVR is used, and if the value of the flag information is 1, determine that MDMVR is used.
[0186] Alternatively, for example, the decoding device may acquire index information related to whether MDMVR is available, and determine whether to use DMVR or MDMVR based on the value of the index information. For example, if the value of the index information is 0, it may determine that DMVR is to be used, and if the value of the index information is 1, it may determine that MDMVR is to be used. Alternatively, as described above, if the value of the index information is 0, it may determine that DMVR is to be used, if the value of the index information is 1, it may determine that single-layer DMVR is to be used, and if the value of the index information is 2, it may determine that multi-layer DMVR (i.e., MDMVR) is to be used.
[0187] If the decoding device determines that the DMVR is to be used based on the information related to the availability of the DMVR, the decoding device can execute the DMVR (S810).
[0188] In this case, the decoding device can obtain DMVR-related information signaled at the PU level and execute DMVR.
[0189] If the decoding device determines that the MDMVR is to be used based on the information related to the availability of the MDMVR, the decoding device can obtain additional control information (S820).
[0190] The additional control information is MDMVR-related control information at a lower level (e.g., PPS) and may indicate, for example, whether an MDMVR-related syntax element is present in the PPS syntax. For example, a value of 0 in the additional control information may indicate that the additional control information (e.g., MDMVR-related syntax element) is not present in the PPS syntax, and a value of 1 in the additional control information may indicate that the additional control information (e.g., MDMVR-related syntax element) is present in the PPS syntax.
[0191] If the value of the additional control information is 0, the decoding device can determine that the MDMVR-related additional control information does not exist in the PPS syntax and can obtain MDMVR-related information in the CU (S830).
[0192] The MDMVR-related information in the CU may indicate whether an MDMVR-related syntax element exists in the CU syntax. For example, if the value of the MDMVR-related information is 1, it indicates that an MDMVR-related syntax element exists in the CU syntax, and the MDMVR-related syntax element can be subsequently signaled / parsed from the CU syntax.
[0193] If the value of the additional control information is 1, the decoding device can determine that the MDMVR-related additional control information exists in the PPS syntax and can obtain the MDMVR-related additional control information (e.g., MDMVR-related syntax elements) from the PPS (S840).
[0194] Thereafter, the decoding device can perform MDMVR based on MDMVR-related information signaled from the PPS or CU.
[0195] The following drawings are created to explain a specific example of the present document. The names of specific devices and specific terms / names (e.g., names of syntax / syntax elements) shown in the drawings are provided for illustrative purposes only, and the technical features of the present document are not limited to the specific terms / names used in the following drawings.
[0196] 9 and 10 illustrate an example of a video / image encoding method and associated components according to embodiment(s) of the present document.
[0197] The method disclosed in FIG. 9 may be performed by the encoding apparatus 200 disclosed in FIG. 2 or 10. Here, the encoding apparatus 200 disclosed in FIG. 10 is a simplified version of the encoding apparatus 200 disclosed in FIG. 2. Specifically, steps S900 to S920 of FIG. 9 may be performed by the prediction unit 220 disclosed in FIG. 10, step S930 of FIG. 9 may be performed by the residual processing unit 230 disclosed in FIG. 10, and step S940 of FIG. 9 may be performed by the entropy encoding unit 240 disclosed in FIG. 10. Also, although not shown, a process of generating reconstructed samples and reconstructed pictures for the current block based on residual samples and predicted samples for the current block may be performed by the adder 250 of the encoding apparatus 200, and a process of encoding prediction information for the current block may be performed by the entropy encoding unit 240 of the encoding apparatus 200. 9 may be implemented in accordance with the embodiments detailed herein, and therefore, detailed descriptions of the same content as those of the above-described embodiments will be omitted or simplified.
[0198] Referring to FIG. 9, an encoding apparatus may determine whether to use multi-layer Decoder-side Motion Vector Refinement (MDMVR) for a current block (S900).
[0199] As mentioned above, the multi-layer DMVR may be referred to as an MDMVR, and may be used interchangeably with or in place of a multi-path DMVR, a multi-level DMVR, or a multi-step DMVR.
[0200] That is, the encoding device can decide whether to apply DMVR using multiple layers (or multiple passes, multiple levels, multiple steps, etc.) to derive a final refined motion vector for the current block.
[0201] At this time, the encoding apparatus can determine whether to use MDMVR for the current block according to the above-described embodiment.
[0202] In one embodiment, whether MDMVR is enabled may be determined based on information signaled in multiple parameter sets. For example, whether MDMVR is enabled may be determined based on first flag information associated with indicating whether MDMVR is enabled. The first flag information may be signaled in a Sequence Parameter Set (SPS). In this case, whether MDMVR is enabled may be determined for the entire sequence, or additional control information may be signaled at the sequence level, and whether MDMVR is enabled may be determined at a lower level (e.g., Picture Parameter Set (PPS) / Picture Header (PH) / Slice Header (SH) / Coding Unit (CU) and / or other appropriate headers). For example, based on the first flag information signaled in the SPS (i.e., based on the first flag information associated with the use of MDMVR), for example, if the value of the first flag information is 1, second flag information associated with MDMVR control may be signaled. The second flag information is information for controlling whether MDMVR can be used at a lower level, and is information related to whether syntax elements related to MDMVR exist in the PPS, PH, SH, CU, or other appropriate header.
[0203] As a specific example, as described with reference to Figures 7 and 8, second flag information related to MDMVR control at the PPS / PH may be signaled based on first flag information related to whether MDMVR at the SPS level is used (e.g., when the value of the first flag information is 1). In this case, if the second flag information indicates that syntax elements related to MDMVR control at the PPS / PH exist, syntax elements related to MDMVR control from the PPS / PH may be further signaled. Alternatively, if the second flag information indicates that syntax elements related to MDMVR control at the PPS / PH do not exist, syntax elements related to MDMVR at a lower level (e.g., CU, SH) may be signaled.
[0204] Also, as an example, the first flag information signaled at a higher level (e.g., SPS) as described above may be binarized by fixed length coding. Also, when MDMVR-related syntax elements are signaled at a lower level such as PPS, PH, SH, or CU based on the first flag and / or the second flag, the MDMVR-related syntax elements may be derived (i.e., parsed) based on context coding. For example, the MDMVR-related syntax elements may include detailed information necessary for MDMVR execution, such as the number of layers, the granularity of MDMVR application, whether early termination is explicitly used, etc. Also, when performing context coding of the MDMVR-related syntax elements, the number of context models and initialization values may be determined, and this may be determined taking into account related aspects of block statistics.
[0205] In addition, as one embodiment, the current block is a coding unit (CU) including at least one prediction unit (PU), and in this case, the number of MDMVR layers can be determined for at least one PU. For example, MDMVR can be applied using two layers to the first PU in the current block, and MDMVR can be applied using three layers to the second PU in the current block. Alternatively, the number of MDMVR layers can be determined for each CU. In this case, all PUs in the CU can have the same number of layers.
[0206] In addition, in one embodiment, whether to use MDMVR may be determined based on MVD (Motion Vector Difference) information. For example, whether to use DMVR may be determined first based on whether the MVD exceeds a predetermined threshold. For example, if the MVD exceeds a predetermined threshold, it may be determined that DMVR is not to be applied. In this case, the threshold may be selected in 1 / 4 pixel (pel), 1 / 2 pixel (pel), 1 pixel (pel), 4 pixels (pel), and / or other appropriate MVD units. Alternatively, if the threshold cannot be met for a specific block size, DMVR is not applied. Alternatively, for example, it may be determined whether to apply single layer or multiple layers to DMVR based on the threshold. For example, if the MVD is within a specific threshold range, DMVR may be performed at a 16x16 level rather than an 8x8 level. In other words, whether to apply MDMVR to a current block may be determined based on whether the MVD exceeds a predetermined threshold. For example, if the MVD exceeds a predetermined threshold, it may be determined that MDMVR is not to be applied. Also, the number of layers of the MDMVR can be determined based on whether the MVD information is within a predetermined threshold range.
[0207] Also, as an embodiment, based on whether MDMVR is used for the current block (i.e., if MDMVR is performed on the current block), a search pattern and a size of a search area can be determined for each layer of MDMVR.
[0208] For example, each layer may have the same search pattern or different search patterns. Also, for example, the search pattern may include a square, a diamond, a cross, a rectangle, and / or other suitable shapes for capturing the basic motion of the block. In this case, the search pattern may be determined based on the x-direction motion vector and the y-direction motion vector, or based on the vertical motion and the horizontal motion. For example, if the x-direction motion vector is larger than the y-direction motion vector, a rectangular search pattern may be more suitable for capturing the basic motion of the block. Alternatively, for example, a diamond / cross search pattern may be more suitable for blocks with more vertical motion.
[0209] Also, for example, the size of the search area may be determined based on the block size. For example, if large blocks are used, a 7x7 / 8x8 or larger, or other suitable square / diamond search pattern may be used. Alternatively, for example, if small blocks are used, the search area may have a size smaller than 5x5. Alternatively, for example, the size of the search area may be determined based on basic motion information, such as the motion characteristics of available neighboring blocks. For example, if the block MVD is greater than a predetermined threshold T, a 7x7 search area may be used.
[0210] Also, for example, search points may be determined based on a search pattern. For example, if there is an initial correlation with the search points, a larger search pattern may be used. Thereafter, if a smaller search pattern and a smaller number of search points are considered, the search area may be reduced by layer.
[0211] Also, for example, the accuracy of the initial search pattern can be determined on an integer basis, and the accuracy of the additional layers can be determined on a 1 / 2-pel or 1 / 4-pel basis.
[0212] In addition, as an example, the refined motion vector may be derived based on Sum of Absolute Differences (SAD) or Mean Removed-Sum of Absolute Differences (MR-SAD). In other words, the refined motion vector derived by applying MDMVR measures distortion based on SAD or MR-SAD, and the final refined motion vector may be derived based on this. In addition, the L0 norm or Euclidean norm (L2), etc. may be used to derive the refined motion vector.
[0213] In one embodiment, a refined motion vector may be derived based on a minimum distortion cost to which a weighted value is applied. Here, the minimum distortion cost may be calculated by applying a weighted value based on whether the distortion of a specific search point is more important than that of other search points. For example, if an initial search point is more important than other points within the search range, a weighted value may be applied to an initial error / distortion metric to calculate the minimum distortion cost for the initial value(s). Alternatively, for example, the minimum distortion cost may be calculated by determining a weighted value in consideration of available motion information of neighboring blocks.
[0214] In addition, as an embodiment, based on whether MDMVR is used for the current block, it may be possible to determine whether a termination condition is met for each layer of MDMVR and then perform MDMVR. The termination condition may be determined based on whether the distance, sample-based difference, or SAD-based difference between the initial motion vector and the refined motion vector is smaller than a threshold. For example, if the distance between the initial starting MV and the MV at a point between repetitions (i.e., the refined MV) is smaller than a threshold T, it may be determined that the termination condition is met and MDMVR may be terminated. In addition, in determining whether the termination condition is met, it may be checked between layers of MDMVR. For example, if it is determined that the termination condition is met after the first layer of MDMVR, MDMVR may be terminated early without performing MDMVR on the remaining layers.
[0215] In one embodiment, reference samples can be pre-fetched and stored in memory while awaiting processing. For example, if samples are pre-fetched, they can be reference sample padded.
[0216] The encoding apparatus may derive a refined motion vector for the current block based on the MDMVR used for the current block (S910).
[0217] That is, whether MDMVR is to be used for the current block can be determined according to the above-described embodiment(s). If it is determined that MDMVR is to be used, the encoding apparatus can derive a refined motion vector by applying MDMVR to the current block.
[0218] In one embodiment, when inter-prediction is performed on a current block, the encoding apparatus may first derive motion information (such as a motion vector and a reference picture index) of the current block. For example, the encoding apparatus may search for blocks similar to the current block within a certain region (search region) of a reference picture through motion estimation, and derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion. Based on this, the encoding apparatus may derive a reference picture index indicating the reference picture in which the reference block is located, and derive a motion vector based on the positional difference between the reference block and the current block.
[0219] In addition, the encoding apparatus may determine an inter prediction mode to be applied to the current block from among various prediction modes, and may compare RD costs for various prediction modes to determine an optimal prediction mode for the current block.
[0220] Thereafter, as described above, if it is determined to apply MDMVR to the current block, the encoding apparatus can apply MDMVR to the motion vector to derive a final refined motion vector.
[0221] The encoding apparatus may derive a prediction sample for the current block based on the refined motion vector (S920), and may derive a residual sample for the current block based on the prediction sample (S930).
[0222] That is, the encoding apparatus may derive residual samples based on original samples for a current block and predicted samples for the current block, and may generate information about the residual samples, where the information about the residual samples may include information about values of quantized transform coefficients derived by performing transform and quantization on the residual samples, position information, a transform technique, a transform kernel, a quantization parameter, etc.
[0223] The encoding apparatus may encode image information (or video information) (S940). Here, the image information may include prediction-related information (e.g., prediction mode information). The image information may also include the residual information. That is, the image information may include various information derived during the encoding process, and may be encoded including such various information.
[0224] In one embodiment, an encoding apparatus may encode video information including information about residual samples into a bitstream, and may also encode video information including prediction-related information (e.g., prediction mode information) into a bitstream.
[0225] Video information including the above-mentioned various information can be encoded and output in the form of a bitstream. The bitstream can be transmitted to a decoding device via a network or a (digital) storage medium. Here, the network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0226] 11 and 12 show an example of a video / image decoding method and associated components according to embodiment(s) of the present document.
[0227] The method disclosed in Figure 11 may be performed by the decoding apparatus 300 disclosed in Figure 3 or Figure 12. Here, the decoding apparatus 300 disclosed in Figure 12 is a simplified version of the decoding apparatus 300 disclosed in Figure 3. Specifically, step S1100 of Figure 11 may be performed by the entropy decoding unit 310 and / or the prediction unit 330 disclosed in Figure 12, steps S1110 to S1120 of Figure 11 may be performed by the prediction unit 330 disclosed in Figure 12, and step S1130 of Figure 11 may be performed by the addition unit 340 disclosed in Figure 12. Also, although not shown, the process of receiving prediction information and / or residual information for a current block may be performed by the entropy decoding unit 310 of the decoding apparatus 300, and the process of deriving a residual sample for the current block based on the residual information may be performed by the residual processing unit 320 of the decoding apparatus 300. In addition, the method disclosed in Fig. 11 may be implemented by including the embodiments detailed in this document, and therefore, detailed descriptions of the contents overlapping with the above-described embodiments will be omitted or simplified in Fig. 11.
[0228] Referring to FIG. 11, a decoding apparatus may determine whether to use multi-layer Decoder-side Motion Vector Refinement (MDMVR) for a current block (S1100).
[0229] As mentioned above, the multi-layer DMVR may be referred to as an MDMVR, and may be used interchangeably with or in place of a multi-path DMVR, a multi-level DMVR, or a multi-step DMVR.
[0230] That is, the decoding device can determine whether to apply DMVR using multiple layers (or multiple passes, multiple levels, multiple steps, etc.) to derive a final refined motion vector for the current block.
[0231] At this time, the decoding apparatus can determine whether to use MDMVR for the current block according to the above-described embodiment.
[0232] In one embodiment, whether MDMVR is enabled may be determined based on information signaled in multiple parameter sets. For example, whether MDMVR is enabled may be determined based on first flag information associated with indicating whether MDMVR is enabled. The first flag information may be signaled in a Sequence Parameter Set (SPS). In this case, whether MDMVR is enabled may be determined for the entire sequence, or additional control information may be signaled at the sequence level, and whether MDMVR is enabled may be determined at a lower level (e.g., Picture Parameter Set (PPS) / Picture Header (PH) / Slice Header (SH) / Coding Unit (CU) and / or other appropriate headers). For example, based on the first flag information signaled in the SPS (i.e., based on the first flag information associated with the use of MDMVR), for example, if the value of the first flag information is 1, second flag information associated with MDMVR control may be signaled. The second flag information is information for controlling whether MDMVR can be used at a lower level, and is information related to whether syntax elements related to MDMVR exist in the PPS, PH, SH, CU, or other appropriate header.
[0233] As a specific example, as described with reference to Figures 7 and 8, second flag information related to MDMVR control at the PPS / PH may be signaled based on first flag information related to whether MDMVR at the SPS level is used (e.g., when the value of the first flag information is 1). In this case, if the second flag information indicates that syntax elements related to MDMVR control at the PPS / PH exist, syntax elements related to MDMVR control from the PPS / PH may be further signaled. Alternatively, if the second flag information indicates that syntax elements related to MDMVR control at the PPS / PH do not exist, syntax elements related to MDMVR at a lower level (e.g., CU, SH) may be signaled.
[0234] Also, as an example, the first flag information signaled at a higher level (e.g., SPS) as described above may be binarized by fixed length coding. Also, when MDMVR-related syntax elements are signaled at a lower level such as PPS, PH, SH, or CU based on the first flag and / or the second flag, the MDMVR-related syntax elements may be derived (i.e., parsed) based on context coding. For example, the MDMVR-related syntax elements may include detailed information necessary for MDMVR execution, such as the number of layers, the granularity of MDMVR application, whether early termination is explicitly used, etc. Also, when performing context coding of the MDMVR-related syntax elements, the number of context models and initialization values may be determined, and this may be determined taking into account related aspects of block statistics.
[0235] In addition, as one embodiment, the current block is a coding unit (CU) including at least one prediction unit (PU), and in this case, the number of MDMVR layers can be determined for at least one PU. For example, MDMVR can be applied using two layers to the first PU in the current block, and MDMVR can be applied using three layers to the second PU in the current block. Alternatively, the number of MDMVR layers can be determined for each CU. In this case, all PUs in the CU can have the same number of layers.
[0236] In addition, in one embodiment, whether to use MDMVR may be determined based on MVD (Motion Vector Difference) information. For example, whether to use DMVR may be determined first based on whether the MVD exceeds a predetermined threshold. For example, if the MVD exceeds a predetermined threshold, it may be determined that DMVR is not to be applied. In this case, the threshold may be selected in 1 / 4 pixel (pel), 1 / 2 pixel (pel), 1 pixel (pel), 4 pixels (pel), and / or other appropriate MVD units. Alternatively, if the threshold cannot be met for a specific block size, DMVR is not applied. Alternatively, for example, it may be determined whether to apply single layer or multiple layers to DMVR based on the threshold. For example, if the MVD is within a specific threshold range, DMVR may be performed at a 16x16 level rather than an 8x8 level. In other words, whether to apply MDMVR to a current block may be determined based on whether the MVD exceeds a predetermined threshold. For example, if the MVD exceeds a predetermined threshold, it may be determined that MDMVR is not to be applied. Also, the number of layers of the MDMVR can be determined based on whether the MVD information is within a predetermined threshold range.
[0237] Also, as an embodiment, based on whether MDMVR is used for the current block (i.e., if MDMVR is performed on the current block), a search pattern and a size of a search area can be determined for each layer of MDMVR.
[0238] For example, each layer may have the same search pattern or different search patterns. Also, for example, the search pattern may include a square, a diamond, a cross, a rectangle, and / or other suitable shapes for capturing the basic motion of the block. In this case, the search pattern may be determined based on the x-direction motion vector and the y-direction motion vector, or based on the vertical motion and the horizontal motion. For example, if the x-direction motion vector is larger than the y-direction motion vector, a rectangular search pattern may be more suitable for capturing the basic motion of the block. Alternatively, for example, a diamond / cross search pattern may be more suitable for blocks with more vertical motion.
[0239] Also, for example, the size of the search area may be determined based on the block size. For example, if large blocks are used, a 7x7 / 8x8 or larger, or other suitable square / diamond search pattern may be used. Alternatively, for example, if small blocks are used, the search area may have a size smaller than 5x5. Alternatively, for example, the size of the search area may be determined based on basic motion information, such as the motion characteristics of available neighboring blocks. For example, if the block MVD is greater than a predetermined threshold T, a 7x7 search area may be used.
[0240] Also, for example, search points may be determined based on a search pattern. For example, if there is an initial correlation with the search points, a larger search pattern may be used. Thereafter, if a smaller search pattern and a smaller number of search points are considered, the search area may be reduced by layer.
[0241] Also, for example, the accuracy of the initial search pattern can be determined on an integer basis, and the accuracy of the additional layers can be determined on a 1 / 2-pel or 1 / 4-pel basis.
[0242] In addition, as an example, the refined motion vector may be derived based on Sum of Absolute Differences (SAD) or Mean Removed-Sum of Absolute Differences (MR-SAD). In other words, the refined motion vector derived by applying MDMVR measures distortion based on SAD or MR-SAD, and the final refined motion vector may be derived based on this. In addition, the L0 norm or Euclidean norm (L2), etc. may be used to derive the refined motion vector.
[0243] In one embodiment, a refined motion vector may be derived based on a minimum distortion cost to which a weighted value is applied. Here, the minimum distortion cost may be calculated by applying a weighted value based on whether the distortion of a specific search point is more important than that of other search points. For example, if an initial search point is more important than other points within the search range, a weighted value may be applied to an initial error / distortion metric to calculate the minimum distortion cost for the initial value(s). Alternatively, for example, the minimum distortion cost may be calculated by determining a weighted value in consideration of available motion information of neighboring blocks.
[0244] In one embodiment, based on whether MDMVR is used for the current block, it may be determined whether a termination condition is met for each layer of MDMVR, and then MDMVR may be performed. The termination condition may be determined based on whether the distance, sample-based difference, or SAD-based difference between the initial motion vector and the refined motion vector is smaller than a threshold. For example, if the distance between the initial starting MV and the MV at a point between repetitions (i.e., the refined MV) is smaller than a threshold T, it may be determined that the termination condition is met, and MDMVR may be terminated. In addition, in determining whether the termination condition is met, it may be checked between layers of MDMVR. For example, if it is determined that the termination condition is met after the first layer of MDMVR, MDMVR may be terminated early without performing MDMVR on remaining layers.
[0245] In one embodiment, reference samples can be pre-fetched and stored in memory while awaiting processing. For example, if samples are pre-fetched, they can be reference sample padded.
[0246] The decoding apparatus can derive a refined motion vector for the current block based on the MDMVR used for the current block (S1110).
[0247] That is, whether MDMVR is used for the current block can be determined according to the above-described embodiment(s). If it is determined that MDMVR is used, the decoding apparatus can derive a refined motion vector by applying MDMVR to the current block.
[0248] In one embodiment, a decoding device may first obtain video information including prediction-related information from a bitstream and determine a prediction mode for a current block based on the prediction-related information. Then, the decoding device may derive motion information (motion vector, reference picture index, etc.) of the current block based on the prediction mode. Here, the prediction mode may include skip mode, merge mode, (A)MVP mode, etc.
[0249] Thereafter, as described above, if it is determined to apply MDMVR to the current block, the decoding apparatus can apply MDMVR to the motion vector to derive a final refined motion vector.
[0250] The decoding apparatus may derive prediction samples for the current block based on the refined motion vector (S1120), and may generate reconstructed samples for the current block based on the prediction samples (S1130).
[0251] In one embodiment, the decoding apparatus may directly use the predicted samples as reconstructed samples according to a prediction mode, or may generate reconstructed samples by adding residual samples to the predicted samples.
[0252] If residual samples for the current block exist, the decoding apparatus may receive information about the residuals for the current block. The information about the residuals may include transform coefficients related to the residual samples. The decoding apparatus may derive residual samples (or residual sample arrays) for the current block based on the residual information. Specifically, the decoding apparatus may derive quantized transform coefficients based on the residual information. The quantized transform coefficients may have a one-dimensional vector form based on a coefficient scanning order. The decoding apparatus may derive transform coefficients based on a dequantization procedure for the quantized transform coefficients. The decoding apparatus may derive residual samples based on the transform coefficients.
[0253] The decoding apparatus may generate reconstructed samples based on predicted samples and residual samples, and derive reconstructed blocks or pictures based on the reconstructed samples. Specifically, the decoding apparatus may generate reconstructed samples based on the sum of predicted samples and residual samples. As described above, the decoding apparatus may then apply an in-loop filtering procedure, such as deblocking filtering and / or an SAO procedure, to the reconstructed pictures as needed to improve subjective / objective image quality.
[0254] In the above-described embodiments, the method is described based on a flow chart with a series of steps or blocks, but the embodiments herein are not limited to the order of steps, and certain steps may occur in a different order or simultaneously with other steps than those described. Furthermore, those skilled in the art will understand that the steps shown in the flow charts are not exclusive, and other steps may be included, or one or more steps in the flow charts may be deleted without affecting the scope of this document.
[0255] The method according to the present document described above can be implemented in software form, and the encoding device and / or decoding device according to the present document can be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, or a display device.
[0256] In this document, when an embodiment is implemented in software, the method described above may be implemented with modules (processes, functions, etc.) that perform the functions described above. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described herein may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information on instructions) or algorithms for implementation may be stored on a digital storage medium.
[0257] In addition, the decoding device and encoding device to which the embodiment(s) of this document are applied may be included in a multimedia broadcast transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video interaction device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a custom video (VoD) service providing device, an over-the-top (OTT) video (over-the-top) device, an internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, an image telephone video device, a vehicle terminal (e.g., a vehicle terminal (including an autonomous vehicle), an airplane terminal, a ship terminal, etc.), a medical video device, etc., and may be used to process a video signal or a data signal. For example, an over-the-top (OTT) video (over-the-top) device may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0258] In addition, a processing method to which the embodiment(s) of this document is applied may be produced in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of this document may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. Examples of the computer-readable recording medium include Blu-ray Discs (BDs), Universal Serial Buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission via the Internet). A bitstream generated by the encoding method may be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0259] Furthermore, the embodiment(s) of this document may be embodied in a computer program product by program code, which may be executed by a computer in accordance with the embodiment(s) of this document, and which may be stored on a computer-readable carrier.
[0260] FIG. 13 illustrates an example of a content streaming system in which the embodiments disclosed herein can be applied.
[0261] Referring to FIG. 13, a content streaming system to which the embodiments of this document are applied may broadly include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.
[0262] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0263] The bitstream can be generated by an encoding method or a bitstream generation method to which an embodiment of this document is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0264] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. The content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.
[0265] The streaming server can receive content from a media repository and / or an encoding server. For example, if content is received from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.
[0266] Examples of the user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, and head-mounted displays (HMDs)), digital TVs, desktop computers, and digital signage.
[0267] Each server in the content streaming system can be operated as a distributed server, and in this case, data received by each server can be processed in a distributed manner.
[0268] The claims herein may be combined in various ways. For example, the technical features of the method claims herein may be combined and embodied in an apparatus, and the technical features of the apparatus claims herein may be combined and embodied in a method. Furthermore, the technical features of the method claims herein and the technical features of the apparatus claims herein may be combined and embodied in an apparatus, and the technical features of the method claims herein and the technical features of the apparatus claims herein may be combined and embodied in a method.
Claims
1. A video decoding method performed by a decoding device, comprising: determining whether multi-layer decoder-side motion vector refinement (MDMVR) is used for the current block; deriving a refined motion vector for the current block based on the MDMVR used for the current block; deriving prediction samples for the current block based on the refined motion vector; generating reconstructed samples for the current block based on the predicted samples; Whether the MDMVR is to be used is determined based on first flag information related to whether the MDMVR is to be used; The video decoding method, wherein the first flag information is signaled by a Sequence Parameter Set (SPS).
2. second flag information related to control of the MDMVR is signaled based on the first flag information related to use of the MDMVR; The video decoding method of claim 1, wherein the second flag information is information related to whether a syntax element related to the MDMVR exists in a PPS (Picture Parameter Set), a PH (Picture header), a SH (Slice header), or a CU (Coding unit).
3. The video decoding method of claim 1 , wherein the first flag information is binarized by fixed length coding.
4. Based on the second flag information, the syntax element associated with the MDMVR is signaled in the PPS, the PH, the SH, or the CU; The video decoding method of claim 2 , wherein the syntax elements associated with the MDMVR are derived based on context coding.
5. The current block includes at least one prediction unit (PU), The video decoding method of claim 1 , wherein the number of layers of the MDMVR is determined for the at least one PU.
6. Whether the MDMVR is used is determined based on MVD (Motion Vector Difference) information, The video decoding method of claim 1 , wherein whether the MDMVR is applied to the current block is determined based on whether the MVD information exceeds a predetermined threshold.
7. The video decoding method of claim 6 , wherein the number of layers of the MDMVR is determined based on whether the MVD information is within a predetermined threshold range.
8. A search pattern and a size of a search area are determined for each layer of the MDMVR based on the MDMVR being used for the current block; the search pattern is determined based on an x-direction motion vector and a y-direction motion vector, or based on a vertical motion and a horizontal motion; The video decoding method of claim 1 , wherein the size of the search area is determined based on a block size.
9. The video decoding method of claim 1 , wherein the refined motion vector is derived based on a sum of absolute differences (SAD) or a mean removed-sum of absolute differences (MR-SAD).
10. the refined motion vector is derived based on a weighted minimum distortion cost; The image decoding method of claim 1 , wherein the minimum distortion cost is calculated by applying the weighting value based on whether distortion of a particular search point is prioritized.
11. determining whether a termination condition is met for each layer of the MDMVR based on the MDMVR being used for the current block; 2. The video decoding method of claim 1, wherein the termination condition determines termination of MDMVR for each layer based on whether a distance, a sample-based difference, or a SAD-based difference between an initial motion vector and a refined motion vector is smaller than a threshold.
12. A video encoding method performed by an encoding device, comprising: determining whether multi-layer decoder-side motion vector refinement (MDMVR) is used for the current block; deriving a refined motion vector for the current block based on the MDMVR used for the current block; deriving prediction samples for the current block based on the refined motion vector; deriving a residual sample based on the predicted sample; generating a bitstream by encoding video information including information about the residual samples; A value of first flag information related to whether the MDMVR is used is determined based on whether the MDMVR is used; The video encoding method, wherein the first flag information is signaled by a Sequence Parameter Set (SPS).
13. 1. A method for transmitting data including a bitstream of video information, comprising: obtaining the bitstream of the video information, the bitstream comprising: determining whether multi-layer decoder-side motion vector refinement (MDMVR) is used for the current block; deriving a refined motion vector for the current block based on the MDMVR used for the current block; deriving prediction samples for the current block based on the refined motion vector; deriving a residual sample based on the predicted sample; encoding video information including information about the residual samples; transmitting the data including the bitstream; A value of first flag information related to whether the MDMVR is used is determined based on whether the MDMVR is used; A transmission method, wherein the first flag information is signaled by a sequence parameter set (SPS).