Joint MVD-based video coding method and apparatus therefor

WO2026205916A1PCT designated stage Publication Date: 2026-10-01LX SEMICON CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/004597
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-23
Publication Date
2026-10-01

Smart Images

  • Figure KR2026004597_01102026_PF_FP_ABST
    Figure KR2026004597_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A video decoding method according to an embodiment of the present disclosure comprises the steps of: obtaining prediction-related information including joint MVD information of a current block; deriving, on the basis of the prediction-related information, MVP0 for reference frame 0 and MVP1 for reference frame 1 of the current block; deriving a joint MVD of the current block on the basis of the joint MVD information; deriving MVD0 and MVD1 of the current block on the basis of the joint MVD; deriving MV0 of the current block on the basis of the MVP0 and the MVD0; deriving MV1 of the current block on the basis of the MVP1 and the MVD1; and deriving a prediction sample of the current block on the basis of the MV0 and the MV1.
Need to check novelty before this filing date? Find Prior Art

Description

Joint MVD-based image coding method and apparatus

[0001] The present disclosure relates to an image / video coding method and an apparatus thereof.

[0002] As the utilization of multimedia data increases, the need for efficient video compression technology is growing. Video compression technology is essential for the efficient transmission of high-quality video data within limited network bandwidth, and to this end, various video codec technologies have been developed.

[0003] Video codec technologies include MPEG-2, H.264 / AVC, H.265 / HEVC, H.266 / VVC, AV1 (AOMedia video 1), and are expected to be widely used in various application fields such as internet-based video streaming, video calls, virtual reality (VR), and augmented reality (AR).

[0004] As video becomes higher in resolution and quality, its data size increases, and consequently, the amount of information or bits transmitted rises. Consequently, transmitting video data using existing wired or wireless broadband lines or storing it on conventional storage media leads to increased transmission and storage costs. To this end, there is a growing need for subsequent codec technologies, such as H.267 and AV2, to provide improved compression efficiency and video quality. In other words, high-efficiency video compression technology is required to effectively compress, transmit, store, and play back high-resolution, high-quality video information.

[0005] According to one embodiment of the present disclosure, a method and apparatus for increasing video / image coding efficiency are provided.

[0006] According to one embodiment of the present disclosure, a method and apparatus for providing next-generation high-definition video services are provided.

[0007] According to one embodiment of the present disclosure, a method and apparatus for providing improved performance in a real-time streaming environment are provided.

[0008] According to one embodiment of the present disclosure, an inter-prediction-based video / image coding method and apparatus are provided.

[0009] According to one embodiment of the present disclosure, an image decoding method performed by a decoding device is provided. The method is characterized by comprising the steps of: obtaining prediction-related information including joint motion vector difference (MVD) information of a current block; deriving an MVP0 (motion vector predictor 0) for reference frame 0 and an MVP1 for reference frame 1 of the current block based on the prediction-related information; deriving a joint motion vector difference (MVD) of the current block based on the joint MVD information; deriving MVD0 and MVD1 of the current block based on the joint MVD; deriving an MV0 (motion vector 0) of the current block based on the MVP0 and MVD0; deriving an MV1 of the current block based on the MVP1 and MVD1; and deriving a prediction sample of the current block based on the MV0 and MV1.

[0010] According to one embodiment of the present disclosure, an image encoding method is provided by an encoding device. The method is characterized by comprising the steps of: deriving an MVP0 (motion vector predictor 0) for reference frame 0 of a current block and an MVP1 for reference frame 1; deriving an MVD0 and an MVD1 of the current block based on a joint MVD (joint motion vector difference) of the current block; deriving an MV0 (motion vector 0) of the current block based on the MVP0 and the MVD0; deriving an MV1 of the current block based on the MVP1 and the MVD1; deriving a prediction sample of the current block based on the MV0 and the MV1; generating prediction-related information including joint MVD information of the current block; and encoding image information including the prediction-related information.

[0011] According to one embodiment of the present disclosure, a decoding device for image decoding is provided. The decoding device comprises a memory and at least one processor connected to the memory, wherein the at least one processor is configured to perform the following steps: acquiring prediction-related information including joint motion vector difference (MVD) information of a current block; deriving an MVP0 (motion vector predictor 0) for reference frame 0 and an MVP1 for reference frame 1 of the current block based on the prediction-related information; deriving a joint motion vector difference (MVD) of the current block based on the joint MVD information; deriving MVD0 and MVD1 of the current block based on the joint MVD; deriving an MV0 (motion vector 0) of the current block based on the MVP0 and MVD0; deriving an MV1 of the current block based on the MVP1 and MVD1; and deriving a prediction sample of the current block based on the MV0 and MV1.

[0012] According to one embodiment of the present disclosure, an encoding device for video encoding is provided. The encoding device comprises a memory and at least one processor connected to the memory, wherein the at least one processor is configured to perform the steps of: deriving an MVP0 (motion vector predictor 0) for reference frame 0 of a current block and an MVP1 for reference frame 1; deriving an MVD0 and an MVD1 of the current block based on a joint MVD (joint motion vector difference) of the current block; deriving an MV0 (motion vector 0) of the current block based on the MVP0 and the MVD0; deriving an MV1 of the current block based on the MVP1 and the MVD1; deriving a prediction sample of the current block based on the MV0 and the MV1; generating prediction-related information including joint MVD information of the current block; and encoding video information including the prediction-related information.

[0013] According to one embodiment of the present disclosure, a method for storing or transmitting video / image data including a bitstream generated according to a video / image encoding method according to at least one of the embodiments of the present disclosure is provided.

[0014] According to one embodiment of the present disclosure, an apparatus for storing or transmitting video / image data including a bitstream generated according to a video / image encoding method according to at least one of the embodiments of the present disclosure is provided.

[0015] According to one embodiment of the present disclosure, a computer-readable storage medium may be provided that stores a program for performing a method according to at least one of the embodiments of the present disclosure.

[0016] According to one embodiment of the present disclosure, a computer-readable digital storage medium is provided that stores encoded video / image information generated according to a video / image encoding method according to at least one of the embodiments of the present disclosure.

[0017] According to one embodiment of the present disclosure, a computer-readable digital storage medium is provided that stores encoded information or encoded video / image information, which causes a video / image decoding method according to at least one of the embodiments of the present disclosure to be performed by a decoding device.

[0018] According to one embodiment of the present disclosure, overall video / image compression efficiency can be increased.

[0019] According to one embodiment of the present disclosure, the prediction performance for the current block can be improved.

[0020] According to one embodiment of the present disclosure, the MVD signaling efficiency of the current block can be increased.

[0021] FIG. 1 schematically illustrates an example of a video / image coding system to which embodiments of the present disclosure may be applied.

[0022] FIG. 2 is a diagram schematically illustrating the configuration of a video / image encoding device to which embodiments of the present disclosure can be applied.

[0023] FIG. 3 is a diagram schematically illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure can be applied.

[0024] Figure 4 illustrates an exemplary inter prediction procedure.

[0025] Figure 5 shows examples of inter-prediction-based video / image encoding methods.

[0026] Figure 6 shows examples of inter-prediction-based video / image decoding methods.

[0027] Figure 7 shows an example of a method for determining the inter-prediction mode / type of the current block.

[0028] FIG. 8 shows an example of performing inter prediction based on the TIP reference frame of the current frame.

[0029] FIG. 9 shows an embodiment of configuring a reference frame list based on cost.

[0030] FIG. 10 shows an example of deriving MVDs for reference frames based on a joint MVD.

[0031] FIG. 11 schematically illustrates a video / image encoding method according to the embodiment(s) of the present disclosure.

[0032] FIG. 12 schematically illustrates a video / image decoding method according to the embodiment(s) of the present disclosure.

[0033] As the present disclosure is subject to various modifications and may have various embodiments, specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the embodiments of the present disclosure to specific embodiments. Terms used in the present disclosure are used merely to describe specific embodiments and are not intended to limit the technical scope of the present disclosure. The singular forms used in the present disclosure are intended to include the plural forms unless the context clearly indicates otherwise. The term "and / or" used in the present disclosure includes any one or more combinations of the related listing items. The terms "comprising," "composing," and "holding" used in this specification specify the presence of the stated features, numbers, actions, elements, components, and / or combinations thereof, but do not exclude the presence or addition of one or more other features, numbers, actions, elements, components, and / or combinations thereof. In this disclosure, the use of the term “may” in relation to examples or embodiments (e.g., what an example or embodiment may include or implement) means that there exists at least one example or embodiment in which such feature is included or implemented, but not all examples are limited thereto, and such feature or configuration may be omitted.

[0034] Meanwhile, each component in the drawings described in this disclosure is depicted independently for the convenience of explaining different characteristic functions and does not imply that each component is implemented in separate hardware or separate software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of this disclosure as long as they do not deviate from the essence of this disclosure.

[0035] In the present disclosure, "A or B" may mean "only A," "only B," or "both A and B." Alternatively, in the present disclosure, "A or B" may be interpreted as "A and / or B." For example, in the present disclosure, "A, B or C" may mean "only A," "only B," "only C," or "any combination of A, B and C."

[0036] A slash ( / ) or a comma used in the present disclosure may mean "and / or." For example, "A / B" may mean "A and / or B." Accordingly, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B or C."

[0037] In the present disclosure, "at least one of A and B" may mean "only A," "only B," or "both A and B." Additionally, in the present disclosure, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted as synonymous with "at least one of A and B."

[0038] Additionally, in the present disclosure, "at least one of A, B and C" may mean "only A," "only B," "only C," or "any combination of A, B and C." Additionally, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."

[0039] Additionally, parentheses used in the present disclosure may mean "for example." Specifically, when indicated as "prediction (intra-prediction)," "intra-prediction" may be proposed as an example of "prediction." In other words, the "prediction" of the present disclosure is not limited to "intra-prediction," and "intra-prediction" may be proposed as an example of "prediction." Furthermore, even when indicated as "prediction (i.e., intra-prediction)," "intra-prediction" may be proposed as an example of "prediction."

[0040] Technical features described individually within one drawing in this disclosure may be implemented individually or simultaneously.

[0041] The present disclosure relates to video / image coding. For example, the methods / exemplars described in the present disclosure may be applied to methods disclosed in the AV2 (AOMedia Video 2) standard. Additionally, the methods / exemplars disclosed in this disclosure may be applied to methods disclosed in the ECM (enhanced compression model) or H.267 standard, or next-generation video / image coding standards (e.g., H.268, H.269, etc.).

[0042] In the present disclosure, coding may include encoding and / or decoding. In the present disclosure, image coding may be used interchangeably with video coding.

[0043] In the present disclosure, "video" may refer to a set of a series of images over time. "Frame" generally refers to a unit representing a single image at a specific time, and "slice" or "tile" is a unit that constitutes a part of a frame in coding. A slice or tile may include one or more superblocks. A single frame may be composed of one or more slices or tiles. A tile may represent a rectangular area of ​​superblocks within a specific tile row or a specific tile column within a frame. Meanwhile, a single frame may be divided into two or more subframes.

[0044] A pixel or pel can refer to the smallest unit that constitutes a single frame (or image). Additionally, the term 'sample' may be used as a counterpart to pixel. Generally, a sample can represent a pixel or its value, and it may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component.

[0045] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a frame and information related to that region. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.

[0046] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the attached drawings. Hereinafter, the same reference numerals may be used for identical components in the drawings, and redundant descriptions of identical components may be omitted.

[0047] FIG. 1 schematically illustrates an example of a video / image coding system to which embodiments of the present disclosure may be applied.

[0048] Referring to FIG. 1, a video / image coding system may include a first device (encoding device) and a second device (decoding device). The first device may transmit encoded video / image information or data to the second device in the form of a file or streaming via a digital storage medium or a network.

[0049] The video / image coding system may further include a video / image acquisition device and a video / image renderer. The video / image acquisition device may be included in the encoding device, or it may be configured as a separate device or an external component. The video / image renderer may be included in the decoding device, or it may be configured as a separate device or an external component.

[0050] The first device may include a transmission unit as an internal component, or it may include a separate device or an external component.

[0051] The second device may include a receiver as an internal component, or it may include a separate device or an external component.

[0052] The encoder may be called an encoding device, and the decoder may be called a decoding device. The transmission unit may be included in the encoding device. The reception unit may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.

[0053] The decoding device and encoding device to which the embodiment(s) of the present disclosure are applied may be included in multimedia broadcasting transmission and reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, Video on Demand (VoD) service providers, Over-the-top video (OTT) devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, video phone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video devices, and may be used to process video signals or data signals. For example, Over-the-top video (OTT) devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, Digital Video Recorders (DVRs), etc.

[0054] A video / image acquisition device can acquire a video / image source. The video / image acquisition device can acquire video / image through processes such as video / image capture, synthesis, or generation. The video / image acquisition device may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / image, etc. The video / image generation device may include, for example, a camcorder, a computer, a tablet, and a smartphone, etc., and can generate video / image (electronically). For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data. The video / image source may perform a video / image preprocessing process to input the optimized video / image into an encoder.

[0055] An encoding device can encode an input video / image. The encoding device can encode the input video / image through the encoding method presented in this disclosure. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0056] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a network in the form of a file or streaming. The encoded video / image information or data output in the form of a bitstream may also be transmitted to the receiving unit through a streaming server. Digital storage media may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission through a broadcasting / communication network. The receiving unit may receive / extract the bitstream and transmit it to a decoding device. In this disclosure, the transmission unit may be referred to as a transmission device, and the receiving unit may be referred to as a receiving device. As video becomes high-resolution and high-quality, the size of the original video / image data increases, and by acquiring and storing / transmitting a bitstream (or data including a bitstream) generated through an efficient encoding method according to this disclosure, storage / transmission efficiency can be increased and low-latency / real-time transmission can be supported.

[0057] The streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream. The streaming server transmits multimedia data to a user device based on a user request through a web server, and the web server acts as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server forwards this request to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server plays the role of controlling commands and responses between each device within the content streaming system.

[0058] The streaming server can receive content from a media storage and / or encoding device. For example, when receiving content from the encoding device, the content can be received in real time. In this case, to provide a seamless streaming service, the streaming server can store the bitstream for a certain period of time.

[0059] A decoding device can decode a video / image by performing a series of procedures, such as inverse quantization, inverse transform, and prediction, corresponding to the operation of an encoding device. The decoding device can decode a video / image through the decoding method presented in this disclosure.

[0060] The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit.

[0061] FIG. 2 is a diagram schematically illustrating the configuration of a video / image encoding device to which embodiments of the present disclosure may be applied. The term "encoding device" below may include an image encoding device and / or a video encoding device.

[0062] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-predictor and an intra-predictor. The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit (250) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) may be configured by one or more hardware components (e.g., an encoder chipset or processor) according to the embodiment. Additionally, the memory (270) may include a frame buffer and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.

[0063] The image segmentation unit (210) can divide an input image (or picture, frame) input to the encoding device (200) into one or more processing units. For example, the processing unit may include a superblock or a coding block. A single frame may be divided into a plurality of tiles. The tiles may be rectangular in shape. Uniform or non-uniform tile sizes may be determined on a frame-by-frame basis. In this disclosure, the terms frame and picture may be used interchangeably. A tile may consist of an integer number of superblocks. Superblocks within a tile may be coded in raster scan order. A superblock may be divided into one or more coding blocks. For example, based on the luminance component, the size of the superblock may be 128x128 or 64x64. Alternatively, based on the luminance component, the size of the superblock may include 256x256. A superblock may have dependencies only on specific surrounding superblocks. For example, a superblock may have dependencies only on the surrounding superblocks to the left and / or above.

[0064] A superblock can be recursively partitioned. For example, a superblock can be derived into a single coding block, or the superblock can be partitioned into coding blocks based on a binary tree, a terminal tree, or a quad tree. A single coding block can be recursively partitioned into multiple coding blocks of a deeper depth based on a binary tree, a terminal tree, or a quad tree. For example, recursive partitioning may be possible when a coding block is partitioned into PARTITION_NxN. A coding procedure according to the present disclosure may be performed based on a final coding block that is no longer partitioned. In this case, based on coding efficiency according to image characteristics, the superblock may be used directly as the final coding block, or, if necessary, the coding block may be recursively partitioned into coding blocks of a deeper depth so that a coding block of an optimal size is used as the final coding block. Here, the term "coding procedure" may include procedures such as prediction, transformation, and restoration described below. Meanwhile, the processing unit may further include a prediction block (PB) or a transform block (TB). In this case, the prediction block or the transform block may be divided or partitioned from the final coding block described above. For example, a single transform block of the same size as the coding block may be derived, or multiple transform blocks may be derived from the coding block based on a quad tree or binary tree. The transform block may also be recursively divided into transform blocks of a deeper depth. The prediction block may be a unit for prediction (or for deriving a prediction mode), and the transform block may be a unit for performing a transformation, a unit for deriving transformation coefficients, and / or a unit for deriving a residual signal from transformation coefficients.For example, the derivation of a prediction mode for intra-prediction can be performed at the level of the coding block or the prediction block, and the derivation of a prediction sample through the intra-prediction procedure can be performed at the level of a transformation block. According to the coding order, the block currently subject to processing may be called the current block.

[0065] The term "block" may be used interchangeably with terms such as "unit" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. The term "sample" may be used as a counterpart to a pixel or pel of a frame (or image).

[0066] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting the predicted signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit (232). In this case, as illustrated, the configuration for subtracting the predicted signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoder (200) may be called a subtraction unit (231). The prediction unit performs a prediction for a block to be processed (hereinafter referred to as the current block) and can generate a predicted block containing prediction samples for the current block. The prediction unit can determine whether intra prediction is applied or inter prediction is applied on a current block or coding block basis. The prediction unit can generate various information regarding prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit it to the entropy encoding unit (240). The information regarding prediction can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0067] The intra prediction unit can predict the current block by referencing samples within the current frame. Depending on the prediction mode, the referenced samples may be located next to the current block or apart from it. Intra prediction can be performed on a per-transform block basis. If multiple transform blocks exist within a coding block, intra prediction can be performed sequentially in the raster order of the transform blocks. In this case, the procedure for deriving neighboring reference samples for intra prediction can be performed based on the transform block. Multiple prediction modes may be considered for intra prediction. These prediction modes may include multiple non-directional modes and multiple directional modes. The prediction mode used in the current block may be signaled from the encoding device to the decoding device; for example, the prediction modes may include a DC intra prediction mode, multiple directional intra prediction modes, multiple SMOOTH intra prediction modes, and / or PAETH intra prediction modes. These prediction modes may include an intra block copy (intrabc) mode. Whether the above-mentioned intra-block copy mode is applied can be signaled separately. Directional prediction modes may include, for example, eight or more prediction modes depending on the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional prediction modes may be used. The intra-prediction unit may determine the prediction mode applied to the current block based on the prediction mode applied to surrounding blocks. In the intra-block copy mode, a reference block is derived based on a vector, similar to the inter-prediction mode described later, and the current frame is used as the reference frame. The vector used to derive the reference block in the above-mentioned intra-block copy may be called a block vector.

[0068] The inter-prediction unit can derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on a reference frame. In this case, to reduce the amount of motion information transmitted in the inter-prediction mode, motion information can be predicted in block, sub-block, or sample units based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and / or a reference frame index. The motion information may further include information on the inter-prediction direction (L0 prediction, L1 prediction, compound prediction, etc.). In the case of inter-prediction, neighboring blocks may include spatial neighboring blocks existing within the current frame and temporal neighboring blocks existing in the reference frame. The reference frame containing the reference blocks and the reference frame containing the temporal neighboring blocks may be the same or different. The above temporal surrounding blocks may be referred to by names such as co-located reference blocks or colblocks, and a reference frame containing the above temporal surrounding blocks may be referred to as a co-located frame or colframe. For example, the inter-prediction unit may construct a motion information stack based on the surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference frame index of the current block. The motion information stack may be referred to as a motion information list. The motion information stack may include a motion vector stack. The motion vector stack may be referred to as RefStackMv. The motion vector stack may include eight or more candidates.Information regarding the (maximum) number of candidates for the above motion vector stack can be signaled on a frame or sequence basis. Motion modes may be further considered for inter-prediction. The above motion modes may include simple mode, OBMC (overlapped block motion compensation) mode and / or local warp mode. In OBMC mode, prediction performance can be improved by utilizing motion information of surrounding blocks for the left and / or upper boundaries of the current block, and in local warp mode, an affine model may be applied in addition to translational motion compensation.

[0069] Motion vectors can be derived based on various prediction modes; for example, in NEWMV mode, the motion vector of the current block can be indicated based on the motion vector of the reference stack and the motion vector difference. In some cases, in NEWMV mode, the motion vector of the current block can be indicated based on the motion vector difference without the motion vector of the reference stack. Information regarding the motion vector difference can be generated by an encoding device and signaled to a decoding device. ZEROMV mode may indicate that a zero vector or a default vector is used as the motion vector of the current block. REFMV mode may indicate that the motion vector of the motion information stack is used as the motion vector of the current block. However, these names are merely examples, and the motion vector of the current block may be indicated by various other names. For example, GLOBALMV mode may indicate that global motion information is used for the current block. For example, NEARSTMV mode may indicate that the first candidate (candidate index 0) of the motion information stack is used for the current block. For example, the NEARMV mode may indicate that a specific candidate of the motion information stack (a candidate indicated by refMVidx) is used in the current block. However, the above mode names are examples, and depending on the case, other names such as the first mode, second mode, etc. may be used in this disclosure.

[0070] When compound prediction is applied, inter-prediction can be performed using both reference frame lists L0 and L1. In this case, a motion vector for the L0 direction and a motion vector for the L1 direction can be derived, respectively. Additionally, an L0 reference frame for the L0 direction and an L1 reference frame for the L1 direction can be derived, respectively. In some cases, a first motion vector and a second motion vector may be used instead of the L0 motion vector and the L1 motion vector. In some cases, a first reference frame and a second reference frame may be used instead of the L0 reference frame and the L1 reference frame.

[0071] The prediction unit (220) can generate a prediction signal based on various prediction methods. For example, the prediction unit may apply intra prediction or inter prediction for the prediction of a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called compound inter-intra prediction. In this case, the mode used as the intra prediction mode may include DC prediction mode, vertical prediction mode, horizontal prediction mode, and SMOOTH prediction mode. Additionally, the prediction unit may be based on the intra block copy mode described above or on the palette mode for the prediction of a block. As described above, the intra block copy mode is a type of intra prediction that basically performs prediction within the current frame, but can be performed similarly to inter prediction in that it derives a reference block based on a vector using the current frame as a reference frame. However, since the current frame is used as a reference frame, the term block vector may be used instead of motion vector for the vector for movement. That is, the intra-block copy mode may utilize at least one of the inter-prediction techniques described in this disclosure. In this case, a motion vector derived through a surrounding block or a motion information stack may be referenced to derive the block vector of the current block. For example, when the intra-block copy mode is applied, the NEWMV mode described above may be used to derive the block vector of the current block.

[0072] The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or to generate a residual signal. The transformation unit (232) can generate transform coefficients by applying a transformation technique to the residual signal. For example, the transformation technique may include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Additionally, for example, the transformation technique may include DCT, ADST (Asymmetric Discrete Sine Transform), FLIPADST (Flipped ADST), IDTX (Identity Transform), WHT (Walsh-Hadamard Transform), V_DCT, H_DCT, etc. DCT is one of the most widely used transformation techniques in image compression and primarily serves to convert the image signal into the frequency domain to concentrate energy on low-frequency components. ADST is similar to DCT but is a transformation method designed to ensure smooth signal connection at block boundaries. FLIPADST is a variation of ADST that applies the transformation direction in reverse, allowing for more efficient signal compression in specific block patterns. IDTX is an identity transformation method that preserves the signal without transformation. Identity transformation can be applied to only one of the vertical or horizontal transformations, or to both. For example, if a transformation method is specified for only one of the vertical or horizontal transformations, identity transformation may be implicitly applied to the transformation in the other direction.WHT is a linear transformation similar to the Discrete Fourier Transform (DFT) or Discrete Cosine Transform (DCT) that represents signals or data by transforming them into different bases. WHT does not use trigonometric functions (sine and cosine); instead, it can use orthogonal basis matrices composed of +1 and -1. V_DCT and H_DCT indicate that the DCT transformation is performed on the vertical direction of the block and the horizontal direction of the block, respectively. For example, DCT can be used to increase compression ratios in simple blocks, ADST or FLIPADST can be used when smooth connections are required at boundaries, and IDTX can be applied to blocks where the signal remains almost unchanged.

[0073] The quantization unit (233) quantizes the transformation coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (233) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients. The entropy encoding unit (240) can perform various encoding methods, such as, for example, CDF (Cumulative Distribution Function) and CABAC (Context-Adaptive Binary Arithmetic Coding). CDF is a method of storing the cumulative value of a probability distribution, and using CDF allows a symbol to be represented with fewer bits than directly storing the probability. In CDF, the probability of a symbol can be adaptively adjusted based on the context. The entropy encoding unit (240) may encode information necessary for video / image restoration (e.g., values ​​of syntax elements) together or separately, in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be packetized into open bitstream units (OBU) in the form of a bitstream and transmitted or stored. The video / image information may further include information commonly applied to a certain range, such as tile headers, frame headers, and sequence headers. In this disclosure, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream.The above bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) that transmits the signal output from the entropy encoding unit (240) and / or a storage unit (not shown) that stores it may be configured as internal / external elements of the encoding device (200), or the transmission unit may be included in the entropy encoding unit (240).

[0074] Quantized transformation coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transformation coefficients through the inverse quantization unit (234) and the inverse transformation unit (235). The addition unit (250) can generate a reconstructed signal (reconstructed frame, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the prediction unit. In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit (250) may be called a reconstructed unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current frame, and can also be used for inter prediction of the next frame after undergoing filtering as described below.

[0075] The filtering unit (260) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (260) can generate a modified restored frame by applying various filtering methods to the restored frame, and can store the modified restored frame in memory (270), specifically in the frame buffer of memory (270). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0076] The modified restored frame transmitted to the memory (270) can be used as a reference frame in the inter-prediction unit. Through this, when inter-prediction is applied, the encoding device can avoid prediction mismatches between the encoding device (200) and the decoding device, and can also improve encoding efficiency.

[0077] Memory (270) can store a modified restored frame to be used as a reference frame in the inter-prediction unit. Memory (270) can store motion information of blocks from which motion information is derived (or encoded) within the current frame and / or motion information of blocks within the already restored frame. The stored motion information can be transmitted to the inter-prediction unit to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory (270) can store restoration samples of blocks restored within the current frame and transmit them to the intra-prediction unit.

[0078] FIG. 3 is a diagram schematically illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure may be applied. The term "decoding device" below may include an image decoding device and / or a video decoding device.

[0079] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-predictor and an intra-predictor. The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The aforementioned entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) may be configured by a single hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Additionally, the memory (360) may include a frame buffer and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.

[0080] When a bitstream containing video / image information is input, the decoding device (300) can restore the image in correspondence with the process in which the video / image information is processed by the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied by the encoding device. Thus, the processing unit for decoding may be, for example, a super block or a coding block, and the coding block may be divided from the super block according to a quad tree structure, a binary tree structure and / or a binary tree structure, etc. One or more prediction blocks or transformation blocks may be derived from the coding unit. And, the restored image signal decoded and output through the decoding device (300) can be played back through a playback device.

[0081] The decoding device (300) can receive a signal output from the encoding device in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or frame restoration). The video / image information may further include information commonly applied to a certain range, such as a tile header, a frame header, and a sequence header. The decoding device can decode the frame based on the header information. The signaling / received information and / or syntax elements described below in this disclosure can be obtained from the bitstream by decoding through the decoding procedure. For example, the entropy decoding unit (310) can decode information within the bitstream based on coding methods such as CDF, CABAC, etc., and output the values ​​of syntax elements necessary for image restoration and the quantized values ​​of conversion coefficients regarding residuals. More specifically, the CDF-based coding method may include a procedure for storing the probability distribution of a symbol in the form of a CDF, analyzing the context of the current block to select an appropriate CDF, encoding the symbol in the encoding stage based on the CDF, and restoring the syntax elements / information in the decoding stage using the same CDF. Additionally, the CABAC coding method may determine a context model using the target syntax elements / information, coding information of surrounding and target blocks, or information on symbols / bins coded in the previous stage; in the encoding stage, predicting the probability of bin occurrence according to the determined context model and performing arithmetic encoding of the bin to generate a bit sequence; and in the decoding stage, predicting the probability of bin occurrence according to the determined context model and performing arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, after determining the context model, the context model may be updated using the information on the coded symbols / bins for the context model of the next symbol / bin.Information regarding prediction among the information decoded in the entropy decoding unit (310) is provided to the prediction unit (330), and residual values, i.e., quantized transformation coefficients and related parameter information, for which entropy decoding has been performed in the entropy decoding unit (310) can be input to the residual processing unit (320). The residual processing unit (320) can derive residual signals (residual blocks, residual samples, residual sample arrays). Additionally, information regarding filtering among the information decoded in the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310). Meanwhile, the decoding device according to the present disclosure may be called a video / image / frame decoding device, and the decoding device may be divided into an information decoder (video / image / frame information decoder) and a sample decoder (video / image / frame sample decoder). The information decoder may include the entropy decoding unit (310), and the sample decoder may include at least one of the inverse quantization unit (321), inverse transform unit (322), adder (340), filtering unit (350), memory (360), and prediction unit (330).

[0082] In the inverse quantization unit (321), the quantized transformation coefficients can be inversely quantized to output transformation coefficients. The inverse quantization unit (321) can rearrange the quantized transformation coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain transformation coefficients.

[0083] In the inverse conversion unit (322), the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0084] The prediction unit performs a prediction for the current block and can generate a predicted block containing prediction samples for the current block. Based on information regarding the prediction output from the entropy decoding unit (310), the prediction unit can determine whether an intra prediction or an inter prediction is applied to the current block and can determine a specific intra / inter prediction mode.

[0085] The prediction unit (330) can generate a prediction signal based on various prediction methods. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This can be called compound inter-intra prediction. In this case, the mode used as the intra prediction mode may include DC prediction mode, vertical prediction mode, horizontal prediction mode, and SMOOTH prediction mode. Additionally, the prediction unit may be based on the intra block copy mode described above or on the palette mode for predicting a block. As described above, the intra block copy mode is a type of intra prediction that basically performs prediction within the current frame, but can be performed similarly to inter prediction in that it derives a reference block based on a vector using the current frame as a reference frame. However, since the current frame is used as a reference frame, the term block vector may be used instead of motion vector for the vector for movement. That is, the intra-block copy mode may utilize at least one of the inter-prediction techniques described in this disclosure. In this case, a motion vector derived through a surrounding block or a motion information stack may be referenced to derive the block vector of the current block. For example, when the intra-block copy mode is applied, the NEWMV mode described above may be used to derive the block vector of the current block.

[0086] The intra prediction unit can predict the current block by referencing samples within the current frame. Depending on the prediction mode, the referenced samples may be located next to the current block or apart from it. Intra prediction can be performed on a per-transform block basis. If multiple transform blocks exist within a coding block, intra prediction can be performed sequentially in the raster order of the transform blocks. In this case, the procedure for deriving neighboring reference samples for intra prediction can be performed based on the transform block. Multiple prediction modes may be considered for intra prediction. These prediction modes may include multiple non-directional modes and multiple directional modes. The prediction mode used in the current block may be signaled from the encoding device to the decoding device; for example, the prediction modes may include a DC intra prediction mode, multiple directional intra prediction modes, multiple SMOOTH intra prediction modes, and / or PAETH intra prediction modes. These prediction modes may include an intra block copy (intrabc) mode. Whether the above-mentioned intra-block copy mode is applied can be signaled separately. Directional prediction modes may include, for example, eight or more prediction modes depending on the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional prediction modes may be used. The intra-prediction unit may determine the prediction mode applied to the current block based on the prediction mode applied to surrounding blocks. In the intra-block copy mode, a reference block is derived based on a vector, similar to the inter-prediction mode described later, and the current frame is used as the reference frame. The vector used to derive the reference block in the above-mentioned intra-block copy may be called a block vector.Intra prediction can be performed based on various prediction modes, and information regarding predictions obtainable through the bitstream may include information indicating an intra prediction mode for the current block.

[0087] The inter-prediction unit can derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on a reference frame. At this time, to reduce the amount of motion information transmitted in the inter-prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and / or a reference frame index. The motion information may further include information on the inter-prediction direction (L0 prediction, L1 prediction, compound prediction, etc.). In the case of inter-prediction, neighboring blocks may include spatial neighboring blocks existing within the current frame and temporal neighboring blocks existing in the reference frame. The reference frame containing the reference blocks and the reference frame containing the temporal neighboring blocks may be the same or different. The temporal neighboring blocks may be referred to by names such as co-located reference blocks or co-located blocks, and the reference frame containing the temporal neighboring blocks may be referred to as co-located frames. For example, the inter-prediction unit may construct a motion information stack based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference frame index of the current block. The motion information stack may be called a motion information list. The motion information stack may include a motion vector stack. The motion vector stack may be called RefStackMv. The motion vector stack may include eight or more candidates. Information regarding the (maximum) number of candidates in the motion vector stack may be signaled on a frame or sequence basis.Motion modes may be further considered for inter prediction. The motion modes may include simple mode, OBMC (overlapped block motion compensation) mode, and / or local warp mode. In OBMC mode, prediction performance can be improved by utilizing motion information of surrounding blocks for the left and / or upper boundaries of the current block, and in local warp mode, an affine model may be applied in addition to translational motion compensation. Inter prediction can be performed based on various prediction modes, and the prediction information obtainable through the bitstream may include information indicating the mode for inter prediction regarding the current block. For example, the inter prediction unit may construct a motion information stack based on surrounding blocks and derive the motion vector and / or reference frame index of the current block based on received candidate selection information (e.g., reference motion vector index and / or reference frame index).

[0088] The adder (340) can generate a restoration signal (restoration frame, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (predicted block, prediction sample array) output from the prediction unit (330). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the restoration block.

[0089] The adder (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current frame, may be output after filtering as described below, or may be used for inter prediction of the next frame.

[0090] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can generate a modified restored frame by applying various filtering methods to the restored frame, and can transmit the modified restored frame to memory (360), specifically to the frame buffer of memory (360). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0091] The (modified) restored frame stored in memory (360) can be used as a reference frame in the inter-prediction unit. Memory (360) can store movement information of blocks from which movement information within the current frame has been derived (or decoded) and / or movement information of blocks within the already restored frame. The stored movement information can be transmitted to the inter-prediction unit to be used as movement information of spatially surrounding blocks or movement information of temporally surrounding blocks. Memory (360) can store restoration samples of blocks restored within the current frame and transmit them to the intra-prediction unit.

[0092] In this specification, the embodiments described in each configuration of the encoding device (200) may be applied to the corresponding configuration of the decoding device (300) in the same or corresponding manner.

[0093] As described above, prediction is performed to increase compression efficiency during video coding. Through this, a predicted block containing predicted samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived identically by both the encoding device and the decoding device, and the encoding device can increase video coding efficiency by signaling to the decoding device information regarding the residual between the original block and the predicted block (residual information), rather than the original sample value of the original block itself. The decoding device derives a residual block containing residual samples based on the residual information, can generate a restored block containing restored samples by combining the residual block and the predicted block, and can generate a restored frame containing the restored blocks.

[0094] The above residual information can be generated through transformation and quantization procedures. For example, an encoding device may derive a residual block between the original block and the predicted block, perform a transformation procedure on residual samples (residual sample array) included in the residual block to derive transformation coefficients, perform a quantization procedure on the transformation coefficients to derive quantized transformation coefficients, and signal the related residual information to a decoding device (via a bitstream). Here, the residual information may include information such as value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transformation coefficients. The decoding device may perform an inverse quantization / inverse transformation procedure based on the residual information and derive residual samples (or residual blocks). The decoding device may generate a restored frame based on the predicted block and the residual block. The encoding device can also derive a residual block by inversely quantizing / inversely transforming the quantized transform coefficients for reference to inter-prediction of subsequent frames, and generate a restored frame based thereon.

[0095] In the present disclosure, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. If the quantization / inverse quantization is omitted, the quantized transformation coefficient may be referred to as a transformation coefficient. If the transformation / inverse transformation is omitted, the transformation coefficient may be referred to as a coefficient or residual coefficient, or may still be referred to as a transformation coefficient for the sake of consistency of expression.

[0096] Additionally, in this disclosure, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information regarding transform coefficient(s), and information regarding said transform coefficient(s) may be signaled through residual coding syntax. Transform coefficients may be derived based on said residual information (or information regarding said transform coefficient(s), and scaled transform coefficients may be derived through inverse transform (scaling) of said transform coefficients. Residual samples may be derived based on inverse transform (transform) of said scaled transform coefficients. This may be similarly applied / expressed in other parts of this disclosure.

[0097] As described above, the prediction unit of the encoding device / decoding device can derive prediction samples by performing inter-prediction on a block-by-block basis. Inter-prediction can represent a prediction derived in a manner dependent on data elements (i.e., sample values, or motion information, etc.) of picture(s) other than the current picture. When inter-prediction is applied to the current block, a predicted block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on the reference picture pointed to by the reference picture index. At this time, to reduce the amount of motion information transmitted in the inter-prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample basis, based on the correlation of motion information between surrounding blocks and the current block.

[0098] The above motion information may further include information on inter-prediction directions (L0 prediction, L1 prediction, compound prediction, etc.). In the case of inter-prediction, neighboring blocks may include spatial neighboring blocks existing within the current frame and temporal neighboring blocks existing in the reference frame. The reference frame containing the reference block and the reference frame containing the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to by names such as co-located reference block or co-located block, and the reference frame containing the temporal neighboring block may be referred to as a co-located frame. For example, the inter-prediction unit may construct a motion information stack based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference frame index of the current block. The motion information stack may be referred to as a motion information list. The motion information stack may include a motion vector stack. The motion vector stack may be referred to as RefStackMv. The motion vector stack may include eight or more candidates. Information regarding the (maximum) number of candidates in the motion vector stack may be signaled on a frame or sequence basis. Motion modes may be further considered for inter-prediction. The motion modes may include simple mode, OBMC (overlapped block motion compensation) mode and / or local warp mode, etc.In OBMC mode, prediction performance can be improved by utilizing motion information of surrounding blocks for the left and / or upper boundaries of the current block, and in local warp mode, an affine model may be applied in addition to translational motion compensation. Inter-prediction can be performed based on various prediction modes, and the prediction information obtainable through the bitstream may include information indicating the mode for inter-prediction regarding the current block. For example, the inter-prediction unit may construct a motion information stack based on surrounding blocks and derive the motion vector and / or reference frame index of the current block based on received candidate selection information (e.g., reference motion vector index and / or reference frame index, etc.).

[0099] The video / image encoding procedure based on inter prediction and the prediction unit within the encoding device may perform the following, for example, in a general manner for inter prediction.

[0100] Figure 4 illustrates an exemplary inter prediction procedure.

[0101] Referring to FIG. 4, as described above, the inter prediction procedure may include an inter prediction mode / type determination step, a motion vector derivation / refining step, and an inter prediction execution (prediction sample generation) step. The inter prediction procedure may be performed in an encoding device and a decoding device as described above. In this document, the term "coding device" may include an encoding device and / or a decoding device.

[0102] The coding device determines the inter prediction mode / type (S400). The coding device may include an encoding device and / or a decoding device as described above.

[0103] The encoding device can determine an inter prediction mode / type applied to the current block among the various inter prediction modes / types described in the present disclosure and can generate prediction-related information. The prediction-related information may include inter prediction mode information indicating an inter prediction mode applied to the current block and / or inter prediction type information indicating an inter prediction type applied to the current block. The decoding device can determine an inter prediction mode / type applied to the current block based on the prediction-related information.

[0104] The coding device derives / refines the motion vector of the current block (S410). The coding device can derive / refine the motion vector of the current block based on the determined inter-prediction mode / type. Here, to derive / refine the motion vector, motion information of surrounding blocks of the current block or signaled prediction-related information may be used.

[0105] For example, a single reference mode or a composite reference mode may be applied to the current block. The single reference mode may be referred to as a unidirectional inter-prediction mode, and the composite reference mode may be referred to as a bidirectional inter-prediction mode. For example, a compound flag indicating whether a single reference mode or a composite reference mode is applied to the current block may be signaled. If the single reference mode is applied, a single motion vector may be derived for the current block.

[0106] For example, if a NEARMV mode or a NEARESTMV mode is applied to the current block, the decoding device may construct a motion vector candidate list and select one motion vector candidate from among the motion vector candidates included in the motion vector candidate list. The decoding device may derive the motion vector of the current block based on the selected motion vector candidate. Information indicating the selected motion vector candidate (e.g., motion vector candidate index) may be included in the prediction-related information. Additionally, the motion vector candidate list may be referred to as a dynamic reference list or reference stack MV, and the motion vector candidate index may be referred to as a DRL index.

[0107] As another example, when a NEW mode is applied to the current block, the decoding device may construct a motion vector candidate list and derive the motion vector of the current block based on the motion vector of a selected motion vector candidate among the motion vector candidates included in the motion vector candidate list and the signaled MVD (motion vector difference). Information indicating the selected motion vector candidate (e.g., motion vector candidate index) may be included in the prediction-related information. In this case, information regarding the MVD as well as the motion vector candidate index may be included in the prediction-related information. Additionally, the motion vector candidate list may be referred to as a dynamic reference list or reference stack MV, and the motion vector candidate index may be referred to as a DRL index.

[0108] Meanwhile, as described below, movement information of the current block can be derived without constructing a candidate list, and in this case, the movement information of the current block can be derived according to the procedure initiated in the prediction mode / type described below. In this case, the construction of the candidate list as described above may be omitted.

[0109] For example, when a GLOBAL mode is applied to the current block, the decoding device can derive global motion parameters for the current frame and a global motion vector for the current block based on global motion parameter information for the current frame, and can perform inter-prediction for the current block based on the global motion parameters and the global motion vector.

[0110] In addition, for example, when the above composite reference mode is applied, multiple motion vectors can be derived for the current block.

[0111] For example, if NEAREST_NEARESTMV mode, NEAR_NEARMV mode, NEAREST_NEWMV mode, NEW_NEARESTMV mode, NEAR_NEWMV mode, NEW_NEARMV mode, or NEW_NEWMV mode is applied to the current block, the decoding device may construct a motion vector candidate list and select at least one motion vector candidate among the motion vector candidates included in the motion vector candidate list. The decoding device may derive motion vectors of the current block based on the selected at least one motion vector candidate. Information indicating the selected at least one motion vector candidate (e.g., motion vector candidate index) may be included in the prediction-related information. Additionally, the motion vector candidate list may be referred to as a dynamic reference list or reference stack MV, and the motion vector candidate index may be referred to as a DRL index.

[0112] Meanwhile, for example, if the NEAREST_NEWMV mode, NEW_NEARESTMV mode, NEAR_NEWMV mode, NEW_NEARMV mode, or NEW_NEWMV mode is applied to the current block, the decoding device can derive the motion vector of the current block based on the motion vector of the selected motion vector candidate and the signaled MVD. In this case, information regarding the MVD as well as the motion vector candidate index may be included in the prediction-related information.

[0113] Meanwhile, for example, when the GLOBAL_GLOBALMV mode is applied to the current block, the decoding device can derive global motion parameters for the current frame and global motion vectors for the current block based on global motion parameter information for the current frame, and can perform inter-prediction for the current block based on the global motion parameters and global motion vectors.

[0114] The coding device predicts the current block (generates a prediction sample) based on the derived / refined motion vector (S420). The coding device can derive a prediction sample of the current block using samples of the reference block pointed to by the motion vector on the reference picture.

[0115] An encoding procedure based on inter-prediction can roughly include, for example, the following.

[0116] An encoding procedure based on inter-prediction can roughly include, for example, the following.

[0117] Figure 5 shows examples of inter-prediction-based video / image encoding methods.

[0118] Referring to FIG. 5, S500 can be performed by the prediction unit of the encoding device, S505 can be performed by the residual processing unit of the encoding device, and S510 or S515 can be performed by the entropy encoding unit of the encoding device. Specifically, the prediction-related information can be derived by the prediction unit and encoded by the entropy encoding unit. The residual information can be derived by the residual processing unit and encoded by the entropy encoding unit. The residual information is information regarding the residual samples. The residual information may include information regarding quantized transformation coefficients for the residual samples. As described above, the residual samples are derived into transformation coefficients through the transformation unit of the encoding device, and the transformation coefficients can be derived into quantized transformation coefficients through the quantization unit. Information regarding the quantized transformation coefficients can be encoded in the entropy encoding unit through a residual coding procedure.

[0119] The encoding device performs inter prediction for the current block (S500). The encoding device derives the inter prediction mode / type and motion information of the current block and can generate prediction samples for the current block. Here, the procedures for determining the inter prediction mode / type, deriving motion information, and generating prediction samples may be performed simultaneously, or one procedure may be performed before the other. For example, the inter prediction unit of the encoding device can search for a block similar to the current block within a certain area (search area) of reference pictures through motion estimation, and derive a reference block whose difference from the current block is minimal or below a certain standard. Based on this, it can derive a reference picture index pointing to the reference picture where the reference block is located, and derive a motion vector based on the positional difference between the reference block and the current block. The encoding device can determine the mode applied to the current block among various prediction modes. The encoding device can compare the RD cost for the various prediction modes and determine the optimal prediction mode for the current block.

[0120] For example, when a NEW mode or NEAR mode is applied to the current block, the encoding device may construct a motion vector candidate list and derive a reference block among the reference blocks pointed to by the motion vector candidates included in the motion vector candidate list, wherein the difference between the current block and the reference block is minimum or below a certain standard. In this case, a motion vector candidate associated with the derived reference block is selected, and motion vector index information pointing to the selected motion vector candidate is generated and signaled to the decoding device. Motion information of the current block can be derived using the motion information of the selected motion vector candidate.

[0121] Alternatively, for example, when the NEAREST mode is applied to the current block, the encoding device may construct a motion vector candidate list and derive a reference block pointed to by a first motion vector candidate among the reference blocks pointed to by the motion vector candidates included in the motion vector candidate list as a reference block in which the difference between the current block and the first motion vector candidate is minimum or below a certain standard. The first motion vector candidate may be a motion vector candidate with an index value of 0 in the motion vector candidate list. In this case, the first motion vector candidate is selected, and the motion vector index information may not be signaled. When the NEAREST mode is applied to the current block, the decoding device may construct a motion vector candidate list and select the first motion vector candidate. The motion information of the current block may be derived using the motion information of the selected motion vector candidate.

[0122] The encoding device can perform residual processing based on the predicted samples (S505). The encoding device can derive residual samples based on the predicted samples. The encoding device can derive the residual samples by comparing the original samples of the current block with the predicted samples. Residual information can be generated based on the residual samples. The residual information may include information regarding quantized transformation coefficients as described above.

[0123] The encoding device encodes video information including prediction-related information and / or residual information (S510 or S515). The encoding device may output the encoded video information in the form of a bitstream. The prediction-related information may include information related to the prediction procedure, such as prediction mode information (e.g., skip_mode, compound_mode, new_mv, zero_mv, ref_mv, etc.) and / or information regarding motion information. The information regarding motion information may include motion type information (e.g., motion_mode) and / or candidate selection information (e.g., drl_mode) which is information for deriving a motion vector. Additionally, the information regarding motion information may include information regarding the MVD described above and / or reference picture index information. Furthermore, for example, the motion type information may indicate whether the motion compensation type of the current block is a SIMPLE type, OBMC (overlapped block motion compensation) type, or LOCALWARP type, which is a general inter-prediction using a motion vector. The SIMPLE type can represent translational motion compensation, the OBMC type can improve prediction performance by utilizing motion information of surrounding blocks for the left and / or upper boundaries of the current block, and the local warp type can apply an affine model in addition to translational motion compensation. The residual information is information regarding the residual samples. The residual information may include information regarding quantized transform coefficients for the residual samples.

[0124] The output bitstream can be stored in a (digital) storage medium and transmitted to a decoding device, or it can be transmitted to a decoding device via a network.

[0125] Meanwhile, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is intended to derive the same prediction results from the encoding device as those performed by the decoding device, thereby increasing coding efficiency. Accordingly, the encoding device can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in memory and utilize it as a reference picture for inter-prediction. As described above, an in-loop filtering procedure, etc., may be further applied to the reconstructed picture.

[0126] The decoding device can perform an operation corresponding to the operation performed by the encoding device. A video / image decoding procedure based on inter-prediction may include, for example, the following.

[0127] Figure 6 shows examples of inter-prediction-based video / image decoding methods.

[0128] Referring to FIG. 6, S600 can be performed by the entropy decoding unit of the decoding device, S610 can be performed by the prediction unit of the decoding device, S615 can be performed by the residual processing unit of the decoding device, and S620 can be performed by the adder or restoration unit of the decoding device.

[0129] Specifically, the decoding device obtains image / video information from the bitstream (S600). The image / video information may include prediction-related information and / or residual information.

[0130] The decoding device performs inter-prediction based on prediction-related information (S610). Based on the prediction-related information, the decoding device can derive an inter-prediction mode / type for the current block, derive / refine movement information of the current block, and generate prediction samples within the current block based on the intra-prediction mode / type and / or the movement information. In this case, the decoding device can perform a prediction sample filtering procedure. Prediction sample filtering may be referred to as post-filtering. Some or all of the prediction samples may be filtered by the prediction sample filtering procedure. In some cases, the prediction sample filtering procedure may be omitted.

[0131] The decoding device performs residual processing based on residual information (S615). The decoding device can derive residual samples for the current block based on the residual information. Specifically, the inverse quantization unit of the residual processing unit derives transformation coefficients by performing inverse quantization based on the quantized transformation coefficients derived based on the residual information, and the inverse transformation unit of the residual processing unit derives residual samples for the current block by performing inverse transformation on the transformation coefficients.

[0132] The decoding device generates a restored block / picture (S620). The decoding device can generate restoration samples for the current block based on the prediction samples and / or the residual samples, and derive a restored block containing the restoration samples. A restored picture for the current picture can be generated based on the restored block. As described above, an in-loop filtering procedure, etc., may be further applied to the restored picture.

[0133] The above prediction-related information can be encoded / decoded through the binarization and coding methods described in this document. For example, the above prediction-related information can be binarized through fixed-length binarization, truncated rice binarization, truncated univariate binarization, etc. For example, the above prediction-related information can be encoded / decoded through entropy coding (e.g., CABAC, CAVLC).

[0134] Meanwhile, if inter prediction is applied to the current block, the inter prediction mode / type of the current block can be derived as follows.

[0135] Figure 7 shows an example of a method for determining the inter-prediction mode / type of the current block.

[0136] The decoding device can determine whether isCompound is 1 (S700). The decoding device can determine whether the value of isCompound is 1. Here, isCompound may be a variable indicating whether a composite prediction is applied to the current block. For example, if the value of isCompound is 1, isCompound may indicate that a composite prediction is applied to the current block, and if the value of isCompound is 0, isCompound may indicate that a single prediction is applied to the current block. The composite prediction may be referred to as a composite inter prediction, and the single prediction may be referred to as a single inter prediction.

[0137] For example, the variable isCompound can be derived based on the value of the reference frame index representing reference frame 1 of the current block. For example, if the reference frame index representing reference frame 1 does not represent NONE, the value of isCompound can be derived as 1, and if the reference frame index representing reference frame 1 represents NONE, the value of isCompound can be derived as 0. For example, if the value of the reference frame index representing reference frame 1 is -1, the reference frame index representing reference frame 1 can represent NONE. The NONE may mean that a single prediction is applied.

[0138] If isCompound is not 1, the decoding device can obtain new_mv (S705) and determine whether new_mv is 0 (S710). If isCompound is not 1, new_mv for the current block can be signaled. The new_mv can indicate whether the MVD of the current block exists. The new_mv can be represented as a newmv mode index.

[0139] When the above new_mv is 0, the decoding device can derive the NEWMV mode as the inter-prediction mode of the current block (S715). When the above new_mv is 0, the MVD of the current block can be signaled.

[0140] When the above new_mv is 1, the decoding device can obtain zero_mv (S720) and determine whether zero_mv is 0 (S725). When new_mv is 1, zero_mv for the current block can be signaled. The above zero_mv may indicate whether the motion vector of the current block is set to be the same as the default motion vector of the current frame. The above zero_mv may be represented as a zeromv mode index or a GLOBALMV mode index. Additionally, the default motion vector of the current frame may also be called a global motion vector.

[0141] When zero_mv is 0, the decoding device can derive the GLOBALMV mode as the inter-prediction mode of the current block (S730). When zero_mv is 0, the motion vector of the current block can be derived as the global motion vector of the current frame.

[0142] When zero_mv is 1, the decoding device can obtain ref_mv (S735) and determine whether ref_mv is 0 (S740). When zero_mv is 1, ref_mv for the current block can be signaled. If ref_mv is 0, it may mean that the most probable motion vector (i.e., NEAREST) ​​is used, and if ref_mv is 1, it may mean that the second most probable motion vector (i.e., NEAR) is used. The ref_mv can be represented as a refmv mode index.

[0143] That is, for example, when the above ref_mv is 0, the decoding device can derive the NEARESTMV mode as the inter prediction mode of the current block (S745), and when the above ref_mv is 1, the decoding device can derive the NEARMV mode as the inter prediction mode of the current block (S750).

[0144] Meanwhile, when isCompound is 1, the decoding device can obtain compound_mode (S755) and derive the NEAREST_NEAREST + compound_mode mode as the inter-prediction mode of the current block (S760). When isCompound is 1, a composite prediction can be applied to the current block. The compound_mode can be represented as a composite mode index.

[0145] The decoding device can derive the composite reference mode of the current block based on compound_mode. For example, the composite reference mode may be as shown in the following table.

[0146]

[0147] For example, the decoding device can derive the prediction mode of the current block by adding an offset to compound_mode. For example, the decoding device can derive the prediction mode of the current block as the value obtained by adding compound_mode to 18, which is the index value representing NEAREST_NEAREST. For example, if the value of compound_mode is n, the 18 + n mode can be derived as the prediction mode of the current block.

[0148] Meanwhile, for example, a TIP (Temporarily Interpolated Prediction) mode may be proposed as an inter-prediction mode. The TIP mode may represent a prediction mode that performs inter-prediction based on a reference frame derived by interpolating reference frames. The reference frame derived by interpolating the reference frames may be referred to as the TIP reference frame.

[0149] For example, when TIP mode is applied, a TIP reference frame can be generated based on the reference frames of the current frame. For example, a TIP reference frame can be generated by interpolating the reference frames of the current frame.

[0150] FIG. 8 shows an example of performing inter prediction based on the TIP reference frame of the current frame.

[0151] For example, referring to FIG. 8, a TIP reference frame of the current frame can be generated based on a backward reference frame of the current frame and a forward reference frame of the current frame. For example, generating a TIP reference frame based on a backward reference frame and a forward reference frame may include using motion vectors of the backward reference frame and the forward reference frame to generate a motion field for the current frame, and using a motion field to fetch a reference block within the TIP reference frame (i.e., through interpolation).

[0152] For example, motion fields based on the backward reference frame and the forward reference frame of the current frame can be used. For example, as shown in FIG. 8, the current frame can be represented as Fi, the backward reference frame as fi-i, and the forward reference frame as Fi+i.

[0153] For example, the reverse reference frame and the forward reference frame may be separated by the same distance from the current frame in the display order of the video sequence. Alternatively, for example, the reverse reference frame and the forward reference frame may be reference frames separated by a different distance from the current frame in the display order. The temporal motion vector predictor of FIG. 8 may be a motion vector predictor pointing from the reverse reference frame to the forward reference frame. A motion vector pointing to the TIP reference frame (a specific block, tile, or area of ​​the TIP reference frame) generated from the current frame (a specific block, tile, or area of ​​the current frame) may represent a motion vector that can be used to predict a specific block, tile, or area of ​​the current frame based on a specific block, tile, or area of ​​the generated TIP reference frame.

[0154] For example, the TIP reference frame may be generated by interpolating the reverse reference frame and the forward reference frame based on a first distance between the reverse reference frame and the current frame and a second distance between the forward reference frame and the current frame. Alternatively, for example, the TIP reference area (or TIP reference tile or TIP reference block) of the TIP reference frame may be generated by interpolating the first area (or first tile or first block) of the reverse reference frame and the second area (or second tile or second block) of the forward reference frame pointed to by the motion vector of the first area of ​​the reverse reference frame based on the first distance and the second distance. The first distance may represent the distance between the reverse reference frame and the current frame, and the second distance may represent the distance between the forward reference frame and the current frame.

[0155] For example, the prediction mode of the current frame can be derived as a TIP prediction mode based on prediction-related information of the current frame through a bitstream, and a TIP reference frame can be generated by interpolating the reverse reference frame and the forward reference frame of the current frame.

[0156] Subsequently, movement information of the current block can be derived based on information related to the inter-prediction of the current block, and a prediction sample of the current block can be derived based on the reference block within the TIP reference frame pointed to by the movement information.

[0157] Alternatively, for example, if the value of the reference frame index of the current block via the bitstream is a specific value, said reference frame index may point to a TIP reference frame. For example, if the value of the reference frame index of the current block is a specific value, TIP prediction may be applied to said current block, and a TIP reference frame may be generated by interpolating the reverse reference frame and the forward reference frame of said current frame. For example, said specific value may be pre-set.

[0158] Subsequently, movement information of the current block can be derived based on prediction-related information of the current block, and a prediction sample of the current block can be derived based on the reference block within the TIP reference frame pointed to by the movement information.

[0159] Alternatively, for example, index information pointing to reference frames for deriving the TIP reference frame may be signaled. For example, based on the index information pointing to the signaled reference frames, the reverse reference frame and the forward reference frame of the current frame may be derived, and the TIP reference frame may be generated by interpolating the reverse reference frame and the forward reference frame of the current frame. For example, the index information pointing to the reference frames for deriving the TIP reference frame may be signaled as a frame header syntax.

[0160] Alternatively, for example, it may not be limited to the interpolation of a reverse reference frame and a forward reference frame. That is, for example, two reference frames may be derived, and a TIP reference frame may be generated by interpolating said reference frames. For example, index information pointing to the reference frames for deriving said TIP reference frame may be signaled. Based on the index information pointing to the signaled reference frames, two reference frames of said current frame may be derived, and a TIP reference frame may be generated by interpolating said reference frames. For example, the index information pointing to the reference frames for deriving said TIP reference frame may be signaled as a frame header syntax.

[0161] In addition, when inter-prediction is applied to the current block as described above, a predicted block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on the reference picture pointed to by the reference picture index. Here, a reference picture list including reference pictures for the current picture may be constructed, and the reference picture index may point to one of the reference pictures in the reference picture list. The current picture may be referred to as the current frame, and the reference picture may be referred to as the reference frame.

[0162] For example, when inter-prediction is applied to the current block, a reference frame list containing reference frames may be constructed, and a predicted block (prediction sample array) for the current block may be derived based on a reference block (reference sample array) specified by a motion vector on a reference frame pointed to by a reference frame index for the current frame.

[0163] For example, as shown in the following table, seven reference frames can be used for the current frame.

[0164]

[0165] Referring to Table 2 above, a reference frame list including LAST_FRAME, LAST2_FRAME, LAST3_FRAME, GOLDEN_FRAME, BWDREF_FRAME, ALTREF2_FRAME, and ALTREF_FRAME of the current frame may be constructed. As shown in Table 2, LAST_FRAME may represent the nearest past frame, LAST2_FRAME may represent the second nearest past frame, LAST3_FRAME may represent the third nearest past frame, GOLDEN_FRAME may represent the distant past frame, BWDREF_FRAME may represent a reverse reference to the nearest frame, ALTREF2_FRAME may represent the next nearest reverse reference, and ALTREF_FRAME may represent a reverse reference to the frame with the highest output order. For example, the reference frame index representing the LAST_FRAME and the reference frame index representing the GOLDEN_FRAME may be explicitly signaled. The reference frame index representing the above LAST_FRAME can be represented as the last frame index, and the reference frame index representing the above GOLDEN_FRAME can be represented as the gold frame index. Reference frames excluding the above LAST_FRAME and the above GOLDEN_FRAME can be derived based on the above last frame index and the above gold frame index. That is, reference frames excluding the above LAST_FRAME and the above GOLDEN_FRAME can be derived based on the above last frame index pointing to the LAST_FRAME and the above gold frame index pointing to the GOLDEN_FRAME.

[0166] On the other hand, if only the distance from the current frame is considered, it may not be possible to reflect cases where the similarity with the current frame changes due to differences in quantization parameters. Therefore, by constructing a reference frame list that takes into account the difference in similarity caused by differences in quantization parameters, reference frames with high similarity can be identified as candidates for small indices, thereby enabling a reduction in the amount of bits of information related to the signaled reference frames.

[0167] For example, the above reference frames may be replaced with ranks 0 through 6. For example, a list of reference frames for the current frame may be constructed based on the cost of the reference frames of the current frames.

[0168] FIG. 9 shows an embodiment of configuring a reference frame list based on cost.

[0169] For example, the number of reference frames (num_total_refs) of the current frame and the ranks of the reference frames may be implicitly derived. For example, up to seven frames may be derived as the reference frames of the current frame from the LAST_FRAME, LAST2_FRAME, LAST3_FRAME, GOLDEN_FRAME, BWDREF_FRAME, ALTREF2_FRAME, and ALTREF_FRAME of the current frame, and the ranks of the reference frames may be derived based on the costs of the reference frames. In this case, for example, the last frame index representing the LAST_FRAME and the gold frame index representing the GOLDEN_FRAME may be explicitly signaled. And regarding the reference frames excluding the above GOLDEN_FRAME, the reference frame pointed to by the above last frame index can be derived as the above LAST_FRAME, and the reference frame pointed to by the above gold frame index can be derived as the above GOLDEN_FRAME, and the reference frames excluding the above LAST_FRAME and the above GOLDEN_FRAME can be derived based on the above last frame index and the above gold frame index.

[0170] Alternatively, for example, the costs of reference frames coded prior to the current frame may be derived, and up to 7 reference frames of rank 0 to rank 6 may be derived based on the costs of said reference frames. For example, up to 7 reference frames of rank 0 to rank 6 may be derived in order of smallest cost among the reference frames. Among said reference frames, the reference frame with the smallest cost may be derived as rank 0, the reference frame with the second smallest cost may be rank 1, the reference frame with the third smallest cost may be rank 2, the reference frame with the fourth smallest cost may be rank 3, the reference frame with the fifth smallest cost may be rank 4, the reference frame with the sixth smallest cost may be rank 5, and the reference frame with the seventh smallest cost may be rank 6.

[0171] Here, the cost of the reference frame can be derived based on 1) the temporal distance of the reference frame and 2) the quantizer of the reference frame. For example, the cost of the reference frame can be derived based on the following mathematical formula.

[0172]

[0173] Here, cost may represent the cost of the reference frame, d may represent the temporal distance of the reference frame, and q may represent the quantization parameter of the reference frame. The distance of the reference frame may be the difference between the picture order count (POC) of the current frame and the POC of the reference frame. Alternatively, the distance of the reference frame may be the difference between the display order count of the current frame and the display order count of the reference frame. The ranks of the reference frames may be derived in order of decreasing cost. That is, the reference frames may be sorted in order of decreasing cost.

[0174] Additionally, for example, the reference frames may be derived from frames among the stored frames whose temporal distance is less than or equal to a specific value. The specific value may be a preset value. Alternatively, for example, the reference frames may be derived from frames among the stored frames whose cost is less than or equal to a specific value. The specific value may be a preset value.

[0175] Subsequently, a reference frame index for the current block can be signaled, and a reference frame of the rank pointed to by the reference frame index can be derived as the reference frame of the current block.

[0176] For example, the above reference frame index can be signaled as a unary codeword. For example, if a composite reference mode is applied to the current block, two reference frames of the current block can be derived based on the above reference frame index.

[0177] For example, if the num_total_refs of the current frame is 4 and a composite reference mode based on a reference frame of rank 2 and a reference frame of rank 4 is applied to the current block, the reference frame index of the current block may be coded as 010. Additionally, for example, if the num_total_refs of the current frame is greater than 4 and a composite reference mode based on a reference frame of rank 2 and a reference frame of rank 4 is applied to the current block, the reference frame index of the current block may be coded as 0101.

[0178] Additionally, for example, the reference frame index may be signaled at the frame header, tile group, tile, or block level. The bits of the reference frame index may be bits transmitted to indicate whether to use them for candidates in the reference frame list.

[0179] Additionally, when the aforementioned composite reference prediction is applied, a motion vector predictor 0 (MVP0) for reference frame 0 and a motion vector predictor 1 (MVP1) for reference frame 1 can be derived, an MVD0 for reference frame 0 and an MVD1 for reference frame 1 can be derived, a motion vector 0 (MV0) for reference frame 0 can be derived based on the MVP0 and MVD0, and a motion vector 1 (MV1) for reference frame 1 can be derived based on the MVP1 and MVD1. Meanwhile, reference frame 0 may be referred to as the first reference frame, and reference frame 1 may be referred to as the second reference frame.

[0180] The above MVD0 and the above MVD1 can be derived based on MVD-related information signaled through a bitstream. According to the various composite reference modes described above, the two MVDs can be signaled individually or jointly in the bitstream.

[0181] Alternatively, for example, Joint MVD coding that derives MVD0 and MVD1 based on signaling MVD-related information may be applied.

[0182] FIG. 10 shows an example of deriving MVDs for reference frames based on a joint MVD.

[0183] For example, in addition to the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes described above, another inter-prediction mode designated as JOINT_NEWMV may be introduced for a mode in which MVDs for reference frame 0 and reference frame 1 are signaled together. Specifically, when the inter-prediction mode is indicated as NEW_NEWMV, MVDs for reference frame 0 and reference frame 1 are signaled individually, whereas when the inter-prediction mode is indicated as JOINT_NEWMV, MVDs for reference frame 0 and reference frame 1 may be signaled together. That is, when the inter-prediction mode is indicated as JOINT_NEWMV, one MVD-related information for the MVD for reference frame 0 and the MVD for reference frame 1 may be signaled.

[0184] For example, if a composite reference mode is applied to the current block, a joint_newmv syntax indicating whether the JOINT_NEWMV is applied may be signaled. If the value of the joint_newmv syntax is 1, the prediction mode of the current block may be derived as JOINT_NEWMV, and if the value of the joint_newmv syntax is 0, the prediction mode of the current block may be derived as NEW_NEWMV.

[0185] Alternatively, for example, if the value obtained by adding an offset to compound_mode is a specific value, the prediction mode of the current block may be derived as JOINT_NEWMV. That is, for example, it can be determined whether the prediction mode of the current block is JOINT_NEWMV based on compound_mode. The offset may be an index value representing NEAREST_NEAREST, and the specific value may be a preset value.

[0186] For example, when JOINT_NEWMV is applied to the current block, joint MVD information for one MVD (e.g., joint_mvd) can be signaled as a bitstream, and MVDs for reference frame 0 and reference frame 1 can be derived from joint_mvd. Then, the derived MVDs can be combined with reference motion vectors from reference frame 0 or reference frame 1 to generate two motion vectors for finding reference blocks for composite inter-prediction. The reference motion vectors can be referred to as MVP (motion vector predictor).

[0187] For example, scaling of the signaled joint MVD may be performed to obtain one or both of the two MVDs (MVD0 and MVD1). In other words, at least one of MVD0 and MVD1 may be derived by scaling the signaled joint MVD. As a result of scaling, the precision or pixel resolution of the scaled MVD may differ from the allowable precision of the motion vector difference.

[0188] In some exemplary implementations, such MVD(s) scaled from jointly signaled MVDs may first be quantized to the allowable precision of the MV(s) for the current picture or slice or tile or superblock or coded block before being added to the reference MV(s) to generate motion vector(s).

[0189] For example, the JOINT_NEWMV mode may be signaled, and if the POC distance between reference frame 0 and the current frame and the POC distance between reference frame 1 and the current frame are different, the joint MVD may be scaled based on the POC distances to derive MVD0 and / or MVD1.

[0190] Specifically, the distance between reference frame 0 and the current frame can be denoted as td0, and the distance between reference frame 1 and the current frame can be denoted as td1. For example, if td0 is greater than or equal to td1, the joint MVD (i.e., joint_mvd) can be derived as MVD0, and MVD1 can be derived by scaling the joint MVD as shown in the following mathematical formula.

[0191]

[0192] Here, joint_mvd can represent a signaled joint MVD, td0 can represent the distance between reference frame 0 and the current frame, and td1 can represent the distance between reference frame 1 and the current frame.

[0193] Alternatively, for example, if td1 is greater than or equal to td0, the joint MVD can be derived as MVD1, and MVD0 can be derived by scaling the joint MVD as shown in the following mathematical formula.

[0194]

[0195] Here, joint_mvd can represent a signaled joint MVD, td0 can represent the distance between reference frame 0 and the current frame, and td1 can represent the distance between reference frame 1 and the current frame.

[0196] Alternatively, for example, if the JOINT_NEWMV mode is applied to the current block, the joint MVD (i.e., joint_mvd) and scaling factor information of the current block may be signaled. For example, the scaling factor information may indicate one of the scaling factor pairs. The scaling factor pair indicated by the scaling factor information may be derived as the scaling factor for MVD0 and the scaling factor for MVD1 of the current block, and the MVD0 may be derived by multiplying the joint MVD by the scaling factor for MVD0, and the MVD1 may be derived by multiplying the joint MVD by the scaling factor for MVD1.

[0197] For example, candidate scaling coefficient pairs may be as follows.

[0198]

[0199] Referring to Table 3, the scaling coefficient pair indicated by the scaling coefficient information may include a scaling coefficient for MVD0 and a scaling coefficient for MVD1. For example, referring to Table 3, if the value of the scaling coefficient information is 0, the scaling coefficient for MVD0 may be derived as 1 and the scaling coefficient for MVD1 as 2. Additionally, for example, if the value of the scaling coefficient information is 1, the scaling coefficient for MVD0 may be derived as 1 and the scaling coefficient for MVD1 as 4. The candidate scaling coefficient pairs shown in the table above are merely examples and are not limited thereto.

[0200] MVD0 can be derived by multiplying the signaled joint MVD by a scaling factor for the derived MVD0, and MVD1 can be derived by multiplying the joint MVD by a scaling factor for MVD1.

[0201] FIG. 11 schematically illustrates a video / image encoding method according to an embodiment(s) of the present disclosure. The method disclosed in FIG. 11 may be performed by the encoding device disclosed in FIG. 2. Specifically, for example, S1100 to S1120 of FIG. 11 may be performed by the prediction unit (220) of the encoding device (200), and S1130 of FIG. 11 may be performed by the entropy encoding unit (240) of the encoding device (200). The method disclosed in FIG. 11 may include the embodiments described above in the present disclosure.

[0202] Referring to FIG. 11, the encoding device derives MVP0 (motion vector predictor 0) for reference frame 0 of the current block and MVP1 for reference frame 1 (S1100).

[0203] The encoding device can derive the reference frame 0 and the reference frame 1 for the current block among the reference frames of the current frame. The encoding device can generate prediction-related information including a first reference frame index representing the reference frame 0 for the current block and a second reference frame index representing the reference frame 1.

[0204] Additionally, for example, the encoding device may construct the motion vector candidate list based on the surrounding blocks of the current block, and may derive the MVP0 for the reference frame 0 and the MVP1 for the reference frame 1 of the current block based on the motion vector candidate list. The surrounding blocks may include spatial surrounding blocks and / or temporal surrounding blocks of the current block. The encoding device may generate prediction-related information including a first motion vector candidate index representing the MVP0 for the reference frame 0 among the motion vector candidate list and a second motion vector candidate index representing the MVP1 for the reference frame 1 among the motion vector candidate list.

[0205] The encoding device derives the MVD0 and MVD1 of the current block based on the joint motion vector difference (joint MVD) of the current block (S1110).

[0206] For example, the encoding device may determine that the JOINT_NEWMV mode is applied to the current block and derive the joint MVD of the current block.

[0207] Subsequently, the encoding device can derive the MVD0 and MVD1 of the current block based on the joint MVD.

[0208] For example, at least one of the above MVD0 and the above MVD1 can be derived by scaling the joint MVD.

[0209] For example, at least one of the MVD0 and the MVD1 can be derived by scaling the joint MVD based on a first temporal distance between the current frame and the reference frame 0 and a second temporal distance between the current frame and the reference frame 1.

[0210] For example, if the first temporal distance is greater than or equal to the second temporal distance, the MVD0 can be derived as the joint MVD, and the MVD1 can be derived by scaling the joint MVD based on the first temporal distance and the second temporal distance. For example, if the first temporal distance is greater than or equal to the second temporal distance, the MVD1 can be derived based on the above-described mathematical formula 2.

[0211] Additionally, for example, if the second temporal distance is greater than or equal to the first temporal distance, the MVD1 can be derived as the joint MVD, and the MVD0 can be derived by scaling the joint MVD based on the first temporal distance and the second temporal distance. For example, if the second temporal distance is greater than or equal to the first temporal distance, the MVD0 can be derived based on the above-described mathematical formula 3.

[0212] Alternatively, as another example, the encoding device may derive one of the scaling factor pairs as a first scaling factor for the MVD0 and a second scaling factor for the MVD1, and may derive the MVD0 and the MVD1 based on the first scaling factor and the second scaling factor. For example, one of the scaling factor pairs may be derived as a first scaling factor for the MVD0 of the current block and a second scaling factor for the MVD1, the MVD0 may be derived by multiplying the joint MVD by the first scaling factor, and the MVD1 may be derived by multiplying the joint MVD by the second scaling factor.

[0213] The encoding device derives the MV0 (motion vector 0) of the current block based on the MVP0 and the MVD0 (S1120). The encoding device can derive the MV0 of the current block by adding the MVD0 to the MVP0.

[0214] The encoding device derives the MV1 of the current block based on the MVP1 and the MVD1 (S1130). The encoding device can derive the MV1 of the current block by adding the MVD1 to the MVP1.

[0215] The encoding device derives a prediction sample of the current block based on the MV0 and the MV1 (S1140). For example, the encoding device may generate a prediction sample of the current block based on a reference sample within the reference frame 0 derived based on the MV0 and a reference sample within the reference frame 1 derived based on the MV1. For example, the encoding device may derive a first prediction block of the current block based on the reference block within the reference frame 0 pointed to by the MV0, derive a second prediction block of the current block based on the reference block within the reference frame 1 pointed to by the MV1, and derive a final prediction block of the current block by weighting the first prediction block and the second prediction block. In this case, as described above, a prediction sample filtering procedure may be further performed on all or some of the prediction samples of the current block depending on the case.

[0216] The encoding device generates prediction-related information including joint MVD information of the current block (S1150). The encoding device may generate prediction-related information including the joint MVD information of the current block. The encoding device may generate joint MVD information for the joint MVD of the current block.

[0217] For example, the prediction-related information may include joint motion vector difference (joint MVD) information of the current block. For example, if the JOINT_NEWMV mode is applied to the current block, the joint MVD information may be signaled.

[0218] For example, the prediction-related information may include a composite mode index of the current block, and if the composite mode index indicates that the JOINT_NEWMV mode is applied to the current block, the joint MVD information may be signaled. That is, for example, the composite mode index may indicate the JOINT_NEWMV mode among a plurality of composite reference modes, and if the composite mode index indicates that the JOINT_NEWMV mode is applied to the current block, the joint MVD information may be signaled.

[0219] Alternatively, for example, the prediction-related information may include a composite mode index of the current block, and if the composite mode index indicates that the NEW_NEWMV mode is applied to the current block, a JOINT_NEWMV mode flag indicating whether the JOINT_NEWMV mode is applied to the current block may be signaled. For example, if the JOINT_NEWMV mode flag indicates that the JOINT_NEWMV mode is applied to the current block, the joint MVD information may be signaled.

[0220] Additionally, as an example, the prediction-related information may include prediction mode information and / or prediction type information of the current block. For example, the prediction mode information may include at least one of a newmv mode index, a zeromv mode index, a refmv mode index, or a compound mode index. The newmv mode index may indicate whether the inter-prediction mode of the current block is NEWMV mode, the zeromv mode index may indicate whether the inter-prediction mode of the current block is GLOBALMV mode, the refmv mode index may indicate whether the inter-prediction mode of the current block is NEARESTMV mode or NEARMV mode, and the compound mode index may indicate one of the compound reference modes as the inter-prediction mode of the current block.

[0221] For example, the zeromv mode index can be signaled when the value of the newmv mode index is 1. Also, for example, the refmv mode index can be signaled when the value of the zeromv mode index is 1.

[0222] Additionally, for example, if a compound reference mode is applied to the current block, the compound mode index may be signaled.

[0223] For example, the prediction-related information may include at least one reference frame index of the current block. For example, the prediction-related information may include a reference frame index representing reference frame 0 of the current block. Or, for example, the prediction-related information may include a reference frame index representing reference frame 0 of the current block and a reference frame index representing reference frame 1 of the current block. If the reference frame index representing reference frame 1 of the current block indicates NONE, a single reference mode may be applied to the current block, and if the reference frame index representing reference frame 1 of the current block does not indicate NONE, a composite reference mode may be applied to the current block.

[0224] Additionally, for example, the prediction-related information may include motion vector information of the current block. For example, the prediction-related information may include a motion vector candidate index representing one motion vector candidate among a list of motion vector candidates.

[0225] In addition, the above image information may include various information according to embodiments of the present disclosure.

[0226] Meanwhile, the above image information may include residual information. The above residual information is information regarding residual samples. The above residual information may include information regarding quantized transformation coefficients for the above residual samples.

[0227] Encoded video information can be output in the form of a bitstream. The bitstream can be transmitted to a decoding device via a network or a storage medium. For example, video data containing the bitstream can be transmitted to a decoding device by a transmission device (or transmission unit). In this case, the video data containing the bitstream can be transmitted to the decoding device via a streaming server.

[0228] In addition, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is intended to derive the same prediction results from the encoding device as those performed by the decoding device, thereby increasing coding efficiency. Accordingly, the encoding device can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in memory and utilize it as a reference picture for inter-prediction. As described above, an in-loop filtering procedure, etc., may be further applied to the reconstructed picture.

[0229] According to the above-described embodiment(s), MVD0 and MVD1 for composite reference prediction can be derived based on joint MVDs, thereby reducing the amount of motion information data for inter prediction. Additionally, the accuracy of deriving MVD0 and MVD1 can be improved by signaling scaling information for joint MVDs, thereby improving the accuracy of inter prediction and reducing the amount of information data for prediction.

[0230] FIG. 12 schematically illustrates a video / image decoding method according to an embodiment(s) of the present disclosure. The method disclosed in FIG. 12 may be performed by the decoding device disclosed in FIG. 3. Specifically, for example, S1200 of FIG. 12 may be performed by the entropy decoding unit (310) of the decoding device (300), and S1210 to S1260 may be performed by the prediction unit (330) of the decoding device (300). The method disclosed in FIG. 12 may include the embodiments described above in the present disclosure.

[0231] Referring to FIG. 12, the decoding device obtains prediction-related information including joint motion vector difference (MVD) information of the current block (S1200). The decoding device can obtain image information including the prediction-related information through a bitstream. The image information may further include residual information as described above.

[0232] For example, the prediction-related information may include prediction mode information and / or prediction type information of the current block. For example, the prediction mode information may include at least one of a newmv mode index, a zeromv mode index, a refmv mode index, or a compound mode index. The newmv mode index may indicate whether the inter prediction mode of the current block is NEWMV mode, the zeromv mode index may indicate whether the inter prediction mode of the current block is GLOBALMV mode, the refmv mode index may indicate whether the inter prediction mode of the current block is NEARESTMV mode or NEARMV mode, and the compound mode index may indicate one of the compound modes as the inter prediction mode of the current block.

[0233] For example, the zeromv mode index can be signaled when the value of the newmv mode index is 1. Also, for example, the refmv mode index can be signaled when the value of the zeromv mode index is 1.

[0234] Additionally, for example, if a compound reference mode is applied to the current block, the compound mode index may be signaled.

[0235] For example, the prediction-related information may include at least one reference frame index of the current block. For example, the prediction-related information may include a reference frame index representing reference frame 0 of the current block. Or, for example, the prediction-related information may include a reference frame index representing reference frame 0 of the current block and a reference frame index representing reference frame 1 of the current block. If the reference frame index representing reference frame 1 of the current block indicates NONE, a single reference mode may be applied to the current block, and if the reference frame index representing reference frame 1 of the current block does not indicate NONE, a composite reference mode may be applied to the current block.

[0236] Additionally, for example, the prediction-related information may include joint motion vector difference (joint MVD) information of the current block. For example, when the JOINT_NEWMV mode is applied to the current block, the joint MVD information may be signaled.

[0237] For example, the prediction-related information may include a composite mode index of the current block, and if the composite mode index indicates that the JOINT_NEWMV mode is applied to the current block, the joint MVD information may be signaled. That is, for example, the composite mode index may indicate the JOINT_NEWMV mode among a plurality of composite reference modes, and if the composite mode index indicates that the JOINT_NEWMV mode is applied to the current block, the joint MVD information may be signaled.

[0238] Alternatively, for example, the prediction-related information may include a composite mode index of the current block, and if the composite mode index indicates that the NEW_NEWMV mode is applied to the current block, a JOINT_NEWMV mode flag indicating whether the JOINT_NEWMV mode is applied to the current block may be signaled. For example, if the JOINT_NEWMV mode flag indicates that the JOINT_NEWMV mode is applied to the current block, the joint MVD information may be signaled.

[0239] Additionally, for example, the prediction-related information may include motion vector information of the current block. For example, the prediction-related information may include a motion vector candidate index representing one motion vector candidate among a list of motion vector candidates.

[0240] The decoding device derives MVP0 (motion vector predictor 0) for reference frame 0 of the current block and MVP1 for reference frame 1 based on the prediction-related information (S1210).

[0241] For example, the prediction-related information may include a first reference frame index representing a reference frame 0 for the current block and a second reference frame index representing a reference frame 1, and the decoding device may derive the reference frame 0 and the reference frame 1 based on the first reference frame index and the second reference frame index. That is, for example, the decoding device may obtain the prediction-related information including a first reference frame index representing a reference frame 0 and a second reference frame index representing a reference frame 1, and may derive the reference frame 0 and the reference frame 1 based on the first reference frame index and the second reference frame index.

[0242] Additionally, for example, the prediction-related information may include motion vector information of the current block. For example, the prediction-related information may include a first motion vector candidate index representing MVP0 for reference frame 0 among the motion vector candidate list and a second motion vector candidate index representing MVP1 for reference frame 1 among the motion vector candidate list. The decoding device may derive the MVP0 and the MVP1 based on the first motion vector candidate index and the second motion vector candidate index. For example, the decoding device may construct the motion vector candidate list based on the surrounding blocks of the current block. The surrounding blocks may include spatial surrounding blocks and / or temporal surrounding blocks of the current block.

[0243] The decoding device derives the joint MVD (joint motion vector difference) of the current block based on the joint MVD information (S1220). The decoding device can derive the joint MVD of the current block based on the joint MVD information. The joint MVD information may represent the joint_mvd described above.

[0244] The decoding device derives the MVD0 and MVD1 of the current block based on the joint MVD (S1230). The decoding device can derive the MVD0 and MVD1 of the current block based on the joint MVD.

[0245] For example, at least one of the above MVD0 and the above MVD1 can be derived by scaling the joint MVD.

[0246] For example, at least one of the MVD0 and the MVD1 can be derived by scaling the joint MVD based on a first temporal distance between the current frame and the reference frame 0 and a second temporal distance between the current frame and the reference frame 1.

[0247] For example, if the first temporal distance is greater than or equal to the second temporal distance, the MVD0 can be derived as the joint MVD, and the MVD1 can be derived by scaling the joint MVD based on the first temporal distance and the second temporal distance. For example, if the first temporal distance is greater than or equal to the second temporal distance, the MVD1 can be derived based on the above-described mathematical formula 2.

[0248] Additionally, for example, if the second temporal distance is greater than or equal to the first temporal distance, the MVD1 can be derived as the joint MVD, and the MVD0 can be derived by scaling the joint MVD based on the first temporal distance and the second temporal distance. For example, if the second temporal distance is greater than or equal to the first temporal distance, the MVD0 can be derived based on the above-described mathematical formula 3.

[0249] Alternatively, as another example, the prediction-related information may include scaling coefficient information, and a first scaling coefficient for the MVD0 and a second scaling coefficient for the MVD1 may be derived based on the scaling coefficient information. For example, the scaling coefficient information may indicate one of the scaling coefficient pairs. The scaling coefficient pair indicated by the scaling coefficient information may be derived as the first scaling coefficient for the MVD0 of the current block and the second scaling coefficient for the MVD1. For example, the candidates for the scaling coefficient pair may be as shown in Table 3 above. For example, the MVD0 may be derived by multiplying the joint MVD by the first scaling coefficient, and the MVD1 may be derived by multiplying the joint MVD by the second scaling coefficient.

[0250] The decoding device derives the MV0 (motion vector 0) of the current block based on the MVP0 and the MVD0 (S1240). The decoding device can derive the MV0 of the current block by adding the MVD0 to the MVP0.

[0251] The decoding device derives the MV1 of the current block based on the MVP1 and the MVD1 (S1250). The decoding device can derive the MV1 of the current block by adding the MVD1 to the MVP1.

[0252] The decoding device derives a predicted sample of the current block based on the MV0 and the MV1 (S1260).

[0253] For example, the decoding device may generate a prediction sample of the current block based on a reference sample within the reference frame 0 derived based on the MV0 and a reference sample within the reference frame 1 derived based on the MV1. For example, the decoding device may derive a first prediction block of the current block based on a reference block within the reference frame 0 pointed to by the MV0, derive a second prediction block of the current block based on a reference block within the reference frame 1 pointed to by the MV1, and derive a final prediction block of the current block by weighting the first prediction block and the second prediction block. In this case, as described above, a prediction sample filtering procedure may be further performed on all or some of the prediction samples of the current block depending on the case.

[0254] The decoding device can generate restoration samples based on predicted samples of the current block. For example, the decoding device can generate the restoration samples for the current block based on the residual samples for the current block and the predicted samples. The residual samples for the current block can be generated based on received residual information. Additionally, the decoding device can generate a restoration picture including the restoration samples, for example. As previously described, an in-loop filtering procedure, etc., may be further applied to the restoration picture.

[0255] According to the above-described embodiment(s), MVD0 and MVD1 for composite reference prediction can be derived based on joint MVDs, thereby reducing the amount of motion information data for inter prediction. Additionally, the accuracy of deriving MVD0 and MVD1 can be improved by signaling scaling information for joint MVDs, thereby improving the accuracy of inter prediction and reducing the amount of information data for prediction.

[0256] In the embodiments described above, methods are described based on flowcharts as a series of steps or blocks, but the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps as described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and other steps may be included, or one or more steps of the flowcharts may be omitted without affecting the scope of the embodiments of the present disclosure.

[0257] The method according to the embodiments of the present disclosure described above may be implemented in the form of software, and the encoding device and / or decoding device according to the present disclosure may be included in a device that performs image processing, such as a TV, computer, smartphone, set-top box, display device, etc.

[0258] The embodiments of the present disclosure described above may also be implemented in the form of a recording medium containing computer-executable (program) instructions, such as program modules executed by a computer. The module may be stored in memory and executed by a processor. The memory may be located inside or outside the processor and may be connected to the processor by various well-known means. A computer-readable medium may be any available medium accessible by a computer and includes both volatile and non-volatile media, and both removable and non-removable media. Additionally, a computer-readable medium may include both computer storage media and communication media. A computer storage medium includes both volatile and non-volatile, removable and non-removable media implemented by any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. A communication medium typically includes computer-readable instructions, data structures, program modules, or other data of modulated data signals such as carrier waves, or other transmission mechanisms, and includes any information transmission medium.

[0259] Additionally, the embodiments of the present disclosure described above may be implemented as a computer program (or computer program product) comprising instructions executable by a computer. The computer program includes programmable machine instructions processed by a processor and may be implemented in a high-level programming language, an object-oriented programming language, an assembly language, or a machine language, etc. Additionally, the computer program may be recorded on a tangible computer-readable recording medium (e.g., memory, a hard disk, a magnetic / optical medium, or a Solid-State Drive (SSD), etc.).

[0260] Accordingly, the embodiments of the present disclosure described above may be implemented by executing a computer program as described above by a computing device. The computing device may include at least some of a processor, memory, a storage device, a high-speed interface connected to the memory and a high-speed expansion port, and a low-speed interface connected to a low-speed bus and a storage device. Each of these components may be connected to one another using various buses and may be mounted on a common motherboard or mounted in other suitable ways.

[0261] Here, the processor can process instructions within the computing device, such as instructions stored in memory or storage devices to display graphic information for providing a Graphic User Interface (GUI) on external input and output devices, such as a display connected to a high-speed interface. In another embodiment, multiple processors and / or multiple buses may be used appropriately with multiple memories and memory types. Additionally, the processor may be implemented as a chipset comprising chips including multiple independent analog and / or digital processors.

[0262] In addition, memory stores information within a computing device. For example, memory may consist of volatile memory units or a set thereof. As another example, memory may consist of non-volatile memory units or a set thereof. Furthermore, memory may be other forms of computer-readable media, such as magnetic or optical discs.

[0263] And the storage device can provide a large amount of storage space to the computing device. The storage device may be a computer-readable medium or a configuration containing such a medium, and may include, for example, devices or other configurations within a Storage Area Network (SAN), and may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, flash memory, or other similar semiconductor memory device or device array.

[0264] In addition, the network can be implemented as a wired network such as a Local Area Network (LAN), Wide Area Network (WAN), or Value Added Network (VAN), or as various types of wireless networks such as a mobile radio communication network or a satellite communication network.

[0265] The present disclosure described above has been explained with reference to the embodiments illustrated in the drawings, but this is merely illustrative, and those skilled in the art will understand that various modifications and variations of the embodiments are possible therefrom. That is, the scope of the present disclosure is not limited to the embodiments described above, and various modifications and improvements by those skilled in the art using the basic concepts of the embodiments defined in the following claims also fall within the scope of the embodiments. Accordingly, the true technical scope of protection of the present disclosure should be determined by the technical concept of the appended claims.

Claims

1. In a video decoding method performed by a decoding device, A step of obtaining prediction-related information including joint motion vector difference (MVD) information of the current block; A step of deriving MVP0 (motion vector predictor 0) for reference frame 0 of the current block and MVP1 for reference frame 1 based on the above prediction-related information; A step of deriving the joint MVD (joint motion vector difference) of the current block based on the joint MVD information; A step of deriving MVD0 and MVD1 of the current block based on the joint MVD; A step of deriving the MV0 (motion vector 0) of the current block based on the MVP0 and MVD0; A step of deriving the MV1 of the current block based on the MVP1 and the MVD1; and An image decoding method characterized by including the step of deriving a predicted sample of the current block based on the above MV0 and the above MV1.

2. In Paragraph 1, An image decoding method characterized in that at least one of the above MVD0 and the above MVD1 is derived by scaling the joint MVD.

3. In Paragraph 1, The above prediction-related information includes the composite mode index of the current block, and An image decoding method characterized by signaling the joint MVD information when the above composite mode index indicates that the JOINT_NEWMV mode is applied to the current block.

4. In Paragraph 1, The above prediction-related information includes the composite mode index of the current block, and If the above composite mode index indicates that the NEW_NEWMV mode is applied to the current block, a JOINT_NEWMV mode flag indicating whether the JOINT_NEWMV mode is applied to the current block is signaled, and A video decoding method characterized by signaling the joint MVD information when the above JOINT_NEWMV mode flag indicates that the above JOINT_NEWMV mode is applied to the above current block.

5. In Paragraph 2, An image decoding method characterized by deriving at least one of the MVD0 and the MVD1 by scaling the joint MVD based on a first temporal distance between the current frame and the reference frame 0 and a second temporal distance between the current frame and the reference frame 1.

6. In Paragraph 5, A video decoding method characterized in that when the first temporal distance is greater than or equal to the second temporal distance, the MVD0 is derived as the joint MVD, and the MVD1 is derived by scaling the joint MVD based on the first temporal distance and the second temporal distance.

7. In Paragraph 6, The above MVD1 is derived based on the following mathematical formula, and A video decoding method characterized in that, where td0 is the first temporal distance, td1 is the second temporal distance, and joint_mvd is the joint MVD.

8. In Paragraph 5, A video decoding method characterized in that when the second temporal distance is greater than or equal to the first temporal distance, the MVD1 is derived as the joint MVD, and the MVD0 is derived by scaling the joint MVD based on the first temporal distance and the second temporal distance.

9. In Paragraph 8, The above MVD0 is derived based on the following mathematical formula, and A video decoding method characterized in that, where td0 is the first temporal distance, td1 is the second temporal distance, and joint_mvd is the joint MVD.

10. In Paragraph 2, The above prediction-related information includes scaling factor information, and An image decoding method characterized by deriving a first scaling coefficient for the MVD0 and a second scaling coefficient for the MVD1 based on the above scaling coefficient information.

11. In Paragraph 10, The above MVD0 is derived by multiplying the above joint MVD by the above first scaling coefficient, and An image decoding method characterized in that the above MVD1 is derived by multiplying the joint MVD by the second scaling factor.

12. In a video encoding method performed by an encoding device, A step of deriving MVP0 (motion vector predictor 0) for reference frame 0 of the current block and MVP1 for reference frame 1; A step of deriving MVD0 and MVD1 of the current block based on the joint MVD (joint motion vector difference) of the current block; A step of deriving the MV0 (motion vector 0) of the current block based on the MVP0 and MVD0; A step of deriving the MV1 of the current block based on the MVP1 and the MVD1; A step of deriving a predicted sample of the current block based on the above MV0 and the above MV1; A step of generating prediction-related information including joint MVD information of the current block; and A video encoding method characterized by including the step of encoding video information containing the above-mentioned prediction-related information.

13. In Paragraph 12, An image encoding method characterized in that at least one of the above MVD0 and the above MVD1 is derived by scaling the joint MVD.

14. In Paragraph 12, The above prediction-related information includes the composite mode index of the current block, and A video encoding method characterized by signaling the joint MVD information when the composite mode index indicates that the JOINT_NEWMV mode is applied to the current block.

15. In Paragraph 12, The above prediction-related information includes the composite mode index of the current block, and If the above composite mode index indicates that the NEW_NEWMV mode is applied to the current block, a JOINT_NEWMV mode flag indicating whether the JOINT_NEWMV mode is applied to the current block is signaled, and A video encoding method characterized by signaling the joint MVD information when the above JOINT_NEWMV mode flag indicates that the above JOINT_NEWMV mode is applied to the above current block.

16. In Paragraph 13, A video encoding method characterized by deriving at least one of the MVD0 and the MVD1 by scaling the joint MVD based on a first temporal distance between the current frame and the reference frame 0 and a second temporal distance between the current frame and the reference frame 1.

17. In Paragraph 16, A video encoding method characterized in that when the first temporal distance is greater than or equal to the second temporal distance, the MVD0 is derived as the joint MVD, and the MVD1 is derived by scaling the joint MVD based on the first temporal distance and the second temporal distance.

18. In Paragraph 16, A video encoding method characterized in that when the second temporal distance is greater than or equal to the first temporal distance, the MVD1 is derived as the joint MVD, and the MVD0 is derived by scaling the joint MVD based on the first temporal distance and the second temporal distance.

19. In Paragraph 13, The above prediction-related information includes scaling coefficient information for a first scaling coefficient for the above MVD0 and a second scaling coefficient for the above MVD1, and The above MVD0 is derived by multiplying the above joint MVD by the above first scaling coefficient, and An image encoding method characterized in that the above MVD1 is derived by multiplying the joint MVD by the second scaling factor.

20. In a method for transmitting image data, Acquiring a bitstream generated by a video encoding method, wherein the video encoding method comprises the steps of: deriving an MVP0 (motion vector predictor 0) for reference frame 0 of a current block and an MVP1 for reference frame 1; deriving an MVD0 and an MVD1 of the current block based on a joint MVD (joint motion vector difference) of the current block; deriving an MV0 (motion vector 0) of the current block based on the MVP0 and the MVD0; deriving an MV1 of the current block based on the MVP1 and the MVD1; deriving a prediction sample of the current block based on the MV0 and the MV1; generating prediction-related information including joint MVD information of the current block; and encoding video information including the prediction-related information; and A transmission method characterized by including the step of transmitting image data including the bitstream above.