Motion vector prediction-based image / video coding method and device
By employing motion vector prediction-based inter prediction and efficient signaling techniques, the method addresses the challenges of compressing high-resolution and immersive media, achieving improved coding efficiency and reduced costs.
Patent Information
- Application Number
- JP2025030343
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-06-13
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The increasing demand for high-resolution and high-quality images/videos, such as 4K or UHD, leads to higher data transmission and storage costs, and existing compression technologies struggle to efficiently handle the diverse characteristics of immersive media like VR, AR, and game images.
A method and apparatus for improving video coding efficiency through motion vector prediction-based inter prediction, efficient signaling of motion vector differences, and dual prediction techniques, which reduce bit waste and complexity in the coding system.
The proposed solution enhances overall video compression efficiency, reduces the complexity of signaling information, and minimizes bit waste, thereby lowering transmission and storage costs while supporting high-resolution and diverse media formats.
Smart Images

Figure 2025084865000001_ABST
Abstract
Description
Technical Field
[0001] This specification relates to a motion vector prediction-based video / video coding method and apparatus.
Background Art
[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or UHD (Ultra High Definition) images / videos of 8K or higher has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to the existing image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line, or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.
[0003] In addition, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) contents, and holograms have been increasing, and the broadcast of images / videos having image characteristics different from those of real images, such as game images, has been increasing.
[0004] Accordingly, there is a need for a highly efficient image / video compression technology to effectively compress, transmit, store, and reproduce information of high-resolution and high-quality images / videos having various characteristics as described above.
Summary of the Invention
Problems to be Solved by the Invention
[0005] According to an embodiment of this document, a method and apparatus for increasing video / video coding efficiency are provided.
[0006] According to an embodiment of this document, a method and apparatus for efficiently performing inter prediction in a video / video coding system are provided.
[0007] According to one embodiment of this document, a method and apparatus for performing motion vector prediction-based inter prediction are provided.
[0008] According to one embodiment of this document, a method and apparatus for signaling information related to motion vector difference in inter prediction are provided.
[0009] According to one embodiment of this document, when dual prediction is applied to a current block, a method and apparatus for signaling prediction-related information are provided.
[0010] According to one embodiment of this document, a method and apparatus for signaling an L1 motion vector difference zero flag and / or an SMVD flag are provided.
[0011] According to one embodiment of this document, a video / video decoding method executed by a decoding apparatus is provided.
[0012] According to one embodiment of this document, a decoding apparatus for executing video / video decoding is provided.
[0013] According to one embodiment of this document, a video / video encoding method executed by an encoding apparatus is provided.
[0014] According to one embodiment of this document, an encoding apparatus for executing video / video encoding is provided.
[0015] According to one embodiment of this document, a computer-readable digital storage medium storing encoded video / video information generated by the video / video encoding method disclosed in at least one of the embodiments of this document is provided.
[0016] According to one embodiment of this document, there is provided a computer-readable digital storage medium storing encoded information for causing a decoding device to execute a video / video decoding method disclosed in at least one of the embodiments of this document, or encoded video / video information.
Advantages of the Invention
[0017] According to this document, the overall video / video compression efficiency can be increased.
[0018] According to this document, information regarding motion vector differences can be efficiently signaled.
[0019] According to this document, when dual prediction is applied to the current block, the L1 motion vector difference can be efficiently derived.
[0020] According to this document, the information used to derive the L1 motion vector difference can be efficiently signaled to reduce the complexity of the coding system.
[0021] According to this document, bit waste that can occur in information signaling regarding the inter prediction mode can be avoided, and the overall coding efficiency can be increased.
[0022] The effects obtainable through a specific example of this document are not limited to the effects listed above. For example, there can be various technical effects that can be understood or induced by a person having ordinary skill in the related art from this document. Accordingly, the specific effects of this document can include various effects that can be understood or induced from the technical features of this document, without being limited to those explicitly described in this document.
Brief Description of the Drawings
[0023]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Best Mode for Carrying Out the Invention
[0024] This document can be modified in various ways and can have various embodiments. Specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are merely used to describe specific embodiments and are not used with the intention of limiting the technical idea of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this specification are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should be understood that the presence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof, etc., is not precluded in advance.
[0025] On the other hand, each configuration in the drawings described in this document is independently illustrated for the convenience of explaining different characteristic functions, and it does not mean that each configuration is implemented by separate hardware or separate software. For example, among the configurations, two or more configurations can be combined to form one configuration, and one configuration can also be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are included in the scope of rights of this document as long as they do not deviate from the essence of this document.
[0026] This document relates to video / video coding. For example, the methods / examples disclosed in this document can be applied to the methods disclosed in the VVC (Versatile Video Coding) standard. Also, the methods / examples disclosed in this document can be applied to the methods disclosed in the EVC (essential video coding) standard, the AV1 (AOMedia Video 1) standard, the AVS2 (2nd generation of audio video coding standard), or next-generation video / video coding standards (e.g., H.267 or H.268, etc.).
[0027] This document presents various examples related to video / video coding, and unless otherwise stated, the examples can also be executed in combination with each other.
[0028] Hereinafter, the examples of this document will be described with reference to the attached drawings. Hereinafter, the same reference numerals can be used for the same components in the drawings, and duplicate descriptions for the same components can be omitted.
[0029] FIG. 1 schematically shows an example of a video / video coding system to which the examples of this document can be applied.
[0030] As shown in FIG. 1, the video / video coding system can include a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or a network.
[0031] The source device can include a video source, an encoding device, and a transmission unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device can be called a video / video encoding device, and the decoding device can be called a video / video decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can also include a display unit, and the display unit can also be composed of a separate device or an external component.
[0032] The video source can obtain video / video through processes such as video / video capture, synthesis, or generation. The video source can include a video / video capture device and / or a video / video generation device. The video / video capture device can include, for example, one or more cameras, a video / video archive containing previously captured video / video, etc. The video / video generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / video. For example, virtual video / video can be generated through a computer or the like, in which case the video / video capture process can be replaced by a process in which related data is generated.
[0033] The encoding device can encode the input video / video. The encoding device can execute a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / video information) can be output in the form of a bitstream.
[0034] The transmitting unit can transmit the encoded video / video information or data output in bitstream form to the receiving unit of the receiving device via a digital storage medium or network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.
[0035] The decoding device can decode the video / video by executing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0036] The renderer can render the decoded video / video. The rendered video / video can be displayed via the display unit.
[0037] In this document, video can mean a collection of a series of images over time. Picture generally means a unit representing one image at a specific time period, and slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile can include one or more CTUs (Coding Tree Units). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width that can be specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a width specified by syntax elements in the picture parameter set and a height that can be the same as the height of the picture.A tile scan can represent a specific sequential ordering of CTUs partitioning a picture, where the CTUs can be ordered consecutively in a CTU raster scan within a tile, and tiles in a picture can be ordered consecutively in a raster scan of the tiles of the picture. A slice can include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that may be exclusively contained in a single NAL unit.
[0038] On the other hand, a picture can be divided into two or more sub - pictures. A sub - picture can be a rectangular region of one or more slices within a picture.
[0039] A pixel or pel can mean the smallest unit that makes up a picture (or video). Also, the term "sample" can be used as the term corresponding to a pixel. A sample can generally represent a pixel or the value of a pixel, and can also represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.
[0040] A unit can indicate the basic unit of video processing. A unit can include at least one of a specific area of a picture and information related to that area. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block can include a sample (or, sample array) consisting of M columns and N rows, or a set (or, array) of transform coefficients.
[0041] In this document, "A or B" can mean "only A", "only B", or "both A and B". In other words, in this document, "A or B" can be interpreted as "A and / or B". For example, in this document, "A, B or C" can mean "only A", "only B", "only C", or "any combination of A, B and C".
[0042] The slash ( / ) or comma used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B or C".
[0043] In this document, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, in this document, expressions such as "at least one of A or B" or "at least one of A and / or B" can be analyzed in the same way as "at least one of A and B".
[0044] Also, in this document, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0045] Also, the parentheses used in this document can mean "for example". Specifically, when it is shown as "prediction (intra prediction)", it may be that "intra prediction" is proposed as an example of "prediction". In other words, "prediction" in this document is not limited to "intra prediction", and it may be that "intra prediction" is proposed as an example of "prediction". Also, when it is shown as "prediction (i.e., intra prediction)", it may be that "intra prediction" is proposed as an example of "prediction".
[0046] In this document, the technical features individually described within one drawing can be embodied individually or simultaneously.
[0047] FIG. 2 is a diagram schematically illustrating the configuration of a video / video encoding apparatus to which the embodiments of this document can be applied. Hereinafter, the encoding apparatus can include a video encoding apparatus and / or a video encoding apparatus. Further, the video encoding method / apparatus can include the video encoding method / apparatus. Alternatively, the video encoding method / apparatus can include the video encoding method / apparatus.
[0048] As shown in FIG. 2, the encoding apparatus 200 can include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter predictor 221 and an intra predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor (231). The adder 250 can be called a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to the embodiment. Further, the memory 270 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.
[0049] The image segmentation unit 210 can divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into a plurality of coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and the binary-tree structure and / or the ternary structure can be applied thereafter. Or, the binary-tree structure can also be applied first. The coding procedure according to the present disclosure can be performed based on the final coding unit that is no longer divided. In this case, based on the coding efficiency according to the image characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth so that the coding unit with an optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, conversion, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can each be divided or partitioned from the above-described final coding unit.The prediction unit can be a unit of sample prediction, and the conversion unit can be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.
[0050] The unit can, in some cases, be used interchangeably with terms such as block or area. In general, an M×N block can represent a set of samples or transform coefficients, etc., consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel for one picture (or image).
[0051] The encoding device 200 can subtract the predicted signal (predicted block, predicted sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from the input video signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the predicted signal (predicted block, predicted sample array) from the input video signal (original block, original sample array) within the encoder 200 can be called the subtraction unit 231. The prediction unit can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit can generate various information related to prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit it to the entropy encoding unit 240. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.
[0052] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to the current block depending on the prediction mode, or can also be located far away. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of fineness of the prediction direction. However, this is an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the adjacent block.
[0053] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring blocks can be called by names such as collocated reference blocks and collocated CUs (col CUs), and the reference picture including the temporal neighboring blocks can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and by signaling the motion vector difference, the motion vector of the current block can be indicated.
[0054] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply not only intra prediction or inter prediction for the prediction of one block, but also can apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can be based on the intra block copy (IBC) prediction mode for the prediction of a block, or can be based on the palette mode. The IBC prediction mode or the palette mode can be used for content video / movie coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information regarding the palette table and the palette index.
[0055] The prediction signal generated through the prediction unit (inter prediction unit 221) and / or the intra prediction unit 222 can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph - Based Transform), or CNT (Conditionally Non - linear Transform). Here, GBT means the conversion obtained from a graph when representing the relationship information between pixels as a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to pixel blocks having the same size of a square, and can also be applied to blocks of variable size that are not square.
[0056] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be called residual information. The quantization unit 233 can reorder the block-form quantized transform coefficients in a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / video information) can be transmitted or stored in units of NAL (network abstraction layer) units in bitstream form. The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information can further include general constraint information. In this document, the information and / or syntax elements transmitted / signaled from the encoding device to the decoding device can be included in the video / video information. The video / video information can be encoded through the encoding procedure described above and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured as an internal / external element of the encoding device 200 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit can also be included in the entropy encoding unit 240.
[0057] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transformation unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 250 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 250 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and, as will be described later, can also be used for inter prediction of the picture after passing through filtering.
[0058] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture encoding and / or restoration process.
[0059] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and can store the modified restored picture in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 260 can generate various information related to filtering as described later in the description of each filtering method and transmit it to the entropy encoding unit 240. The information related to filtering can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0060] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and can also improve the encoding efficiency.
[0061] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the blocks in which the motion information in the current picture has been derived (or encoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.
[0062] FIG. 3 is a diagram schematically illustrating a configuration of a video / video decoding apparatus to which an embodiment of the present document can be applied. Hereinafter, the decoding apparatus can include a video decoding apparatus and / or a video decoding apparatus. Further, the video decoding method / apparatus can include the video decoding method / apparatus. Or, the video decoding method / apparatus can include the video decoding method / apparatus.
[0063] As shown in FIG. 3, the decoding apparatus 300 can include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filtering unit 350, and a memory 360. The predictor 330 can include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 321. The entropy decoding unit 310, the residual processing unit 320, the prediction unit 330, the addition unit 340, and the filtering unit 350 described above can be configured by one hardware component (for example, a decoder chipset or a processor) according to an embodiment. Further, the memory 360 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 360 as an internal / external component.
[0064] If a bitstream including image / video information is input, the decoding device 300 can restore an image corresponding to the process in which the image / video information was processed by the encoding device of FIG. 2. For example, the decoding device 300 can derive units / blocks based on block division related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided according to a quad tree structure, a binary tree structure, and / or a ternary tree structure from a coding tree unit or a maximum coding unit. One or more transform units can be derived from the coding unit. Then, the restored image signal decoded and output via the decoding device 300 can be reproduced via a playback device.
[0065] The decoding device 300 can receive the signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information can further include general constraint information. The decoding device can decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the decoding target syntax element information adjacent to and the decoding information of the decoding target block or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin according to the determined context model, and executes arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit (inter prediction unit 332 and intra prediction unit 331), and the residual value for which entropy decoding is executed by the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive a residual signal (residual block, residual sample, residual sample array). Also, the information related to filtering among the information decoded by the entropy decoding unit 310 can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / video / picture decoding device, and the decoding device can also be classified into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder can include the entropy decoding unit 310, and the sample decoder can include at least one of the inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.
[0066] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) to obtain transform coefficients.
[0067] In the inverse transform unit 322, the transform coefficients are inverse transformed to obtain a residual signal (residual block, residual sample array).
[0068] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.
[0069] The prediction unit 320 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply not only intra prediction or inter prediction for the prediction of one block, but also can apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode can be used for content video / moving picture coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index can be included in and signaled in the video / video information.
[0070] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to or away from the current block depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also utilize the prediction mode applied to an adjacent block to determine the prediction mode to be applied to the current block.
[0071] The inter prediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 331 can construct a motion information candidate list based on the adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be executed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.
[0072] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to a prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block.
[0073] The addition unit 340 can be called a restoration unit or a restoration block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block in the current picture, and as will be described later, it can also be output after filtering, or can be used for inter prediction of the next picture.
[0074] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.
[0075] The filtering unit 350 can apply filtering to the restoration signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 60, specifically, the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0076] The (modified) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the blocks for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and can be transmitted to the intra prediction unit 331.
[0077] In this specification, the embodiments described in the filtering unit 260, the inter prediction unit 221, and the intra prediction unit 222 of the encoding apparatus 200 can also be applied to the filtering unit 350, the inter prediction unit 332, and the intra prediction unit 331 of the decoding apparatus 300 so as to be the same or corresponding thereto.
[0078] As described above, prediction is performed to increase the compression efficiency in performing video coding. Thereby, a predicted block including predicted samples for a current block which is a block to be coded can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived in the same manner in the encoding apparatus and the decoding apparatus, and the encoding apparatus can increase the image coding efficiency by signaling information (residual information) regarding the residual between the original block and the predicted block, which is not the original sample value of the original block, to the decoding apparatus. The decoding apparatus can derive a residual block including residual samples based on the residual information, add the residual block and the predicted block to generate a restored block including restored samples, and generate a restored picture including the restored block.
[0079] The residual information can be generated through conversion and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, execute a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, execute a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, and thus signal the relevant residual information (via a bitstream) to a decoding device. Here, the residual information can include information such as value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can execute an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse transform the quantized conversion coefficients for reference in inter prediction of subsequent pictures to derive a residual block and generate a restored picture based on this.
[0080] The prediction unit of the encoding device / decoding device can derive a prediction sample by performing inter prediction in block units. Inter prediction can be a prediction derived in a manner that is dependent on data elements (e.g., sample values or motion information) of picture(s) other than the current picture (Inter prediction can be a prediction derived in a manner that is dependent on data elements (e.g., sample values or motion information) of picture(s) other than the current picture). When inter prediction is applied to the current block, a predicted block (predicted sample array) for the current block can be induced based on a reference block (reference sample array) specified by a motion vector on a reference picture pointed to by a reference picture index. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the adjacent block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When inter prediction is applied, the adjacent block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block can be called by names such as a collocated reference block and a collocated CU (colCU), and the reference picture including the temporal neighboring block can also be called a collocated picture (colPic).For example, a motion information candidate list can be configured based on adjacent blocks of a current block, and flag or index information indicating which candidate is selected (used) can be signaled to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the motion information of the current block can be the same as that of the selected adjacent block. In the case of skip mode, unlike merge mode, a residual signal cannot be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the selected adjacent block is used as a motion vector predictor, and a motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.
[0081] The motion information can include L0 motion information and / or L1 motion information depending on the inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). The motion vector in the L0 direction can be called the L0 motion vector or MVL0, and the motion vector in the L1 direction can be called the L1 motion vector or MVL1. The prediction based on the L0 motion vector can be called L0 prediction, the prediction based on the L1 motion vector can be called L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector can be called bi (Bi) prediction. Here, the L0 motion vector can represent a motion vector related to the reference picture list L0 (L0), and the L1 motion vector can represent a motion vector related to the reference picture list L1 (L1). The reference picture list L0 can include pictures before the current picture in the output order as reference pictures, and the reference picture list L1 can include pictures after the current picture in the output order. The previous picture can be called a forward (reference) picture, and the subsequent picture can be called a backward (reference) picture. The reference picture list L0 can further include pictures after the current picture in the output order as reference pictures. In this case, the previous pictures are first indexed within the reference picture list L0, and the subsequent pictures can be indexed thereafter. The reference picture list L1 can further include pictures before the current picture in the output order as reference pictures. In this case, the subsequent pictures are first indexed within the reference picture list 1, and the previous pictures can be indexed thereafter. Here, the output order can correspond to the POC (picture order count) order (order).
[0082] The video / video encoding procedure based on inter prediction can generally include, for example, the following.
[0083] FIG. 4 shows an example of an inter prediction-based video / video encoding method.
[0084] The encoding device performs inter prediction on the current block (S400). The encoding device can derive the inter prediction mode and motion information of the current block and generate a prediction sample of the current block. Here, the inter prediction mode determination, motion information derivation, and prediction sample generation procedures can be performed simultaneously, or any one of the procedures can be performed first than the other procedures. For example, the inter prediction unit of the encoding device can include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit can determine the prediction mode for the current block, the motion information derivation unit can derive the motion information of the current block, and the prediction sample derivation unit can derive the prediction sample of the current block. For example, the inter prediction unit of the encoding device can search for a block similar to the current block within a certain area (search area) of the reference picture through motion estimation, and derive a reference block whose difference from the current block is the smallest or below a certain criterion. Based on this, a reference picture index indicating the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine the mode applied to the current block among various prediction modes. The encoding device can compare the RD costs for the various prediction modes and determine the optimal prediction mode for the current block.
[0085] For example, when the skip mode or the merge mode is applied to the current block, the encoding device constructs a merge candidate list described below, and among the reference blocks pointed to by the merge candidates included in the merge candidate list, a reference block whose difference from the current block is the smallest or below a certain criterion can be derived. In this case, a merge candidate associated with the derived reference block is selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.
[0086] As another example, when the (A)MVP mode is applied to the current block, the encoding device constructs an (A)MVP candidate list described below, and among the mvp (motion vector predictor) candidates included in the (A)MVP candidate list, the motion vector of the selected mvp candidate can be used as the mvp of the current block. In this case, for example, the motion vector pointing to the reference block derived by the above-described motion estimation can be used as the motion vector of the current block, and among the mvp candidates, the mvp candidate having the smallest difference from the motion vector of the current block can be the selected mvp candidate. An MVD (motion vector difference), which is the difference obtained by subtracting the mvp from the motion vector of the current block, can be derived. In this case, information regarding the MVD can be signaled to the decoding device. Also, when the (A)MVP mode is applied, the value of the reference picture index can be configured with reference picture index information and separately signaled to the decoding device.
[0087] The encoding device can derive a residual sample based on the prediction sample (S410). The encoding device can derive the residual sample by comparing the original sample of the current block with the prediction sample.
[0088] The encoding device encodes video information including prediction information and residual information (S420). The encoding device can output the encoded video information in the form of a bitstream. The prediction information can include prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and information related to motion information as information related to the prediction procedure. The information related to the motion information can include candidate selection information (e.g., merge index, mvp flag, or mvp index) which is information for deriving a motion vector. Also, the information related to the motion information can include information related to the above-mentioned MVD and / or reference picture index information. Also, the information related to the motion information can include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information related to the residual sample. The residual information can include information related to the quantized transform coefficients for the residual sample.
[0089] The output bitstream can be stored in a (digital) storage medium and transmitted to the decoding device, or can be transmitted to the decoding device via a network.
[0090] On the one hand, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference sample and the residual sample. This is to derive the same prediction result in the encoding device as that performed in the decoding device, and through this, the coding efficiency can be improved. Therefore, the encoding device can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in the memory and utilize it as a reference picture for inter prediction. As described above, loop filtering procedures and the like can be further applied to the reconstructed picture.
[0091] The video / video decoding procedure based on inter prediction can generally include, for example, the following.
[0092] FIG. 5 shows an example of an inter prediction-based video / video decoding method.
[0093] As shown in FIG. 5, the decoding device can perform operations corresponding to the operations performed in the encoding device. The decoding device can perform prediction on the current block based on the received prediction information and derive prediction samples.
[0094] Specifically, the decoding device can determine the prediction mode for the current block based on the received prediction information (S500). The decoding device can determine which inter prediction mode is applied to the current block based on the prediction mode information in the prediction information.
[0095] For example, based on the merge flag, it can be determined whether the merge mode is applied to the current block, or whether (A) the MVP mode is determined. Alternatively, based on the mode index, one of various inter-prediction mode candidates can be selected. The inter-prediction mode candidates can include a skip mode, a merge mode, and / or (A) the MVP mode, or can include various inter-prediction modes described later.
[0096] The decoding device derives motion information of the current block based on the determined inter-prediction mode (S510). For example, when the skip mode or the merge mode is applied to the current block, the decoding device constructs a merge candidate list described later, and can select one of the merge candidates included in the merge candidate list. The selection can be performed based on the selection information (merge index) described above. Using the motion information of the selected merge candidate, the motion information of the current block can be derived. The motion information of the selected merge candidate can be used as the motion information of the current block.
[0097] As another example, when the (A)MVP mode is applied to the current block, the decoding device constructs an (A)MVP candidate list described later, and among the mvp (motion vector predictor) candidates included in the (A)MVP candidate list, the motion vector of the selected mvp candidate can be used as the mvp of the current block. The selection can be performed based on the selection information (mvp flag or mvp index) described above. In this case, the MVD of the current block can be derived based on the information regarding the MVD, and based on the mvp of the current block and the MVD, the motion vector of the current block can be derived. Also, the reference picture index of the current block can be derived based on the reference picture index information. The picture pointed to by the reference picture index within the reference picture list regarding the current block can be derived as the reference picture to be referred to for the inter prediction of the current block.
[0098] On the other hand, as will be described later, the motion information of the current block can be derived without constructing a candidate list, and in this case, the motion information of the current block can be derived according to the procedure disclosed in the prediction mode described later. In this case, the candidate list construction as described above can be omitted.
[0099] The decoding device can generate a prediction sample for the current block based on the motion information of the current block (S520). In this case, the reference picture is derived based on the reference picture index of the current block, and the prediction sample of the current block can be derived using the samples of the reference block pointed to by the motion vector of the current block on the reference picture. In this case, as will be described later, a prediction sample filtering procedure can be further performed for all or part of the prediction samples of the current block depending on the case.
[0100] For example, the inter prediction unit of the decoding device can include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. Based on the prediction mode information received by the prediction mode determination unit, it determines the prediction mode for the current block. Based on the information related to the motion information received by the motion information derivation unit, it derives the motion information (such as motion vectors and / or reference picture indices, etc.) of the current block. And the prediction sample derivation unit can derive the prediction sample of the current block.
[0101] The decoding device generates a residual sample for the current block based on the received residual information (S530). The decoding device can generate a restored sample for the current block based on the prediction sample and the residual sample, and generate a restored picture based on this (S540). As described above, an in-loop filtering procedure or the like can be further applied to the restored picture later.
[0102] FIG. 6 exemplarily shows an inter prediction procedure.
[0103] As shown in FIG. 6, the inter prediction procedure can include an inter prediction mode determination step, a motion information derivation step according to the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information. The inter prediction procedure can be performed by the encoding device and the decoding device as described above. In this document, the coding device can include the encoding device and / or the decoding device.
[0104] As shown in FIG. 6, the coding device determines an inter prediction mode for the current block (S600). A variety of inter prediction modes can be used for predicting the current block within a picture. For example, various modes such as a merge mode, a skip mode, an MVP (motion vector prediction) mode, an Affine mode, a sub-block merge mode, an MMVD (merge with MVD) mode, etc. can be used. A DMVR (Decoder side motion vector refinement) mode, an AMVR (adaptive motion vector resolution) mode, Bi-prediction with CU-level weight (BCW), Bi-directional optical flow (BDOF), etc. can be used additionally or instead as accompanying modes. The Affine mode can also be called an affine motion prediction mode. The MVP mode can also be called an AMVP (advanced motion vector prediction) mode. In this document, some modes and / or motion information candidates derived by some modes can also be included as any one of the motion information related candidates of other modes. For example, an HMVP candidate can be added as a merge candidate of the merge / skip mode, or can be added as an mvp candidate of the MVP mode. When the HMVP candidate is used as a motion information candidate of the merge mode or the skip mode, the HMVP candidate can be called an HMVP merge candidate.
[0105] Prediction mode information indicating the inter prediction mode of the current block can be signaled from the encoding device to the decoding device. The prediction mode information can be included in a bitstream and received by the decoding device. The prediction mode information can include index information indicating any one of a number of candidate modes. Alternatively, the inter prediction mode can also be indicated via hierarchical signaling of flag information. In this case, the prediction mode information can include one or more flags. For example, a skip flag is signaled to indicate whether to apply the skip mode, and when the skip mode is not applied, a merge flag is signaled to indicate whether to apply the merge mode. When the merge mode is not applied, it can be indicated that the MVP mode is applied, or additional flags for further classification can also be signaled. The affine mode can also be signaled as an independent mode, or can be signaled as a mode subordinate to the merge mode or the MVP mode, etc. For example, the affine mode can include an affine merge mode and an affine MVP mode.
[0106] The coding device derives motion information for the current block (S610). The motion information derivation can be derived based on the inter prediction mode.
[0107] The coding device can perform inter prediction using the motion information of the current block. The coding device can derive the optimal motion information for the current block through a motion estimation procedure. For example, the coding device can search for a highly correlated similar reference block within a determined search range in the reference picture in terms of fractional pixels using the original block in the original picture for the current block, and derive motion information through this. The similarity of the blocks can be derived based on the difference in phase-based sample values. For example, the similarity of the blocks can be calculated based on the SAD between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, the motion information can be derived based on the reference block with the smallest SAD within the search area. The derived motion information can be signaled to the decoding device in various ways based on the inter prediction mode.
[0108] The coding device performs inter prediction based on the motion information for the current block (S620). The coding device can derive the prediction sample(s) for the current block based on the motion information. The current block including the prediction sample can be called a predicted block.
[0109] On the one hand, information indicating whether the above-described list 0 (L0) prediction, list 1 (L1) prediction, or bi-prediction is used for the current block (current coding unit) can be signaled. The said information can be called motion prediction direction information, inter prediction direction information, or inter prediction indication information, and can be configured / encoded / signaled, for example, in the form of an inter_pred_idc syntax element. That is, the inter_pred_idc syntax element can indicate whether the above-described list 0 (L0) prediction, list 1 (L1) prediction, or bi-prediction is used for the current block (current coding unit). In this document, for the sake of convenience of explanation, the inter prediction type (L0 prediction, L1 prediction, or BI prediction) indicated by the inter_pred_idc syntax element can be referred to as the motion prediction direction. L0 prediction can also be represented as pred_L0, L1 prediction as pred_L1, and bi-prediction as pred_BI.
[0110] One picture can include one or more slices. A slice can have one of the slice types including intra (I) slice, predictive (P) slice, and bi-predictive (B) slice. The said slice type can be indicated based on slice type information. For blocks within an I slice, inter prediction is not used for prediction, and only intra prediction can be used. Of course, in this case, it is also possible to code and signal the original sample value without prediction. For blocks within a P slice, intra prediction or inter prediction can be used, and when inter prediction is used, only uni-prediction can be used. On the other hand, for blocks within a B slice, intra prediction or inter prediction can be used, and when inter prediction is used, up to bi-prediction can be used.
[0111] L0 and L1 can include reference pictures that have been encoded / decoded prior to the current picture. For example, L0 can include reference pictures before and / or after the current picture in the POC order, and L1 can include reference pictures after and / or before the current picture in the POC order. In this case, a relatively lower reference picture index can be assigned to the reference pictures in L0 that are prior to the current picture in the POC order, and a relatively lower reference picture index can be assigned to the reference pictures in L1 that are after the current picture in the POC order. In the case of a B slice, bi-prediction can be applied. In this case, uni-directional bi-prediction can also be applied, or bi-directional bi-prediction can be applied. Bi-directional bi-prediction can be called true bi-prediction.
[0112] Specifically, for example, information regarding the inter-prediction mode of the current block can be coded and signaled at the level of CU (CU syntax), etc., or can be implicitly determined depending on conditions. In this case, for some modes, it is signaled explicitly, and for some of the remaining modes, it is derived implicitly.
[0113] FIG. 7 exemplarily shows the spatial adjacent blocks used for deriving motion information candidates in the MVP mode.
[0114] The MVP (Motion vector prediction) mode can be called the AMVP (Advanced motion vector prediction) mode. When the MVP mode is applied, a motion vector predictor (mvp) candidate list can be generated by using the motion vectors of the restored spatially adjacent blocks (for example, it can include the adjacent blocks in FIG. 5) and / or the motion vectors corresponding to the temporally adjacent blocks (or Col blocks). That is, the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the temporally adjacent blocks can be used as motion vector predictor candidates. The spatially adjacent blocks can include the lower left corner adjacent block A0, the left adjacent block A1, the upper right corner adjacent block B0, the upper adjacent block B1, and the upper left corner adjacent block B2 of the current block.
[0115] When dual prediction is applied, an mvp candidate list for deriving L0 motion information and an mvp candidate list for deriving L1 motion information can be generated and used individually. The prediction information (or information related to prediction) described above can include selection information (e.g., MVP flag or MVP index) indicating the optimal motion vector predictor candidate selected from among the motion vector predictor candidates included in the list. At this time, the prediction unit can select the motion vector predictor of the current block from among the motion vector predictor candidates included in the motion vector candidate list using the selection information. The prediction unit of the encoding device can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, and encode and output this in the form of a bit stream. That is, the MVD can be obtained as the value obtained by subtracting the motion vector predictor from the motion vector of the current block. At this time, the prediction unit of the decoding device can obtain the motion vector difference included in the information related to the prediction, and derive the motion vector of the current block through the addition of the motion vector difference and the motion vector predictor. The prediction unit of the decoding device can obtain or derive a reference picture index indicating a reference picture, etc. from the information related to the prediction.
[0116] FIG. 8 schematically shows a method of constructing an MVP candidate list according to this document.
[0117] As shown in FIG. 8, in one embodiment, first, a spatial candidate block for motion vector prediction is searched, and an available spatial (MVP) candidate can be inserted into a prediction candidate list (MVP candidate list) (S810). In this case, adjacent blocks can be divided into two groups to derive candidates. For example, one (MVP) candidate can be derived from group A including the lower left corner adjacent block A0 and the left adjacent block A1 of the current block, and one (MVP) candidate can be derived from group B including the upper right corner adjacent block B0, the upper adjacent block B1, and the upper left corner adjacent block B2 of the current block. The MVP candidate derived from group A can be called mvpA, and the MVP candidate derived from group B can be called mvpB. If all adjacent blocks in the group are not available or are intra-coded, no (MVP) candidate may be derived in the corresponding group.
[0118] Thereafter, in one embodiment, it can be determined whether the number of spatial candidates is less than 2 (S820). For example, in one embodiment, when the number of spatial candidates is less than 2, a temporal candidate derived by searching temporal adjacent blocks can be additionally inserted into the prediction candidate list (S830). The temporal candidate derived from the temporal adjacent blocks can be called mvpCol.
[0119] On the other hand, when the MVP mode is applied, the reference picture index can be explicitly signaled. In this case, it can be signaled by being divided into a reference picture index (refidxL0) for L0 prediction and a reference picture index (refidxL1) for L1 prediction. For example, when the MVP mode is applied and bi-prediction (BI prediction) is applied, information regarding refidxL0 and information regarding refidxL1 can both be signaled.
[0120] When dual prediction is applied, an mvp candidate list for deriving L0 motion information and an mvp candidate list for deriving L1 motion information can be individually generated and used. The prediction information (or information related to prediction) described above can include selection information (e.g., MVP flag or MVP index) indicating the optimal motion vector predictor candidate selected from among the motion vector predictor candidates included in the list. At this time, the prediction unit can select the motion vector predictor of the current block from among the motion vector predictor candidates included in the motion vector candidate list using the selection information. The prediction unit of the encoding device can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, and can encode this and output it in the form of a bit stream. That is, the MVD can be obtained as the value obtained by subtracting the motion vector predictor from the motion vector of the current block. At this time, the prediction unit of the decoding device can obtain the motion vector difference included in the information related to the prediction, and can derive the motion vector of the current block through the addition of the motion vector difference and the motion vector predictor. The prediction unit of the decoding device can obtain or derive a reference picture index indicating a reference picture, etc. from the information related to the prediction.
[0121] When the MVP mode is applied, as described above, information related to the MVD derived from the encoding device can be signaled to the decoding device. The information related to the MVD can include, for example, information representing the x and y components of the MVD absolute value and the sign. In this case, information indicating whether the MVD absolute value is greater than 0 and whether it is greater than 1, and information representing the remaining MVD can be signaled step by step. For example, the information indicating whether the MVD absolute value is greater than 1 can be signaled only when the value of the flag information indicating whether the MVD absolute value is greater than 0 is 1.
[0122] For example, information regarding MVD can be configured with a syntax as shown in Table 1, encoded by an encoding device, and signaled to a decoding device.
[0123] [Table 1]
[0124] For example, in Table 1, the abs_mvd_greater0_flag syntax element can represent information on whether the difference (MVD) is greater than 0, and the abs_mvd_greater1_flag syntax element can represent information on whether the difference (MVD) is greater than 1. Also, the abs_mvd_minus2 syntax element can represent information on the value obtained by subtracting 2 from the difference (MVD), and the mvd_sign_flag syntax element can represent information on the sign of the difference (MVD). Further, in Table 1, "[0]" of each syntax element can represent information for L0, and "[1]" can represent information for L1.
[0125] For example, MVD[compIdx] can be derived based on abs_mvd_greater0_flag[compIdx] * (abs_mvd_minus2[compIdx] + 2) * (1 - 2 * mvd_sign_flag[compIdx]). Here, compIdx (or cpIdx) represents the index of each component and can have a value of 0 or 1. CompIdx 0 can refer to the x component, and compIdx 1 can refer to the y component. However, this is for illustration purposes, and values for each component can also be represented using other coordinate systems instead of the x, y coordinate system.
[0126] On the other hand, the MVD for L0 prediction (MVDL0) and the MVD for L1 prediction (MVDL1) can be separated and signaled, and the information regarding the MVD can include the information regarding MVDL0 and / or the information regarding MVDL1. For example, when the MVP mode is applied to the current block and the BI prediction is applied, both the information regarding MVDLO and the information regarding MVDL1 can be signaled.
[0127] FIG. 9 is a diagram for explaining a symmetric MVD.
[0128] On the other hand, when the bi-prediction (BI prediction) is applied, a symmetric MVD can also be used in consideration of coding efficiency. Here, the symmetric MVD can also be called an SMVD. In this case, some of the signaling of the motion information can be omitted. For example, when the symmetric MVD is applied to the current block, the information regarding refidxL0, the information regarding refidxL1, and the information regarding MVDL1 can be internally derived without being signaled from the encoding device to the decoding device. For example, when the MVP mode and the BI prediction are applied to the current block, flag information (e.g., symmetric MVD flag information or sym_mvd_flag syntax element) indicating whether to apply the symmetric MVD can be signaled. When the value of the flag information is 1, the decoding device can determine that the symmetric MVD is applied to the current block.
[0129] When the symmetric MVD mode is applied (i.e., when the value of the symmetric MVD flag information is 1), information regarding mvp_l0_flag, mvp_l1_flag, and MVDL0 can be explicitly signaled. As described above, signaling of information regarding refidxL0, refidxL1, and MVDL1 can be omitted and can be derived internally. For example, refidxL0 can be derived as an index that points to the previous reference picture closest to the current picture in POC order within reference picture list 0 (which can be called list 0 or L0). refidxL1 can be derived as an index that points to the subsequent reference picture closest to the current picture in POC order within reference picture list 1 (which can be called list 1 or L1). Or, for example, both refidxL0 and refidxL1 can be derived as 0 respectively. Or, for example, the refidxL0 and refidxL1 can be derived as the smallest indices having the same POC difference in relation to the current picture respectively. Specifically, for example, when [POC of the current picture] - [POC of the first reference picture indicated by refidxL0] is defined as the first POC difference and [POC of the second reference picture indicated by refidxL1] is defined as the second POC difference, the value of refidxL0 that points to the first reference picture can be derived from the value of refidxL0 of the current block and the value of refidxL1 that points to the second reference picture can be derived from the value of refidxL1 of the current block only when the first POC difference and the second POC difference are the same. Also, for example, when there are multiple sets where the first POC difference and the second POC difference are the same, among them, refidxL0 and refidxL1 of the set with the smallest difference can be derived from refidxL0 and refidxL1 of the current block. When the symmetric MVD mode is applied, the refidxL0 and refidxL1 can be called RefIdxSymL0 and RefIdxSymL1 respectively.As an example, RefIdxSymL0 and RefIdxSymL1 can be determined on a slice-by-slice basis. In this case, the blocks to which SMVD in the current slice is applied can utilize the same RefIdxSymL0 and RefIdxSymL1.
[0130] MVDL1 can be derived from -MVDL0. For example, the final MV for the current block can be derived as shown in Equation 1.
[0131] [Number]
[0132] In Equation 1, mvx 0 and mvy 0 can represent the x and y components of the motion vector for L0 motion information or L0 prediction, and mvx 1 and mvy 1 can represent the x and y components of the motion vector for L1 motion information or L1 prediction. Also, mvpx 0 and mvpy 0 can represent the x and y components of the motion vector predictor for L0 prediction, and mvpx 1 and mvpy 1 can represent the x and y components of the motion vector predictor for L1 prediction. Also, mvdx 0 and mvdy 0 can represent the x and y components of the motion vector difference for L0 prediction.
[0133] On the one hand, a predicted block for the current block can be derived based on the motion information derived by the prediction mode. The predicted block can include the predicted sample (predicted sample array) of the current block. When the motion vector of the current block indicates a fractional sample unit, an interpolation procedure can be performed, through which the predicted sample of the current block can be derived based on the reference samples in the reference picture in fractional sample units. When Affine inter prediction is applied to the current block, predicted samples can be generated based on the sample / sub-block unit MV. When dual prediction is applied, the predicted sample used as the predicted sample of the current block can be the weighted sum or weighted average (according to the phase) of the predicted sample derived based on the L0 prediction (i.e., prediction using the reference picture in the reference picture list L0 and MVL0) and the predicted sample derived based on the L1 prediction (i.e., prediction using the reference picture in the reference picture list L1 and MVL1). When dual prediction is applied, if the reference picture used for the L0 prediction and the reference picture used for the L1 prediction are located in different temporal directions with respect to the current picture (i.e., when it is dual prediction and corresponds to bidirectional prediction), this can be called true dual prediction.
[0134] As described above, a restored sample and a restored picture can be generated based on the derived predicted sample, and subsequent procedures such as in-loop filtering can be performed.
[0135] On the other hand, regarding the configuration related to the symmetric MVD (SMVD) again, the symmetric MVD flag is signaled only for the blocks to which dual prediction is applied. When the symmetric MVD flag is true, only the MVD for L0 can be signaled, and the MVD for L1 can be used by mirroring the MVD signaled for L0. However, in this case, problems can occur depending on the value of the mvd_l1_zero_flag syntax element (for example, when the mvd_l1_zero_flag syntax element is true) in the process of applying the symmetric MVD.
[0136] For example, the mvd_l1_zero_flag syntax element can be signaled based on the syntax such as Tables 2 to 4. That is, when the current slice is a B slice in the slice header, the mvd_l1_zero_flag syntax element can be signaled. Here, Tables 2 to 4 can represent a series of syntaxes (for example, slice header syntax).
[0137] [Table 2]
[0138] [Table 3]
[0139] [Table 4]
[0140] For example, the semantics of the mvd_l1_zero_flag syntax element in Tables 2 to 4 can be as shown in Table 5 below.
[0141] [Table 5]
[0142] Alternatively, for example, the mvd_l1_zero_flag syntax element can represent information regarding whether the mvd_coding syntax for L1 prediction is parsed. For example, when the value of the mvd_l1_zero_flag syntax element is 1, it can represent that the mvd_coding syntax by L1 prediction is not parsed and the MvdL1 value is determined to be 0. Alternatively, for example, when the value of the mvd_l1_zero_flag syntax element is 0, it can represent that the mvd_coding syntax by L1 prediction is parsed. That is, the MvdL1 value can be determined by the mvd_l1_zero_flag syntax element.
[0143] For example, the decoding procedure of symmetric motion vector difference reference indices can be as shown in Table 6 below, but is not limited thereto. For example, the symmetric MVD reference index for L0 prediction can be represented as RefIdxSymL0, and the symmetric MVD reference index for L1 prediction can be represented as RefIdxSymL1.
[0144] [Table 6]
[0145] For example, information regarding the symmetric MVD or the sym_mvd_flag syntax element can be signaled based on a syntax such as Tables 7 to 10 below. Here, Tables 7 to 10 can represent one syntax (for example, coding unit syntax) continuously.
[0146] [Table 7]
[0147]
Table 8
[0148]
Table 9
[0149]
Table 10
[0150] Alternatively, for example, information for a symmetric MVD or the sym_mvd_flag syntax element can be signaled based on the syntax as shown in Tables 11 to 15 below. Here, Tables 11 to 15 can represent one syntax (e.g., coding unit syntax) continuously.
[0151]
Table 11
[0152]
Table 12
[0153]
Table 13
[0154]
Table 14
[0155]
Table 15
[0156] For example, information regarding symmetric MVD or the sym_mvd_flag syntax element can be signaled as shown in Tables 7 to 10 or Tables 11 to 15.
[0157] For example, as shown in Tables 7 to 10 or Tables 11 to 15, for information regarding symmetric MVD or the sym_mvd_flag syntax element, in the case of a block where dual prediction is applied and the current block is not an affine block regardless of the mvd_l1_zero_flag syntax element, if there exists a symmetric mvd reference index derived by the method, it can be signaled.
[0158] That is, the sym_mvd_flag syntax element can be 1 even when the mvd_l1_zero_flag syntax element is 1. In this case, even if the sym_mvd_flag syntax element is 1, MvdL1 can be derived as 0. This causes a problem of unnecessary bit signaling because the sym_mvd_flag syntax element is signaled even though it does not operate based on the sym_mvd_flag syntax element.
[0159] On the other hand, according to an embodiment of this document, when the value of the mvd_l1_zero_flag syntax element is 1 and the value of the sym_mvd_flag syntax element is 1, even if the mvd_l1_zero_flag syntax element is 1, the value of MvdL1 can be derived as a value mirrored from the value of MvdL0.
[0160] For example, information regarding symmetric MVD or the sym_mvd_flag syntax element can be signaled based on at least a part of the coding unit syntax as shown in Tables 16 and 17. Here, Tables 16 and 17 may represent a continuous part of a series of syntaxes.
[0161]
Table 16
[0162]
Table 17
[0163] For example, as shown in Table 16 and Table 17, the value of MvdL1 can be derived based on the value of the mvd_l1_zero_flag syntax element, the value of the sym_mvd_flag syntax element, and whether dual prediction is applied. That is, based on the value of the mvd_l1_zero_flag syntax element, the value of the sym_mvd_flag syntax element, and whether dual prediction is applied, the value of MvdL1 can be derived as 0 or the value of -MvdL0 (the mirrored value of MvdL0).
[0164] For example, in Table 16 and Table 17, the semantics of the mvd_l1_zero_flag syntax element can be as shown in Table 18.
[0165]
Table 18
[0166] On the one hand, according to an embodiment of this document, when the value of the mvd_l1_zero_flag syntax element is 1, the sym_mvd_flag syntax element may not be parsed. That is, when the value of the mvd_l1_zero_flag syntax element is 1, symmetric MVD is not allowed. In other words, when the value of the mvd_l1_zero_flag syntax element is 1, the decoding device does not have to parse the sym_mvd_flag syntax element, and the encoding device can be configured such that the sym_mvd_flag syntax element is not parsed (from the bitstream, video / video information, CU syntax, inter prediction mode information, or prediction-related information). Or, according to an embodiment, when the value of the mvd_l1_zero_flag syntax element is 0, a symmetric MVD index can be derived or induced. Here, the symmetric MVD index can represent a symmetric MVD reference (picture) index. For example, when a symmetric MVD is applied to the current block (ex. sym_mvd_flag == 1), there are several cases where the L0 / L1 reference (picture) index information (ex. ref_idx_l0 and / or ref_idx_l1) for the current block cannot be explicitly signaled, and the symmetric MVD reference index can be derived or induced.
[0167] For example, in this case, the decoding procedure for the symmetric motion vector difference reference indices is as follows in Table 19 below, but is not limited thereto. For example, the symmetric MVD reference index for L0 prediction can be represented as RefIdxSymL0, and the symmetric MVD reference index for L1 prediction can be represented as RefIdxSymL1.
[0168]
Table 19
[0169] Alternatively, for example, information regarding symmetric MVD or the sym_mvd_flag syntax element can be signaled based on at least a part of the coding unit syntax such as in Tables 20 and 21. Here, Tables 20 and 21 may represent a part of one syntax continuously.
[0170]
Table 20
[0171]
Table 21
[0172] For example, as shown in Tables 20 and 21, the SMVD flag information (e.g., the sym_mvd_flag syntax element) can be signaled based on at least one of the inter-prediction type information (e.g., the inter_pred_idc syntax element) indicating whether dual prediction is applied to the current block, the L1 motion vector difference zero flag information (e.g., the mvd_l1_zero_flag or ph_mvd_l1_zero_flag syntax element), the SMVD available flag information (e.g., the sps_smvd_enabled_flag syntax element), and / or the inter-affine flag information (e.g., the inter_affine_flag syntax element).
[0173] Specifically, for example, when the inter-prediction type information represents the dual prediction type, the value of the L1 motion vector difference zero flag information is 0, the value of the SMVD available flag information is 1, and the value of the inter-affine flag information is 0, the SMVD flag information can be explicitly signaled. That is, when the value of the L1 motion vector difference zero flag information is 1, the SMVD flag information can be not explicitly signaled. In this case, the decoding device can implicitly infer that the value of the SMVD flag information is 0. That the SMVD flag information is not signaled can mean that the SMVD flag information is not included in or not parsed from the prediction-related information (or inter-prediction mode information) or the CU syntax.
[0174] Based on the embodiments of the present document described above, the signaling of information related to Symmetric MVD can be optimized to avoid unnecessary bit waste. Through this, the overall coding efficiency can be improved.
[0175] FIG. 10 and FIG. 11 schematically show an example of a video / video encoding method and related components according to the embodiments of the present document.
[0176] The method disclosed in FIG. 10 can be performed by the encoding device disclosed in FIG. 2 or FIG. 11. Specifically, for example, S1000 to S1010 in FIG. 10 can be performed by the prediction unit 220 of the encoding device, and 1020 in FIG. 10 can be performed by the entropy encoding unit 240 of the encoding device. The method disclosed in FIG. 10 can include the embodiments described above in the present document.
[0177] As shown in FIG. 10, the encoding device derives an MVP candidate list for the current block (S1000). The encoding device can derive the MVP candidate list based on the inter-prediction mode of the current block and the adjacent blocks of the current block. The adjacent blocks can include the spatial adjacent blocks and / or the temporal adjacent blocks of the current block as described above. The adjacent blocks can include the lower left corner adjacent block, the left adjacent block, the upper right corner adjacent block, the upper adjacent block, and the upper left corner adjacent block of the current block.
[0178] The encoding device generates information for representing the motion information of the current block based on the MVP candidate list (S1010). The information for representing the motion information of the current block can include the selection information (such as ex.mvp flag information or mvp index information, etc.) described above. The mvp flag information can include the mvp_l0_flag and / or the mvp_l1_flag described above. Also, the information for representing the motion information of the current block can include information related to MVD. For example, the information related to MVD can include at least one of the abs_mvd_greater0_flag syntax element, the abs_mvd_greater1_flag syntax element, the abs_mvd_minus2 syntax element, or the mvd_sign_flag syntax element, but it is not limited thereto since other information can also be further included. The motion information can include at least one of the L0 motion vector for L0 prediction or the L1 motion vector for L1 prediction. The L0 motion vector is represented based on the L0 motion vector predictor and the L0 motion vector difference, and the L1 motion vector can also be represented based on the L1 motion vector predictor and the L1 motion vector difference.
[0179] For example, the encoding device can perform inter prediction on the current block by considering the RD (rate distortion) cost to generate the predicted sample of the current block. Or, for example, the encoding device can determine the inter prediction mode used to generate the predicted sample of the current block and derive motion information. Here, the inter prediction mode can be the MVP (motion vector prediction) mode, but is not limited thereto. Here, the MVP mode can also be called the AMVP (advanced motion vector prediction) mode.
[0180] The encoding device can derive the optimal motion information for the current block through motion estimation. For example, the encoding device can use the original block in the original picture for the current block to search for similar reference blocks with high correlation within a determined search range in the reference picture in units of fractional pixels, and derive motion information through this.
[0181] The encoding device can construct a motion vector predictor candidate list to represent the derived motion information using the motion vector predictor and / or the motion vector difference. For example, the encoding device can construct a motion vector predictor candidate list based on the spatial neighboring candidate blocks and / or the temporal neighboring candidate blocks. For example, when dual prediction is applied to the current block, an L0 motion vector predictor candidate list for L0 prediction and an L1 motion vector predictor candidate list for L1 prediction can be respectively constructed.
[0182] The encoding device can determine a motion vector predictor for a current block based on a motion vector predictor candidate list. For example, the encoding device can determine a motion vector predictor for the current block based on the derived motion information (or motion vector) among the motion vector predictor candidates in the motion vector predictor candidate list. Or, the encoding device can determine, within the motion vector predictor candidate list, a motion vector predictor with the least difference from the derived motion information (or motion vector). For example, when dual prediction is applied to the current block, an L0 motion vector predictor for L0 prediction and an L1 motion vector predictor for L1 prediction can be determined from an L0 motion vector predictor candidate list and an L1 motion vector predictor candidate list, respectively.
[0183] The encoding device can also generate selection information representing the motion vector predictor from among the motion vector predictor candidate list. For example, the selection information can also be called index information and can also be called an MVP flag or an MVP index. That is, the encoding device can generate information indicating the motion vector predictor used to represent the motion vector of the current block within the motion vector predictor candidate list. For example, when dual prediction is applied to the current block, selection information for the L0 motion vector predictor and selection information for the L1 motion vector predictor can be generated, respectively.
[0184] The encoding device can determine the motion vector difference for the current block based on a motion vector predictor. For example, the encoding device can determine the motion vector difference based on the derived motion information (or motion vector) for the current block and the motion vector predictor. Alternatively, the encoding device can determine the motion vector difference based on the difference between the derived motion information (or motion vector) for the current block and the motion vector predictor. For example, when dual prediction is applied to the current block, an L0 motion vector difference and an L1 motion vector difference can each be determined. Here, the L0 motion vector difference can be represented as MvdL0, and the L1 motion vector difference can also be represented as MvdL1. The L0 motion vector can be represented based on the sum of the L0 motion vector predictor and the L0 motion vector difference, and the L1 motion vector can be represented based on the sum of the L1 motion vector predictor and the L1 motion vector difference.
[0185] In addition, the encoding device can derive the motion information of the current block based on the inter prediction mode and generate a prediction sample. The motion information can include an L0 motion vector for L0 prediction and / or an L1 motion vector for L1 prediction.
[0186] When dual prediction is applied, the encoding device can derive an L0 prediction sample and an L1 prediction sample, and derive the prediction sample of the current block based on the weighted sum or weighted average of the L0 prediction sample and the L1 prediction sample. The L0 motion vector can represent the L0 prediction sample on the L0 reference picture, and the L1 motion vector can represent the L1 prediction sample on the L1 reference picture.
[0187] Based on the case where the value of the L1 motion vector difference zero flag information is 0 and the value of the SMVD flag information is 1, the L1 motion vector difference can be derived from the L0 motion vector difference. In this case, the absolute value of the L1 motion vector difference is the same as the absolute value of the L0 motion vector difference, and the sign of the L1 motion vector difference may be different from the sign of the L0 motion vector difference.
[0188] Based on the predicted sample of the current block, the encoding device can derive residual information. The encoding device can derive a residual sample based on the predicted sample. The encoding device can derive a residual sample based on the original sample for the current block and the predicted sample for the current block. The encoding device can derive residual information based on the residual sample. The residual information can include information for the quantized transform coefficients. The encoding device can perform a transform / quantization procedure on the residual sample to derive the quantized transform coefficients.
[0189] The encoding device encodes video / video information (S1020). The video / video information can include at least one of the information related to the inter prediction mode or the information for representing the motion information of the current block. The video / video information can further include the residual information. The encoded video / video information can be output in the form of a bitstream. The bitstream can be transmitted to the decoding device via a network or a storage medium. The bitstream can include the encoded (video / video) information.
[0190] The video / video information can include inter-prediction type information (e.g., inter_pred_idc), general merge flag, SMVD flag information, L1 motion vector difference zero flag information, SMVD availability flag information, and / or inter-affine flag information. For example, when the value of the general merge flag is 0, it can indicate that the MVP mode is applied to the current block. The inter-prediction type information can indicate whether dual prediction is applied to the current block. Specifically, for example, the inter-prediction type information can indicate whether L0 prediction, L1 prediction, or dual prediction is applied to the current block.
[0191] As an example, the video information includes L1 motion vector difference zero flag information, the video information includes the coding unit (CU) syntax for the current block, and based on the L1 motion vector difference zero flag information, it can be determined whether the CU syntax includes SMVD (symmetric motion vector differences) flag information indicating whether SMVD is applied to the current block.
[0192] Specifically, for example, the video information can include header information. The header information can include a slice header or a picture header. The header information includes the L1 motion vector difference zero flag information, and based on the case where the value of the L1 motion vector difference zero flag information is 0, the CU syntax can include the SMVD flag information. For example, when the value of the mvd_l1_zero_flag syntax element is 1, the decoding device does not need to parse the sym_mvd_flag syntax element, and the encoding device can be configured so that the sym_mvd_flag syntax element is not parsed (from the bitstream, video / video information, CU syntax, inter-prediction mode information, or prediction-related information).
[0193] Also, as an example, based on the case where the value of the L1 motion vector difference zero flag information is 1, the CU syntax may not include the SMVD flag information. In this case, the decoding device can infer that the value of the SMVD flag information is 0 without parsing the SMVD flag information.
[0194] Also, as an example, based on the value of the symmetric motion vector difference reference index for the current block, it is determined whether the SMVD flag information is included in the CU syntax. Based on the L1 motion vector difference zero flag information whose value is 1, the value of the symmetric motion vector difference reference index is derived as -1. Based on the symmetric motion vector difference reference index whose value is -1, it can be determined that the SMVD flag information is not included in the CU syntax. For example, the symmetric motion vector difference reference index can include the L0 symmetric motion vector difference reference index and the L1 symmetric motion vector difference reference index, and each can be determined to have a value of -1. Here, the L0 symmetric motion vector difference reference index can be represented by RefIdxSymL0, and the L1 symmetric motion vector difference reference index can be represented by RefIdxSymL1. For example, the SMVD flag information can be included in the inter prediction mode information or the CU syntax when RefIdxSymL0 and RefIdxSymL1 are each greater than -1. Or, for example, based on RefIdxSymL0 greater than -1 and RefIdxSymL1 greater than -1, the inter prediction mode information or the CU syntax can include the symmetric motion vector difference flag.
[0195] Also, as an example, the video information includes a sequence parameter set (SPS), the SPS includes SMVD available flag information, the CU syntax includes prediction type information and inter-affine flag information, the prediction type information indicates whether dual prediction is applied to the current block, and based on the SMVD available flag information, the prediction type information, and the inter-affine flag information, the CU syntax can further include the SMVD flag information.
[0196] Specifically, for example, when the inter-prediction type information represents the dual prediction type, the value of the L1 motion vector difference zero flag information is 0, the value of the SMVD available flag information is 1, and the value of the inter-affine flag information is 0, the SMVD flag information can be included in the CU syntax and explicitly signaled. That is, when the value of the L1 motion vector difference zero flag information is 1, the SMVD flag information can be not explicitly signaled. In this case, the decoding device can implicitly infer that the value of the SMVD flag information is 0. Not signaling the SMVD flag information can mean that the SMVD flag information is not included in the prediction-related information (or inter-prediction mode information) or the CU syntax.
[0197] The video / video information can include information related to the motion vector difference. For example, the information related to the motion vector difference can include information representing the motion vector difference. When SMVD is applied, the information related to the motion vector difference can include only information related to the L0 motion vector difference, and the L1 motion vector difference can be derived based on the L0 motion vector difference as described above.
[0198] The zero flag for the L1 motion vector difference can represent information on whether the L1 motion vector difference is 0, and can be called the L1 motion vector difference zero flag, L1 MVD zero flag, or MVD L1 zero flag. Also, the L1 motion vector difference zero flag can be represented by the mvd_l1_zero_flag syntax element or the ph_mvd_l1_zero_flag syntax element. Also, the SMVD flag information can represent information on whether the L0 motion vector difference and the L1 motion vector difference are symmetric, and can also be represented by the sym_mvd_flag syntax element.
[0199] Also, the encoding device can derive (modified) residual samples based on the residual information, and can also generate restored samples based on the (modified) residual samples and the predicted samples. Also, a restored block and a restored picture can be derived based on the restored samples. The derived restored picture can be referenced for inter prediction of subsequent pictures.
[0200] For example, the encoding device can encode video / video information including all or part of the above-described information (or syntax elements) to generate a bitstream or encoded information. Or it can be output in the form of a bitstream. Also, the bitstream or encoded information can be transmitted to a decoding device via a network or a storage medium. Or the bitstream or encoded information can be stored in a computer-readable storage medium, and the bitstream or the encoded information can be generated by the above-described video / video encoding method.
[0201] FIG. 12 and FIG. 13 schematically show an example of a video / video decoding method and related components according to the embodiment(s) of this document.
[0202] The method disclosed in FIG. 12 can be performed by the decoding apparatus disclosed in FIG. 3 or FIG. 13. Specifically, for example, S1200 in FIG. 12 can be performed by the entropy decoding unit 310 of the decoding apparatus, S1210 to S1040 in FIG. 12 can be performed by the prediction unit 330 of the decoding apparatus, and S1050 in FIG. 12 can be performed by the addition unit 340 or the restoration unit of the decoding apparatus. The method disclosed in FIG. 12 can include the embodiments described above in this document.
[0203] As shown in FIG. 12, the decoding apparatus acquires video / video information from a bitstream (S1200). The video / video information can include at least one of information regarding the inter prediction mode or information for representing motion information of a current block. The video / video information can further include the residual information. For example, the decoding apparatus can acquire the video / video information by parsing or decoding the bitstream. For example, the prediction-related information can include information representing a prediction mode and / or motion information used to generate a prediction sample of the current block. The prediction-related information can include information regarding the selection information or motion vector difference.
[0204] The information for representing the motion information of the current block can include the selection information described above (such as mvp flag information or mvp index information, etc.). The mvp flag information can include the mvp_l0_flag and / or mvp_l1_flag described above. Also, the information for representing the motion information of the current block can include information regarding MVD.
[0205] For example, the selection information can also be referred to as index information and can also be called an MVP flag or an MVP index. That is, the selection information can represent information indicating a motion vector predictor of a current block within a motion vector predictor candidate list. For example, when dual prediction is applied to the current block, the selection information can include selection information for L0 prediction and selection information for L1 prediction. The selection information for L0 prediction can represent information indicating a motion vector predictor within an L0 motion vector predictor candidate list for L0 prediction described later, and the selection information for L1 prediction can represent information indicating a motion vector predictor within an L1 motion vector predictor candidate list for L1 prediction described later.
[0206] For example, the information regarding the motion vector difference (MVD) can include information used to derive the motion vector difference. For example, the information regarding the motion vector difference can include at least one of an abs_mvd_greater0_flag syntax element, an abs_mvd_greater1_flag syntax element, an abs_mvd_minus2 syntax element, or an mvd_sign_flag syntax element, but is not limited thereto since other information can also be further included. For example, when dual prediction is applied to the current block, the information regarding the motion vector difference can include information regarding the motion vector difference for L0 prediction and information regarding the motion vector difference for L1 prediction.
[0207] The video / video information can include inter-prediction type information (e.g., inter_pred_idc), general merge flag, SMVD flag information, L1 motion vector difference zero flag information, SMVD availability flag information, and / or inter-affine flag information as information related to the inter-prediction mode. For example, when the value of the general merge flag is 0, it can indicate that the MVP mode is applied to the current block. The inter-prediction type information can indicate whether dual prediction is applied to the current block. Specifically, for example, the inter-prediction type information can indicate whether L0 prediction, L1 prediction, or dual prediction is applied to the current block. The prediction-related information can include information related to the motion vector difference.
[0208] As an example, the video information includes L1 motion vector difference zero flag information, the video information includes the coding unit (CU) syntax for the current block, and based on the L1 motion vector difference zero flag information, it can be determined whether the CU syntax includes SMVD (symmetric motion vector differences) flag information indicating whether SMVD is applied to the current block.
[0209] Specifically, for example, the video information can include header information. The header information can include a slice header or a picture header. The header information includes the L1 motion vector difference zero flag information, and based on the case where the value of the L1 motion vector difference zero flag information is 0, the CU syntax can include the SMVD flag information. For example, when the value of the mvd_l1_zero_flag syntax element is 1, the decoding device does not need to parse the sym_mvd_flag syntax element, and the encoding device can be configured such that the sym_mvd_flag syntax element is not parsed (from the bitstream, video / video information, CU syntax, inter prediction mode information, or prediction-related information).
[0210] Also, as an example, based on the case where the value of the L1 motion vector difference zero flag information is 1, the CU syntax may not include the SMVD flag information. In this case, the decoding device can infer that the value of the SMVD flag information is 0 without parsing the SMVD flag information.
[0211] Also, as an example, based on the value of the symmetric motion vector difference reference index for the current block, it is determined whether the SMVD flag information is included in the CU syntax. Based on the L1 motion vector difference zero flag information whose value is 1, the value of the symmetric motion vector difference reference index is derived as -1. Based on the symmetric motion vector difference reference index whose value is -1, it can be determined that the SMVD flag information is not included in the CU syntax. For example, the symmetric motion vector difference reference index can include an L0 symmetric motion vector difference reference index and an L1 symmetric motion vector difference reference index, and each can be determined to have a value of -1. Here, the L0 symmetric motion vector difference reference index can be represented by RefIdxSymL0, and the L1 symmetric motion vector difference reference index can be represented by RefIdxSymL1. For example, the SMVD flag information can be included in the inter prediction mode information or the CU syntax when RefIdxSymL0 and RefIdxSymL1 are each greater than -1. Or, for example, based on RefIdxSymL0 greater than -1 and RefIdxSymL1 greater than -1, the inter prediction mode information or the CU syntax can include the symmetric motion vector difference flag.
[0212] Also, as an example, the video information includes a sequence parameter set (SPS), the SPS includes SMVD availability flag information, the CU syntax includes prediction type information and inter-affine flag information, the prediction type information indicates whether dual prediction is applied to the current block, and based on the SMVD availability flag information, the prediction type information, and the inter-affine flag information, the CU syntax can include the SMVD flag information.
[0213] Specifically, for example, when the inter-prediction type information represents a dual-prediction type, the value of the L1 motion vector difference zero flag information is 0, the value of the SMVD available flag information is 1, and the value of the inter-affine flag information is 0, the SMVD flag information can be included in the CU syntax and explicitly signaled. That is, when the value of the L1 motion vector difference zero flag information is 1, the SMVD flag information can be not explicitly signaled. In this case, the decoding device can implicitly infer that the value of the SMVD flag information is 0. The fact that the SMVD flag information is not signaled can mean that the SMVD flag information is not included in the prediction-related information (or inter-prediction mode information) or the CU syntax.
[0214] The video / video information can include information regarding the motion vector difference. For example, the information regarding the motion vector difference can include information representing the motion vector difference. When the SMVD is applied, the information regarding the motion vector difference can include only the information regarding the L0 motion vector difference, and the L1 motion vector difference can be derived based on the L0 motion vector difference as described above.
[0215] The zero flag for the L1 motion vector difference can represent information as to whether the L1 motion vector difference is 0, and can be called the L1 motion vector difference zero flag, the L1 MVD zero flag, or the MVD L1 zero flag. Also, the L1 motion vector difference zero flag can also be represented by the mvd_l1_zero_flag syntax element or the ph_mvd_l1_zero_flag syntax element. Also, the SMVD flag information can represent information as to whether the L0 motion vector difference and the L1 motion vector difference are symmetric, and can also be represented by the sym_mvd_flag syntax element.
[0216] Alternatively, for example, based on the L1 motion vector difference zero flag having a value of 0 or the symmetric motion vector difference flag having a value of 1, the L1 motion vector difference can also be represented from the L0 motion vector difference. That is, when the value of the L1 motion vector difference zero flag is 0 or the value of the symmetric motion vector difference flag is 1, the L1 motion vector difference can also be represented from the L0 motion vector difference. Alternatively, the L1 motion vector difference can also be represented by a mirrored value of the L0 motion vector difference. For example, when the value of the mvd_l1_zero_flag syntax element or the ph_mvd_l1_zero_flag syntax element is 0, or the value of the sym_mvd_flag syntax element is 1, MvdL1 can also be represented based on MvdL0. Alternatively, MvdL1 can also be determined to be -MvdL0. For example, MvdL1 being -MvdL0 can represent that the absolute value, i.e., the magnitude, of the L1 motion vector difference is the same as the absolute value, i.e., the magnitude, of the L0 motion vector difference, and the sign of the L1 motion vector difference is different from the sign of the L0 motion vector difference.
[0217] The decoding device derives an inter prediction mode of the current block based on the video information (S1210). For example, the decoding device can derive the inter prediction mode based on information regarding the inter prediction mode included in the video information. For example, the inter prediction mode can be, but is not limited to, the MVP (motion vector prediction) mode. Here, the MVP mode can also be called the AMVP (advanced motion vector prediction) mode.
[0218] The decoding device derives an MVP candidate list based on the inter prediction mode and the adjacent blocks of the current block (S1220). For example, the decoding device can configure an MVP candidate list based on spatially adjacent candidate blocks and / or temporally adjacent candidate blocks. Here, the configured MVP candidate list may be the same as the MVP candidate list configured by the encoding device. The adjacent blocks may include the lower left corner adjacent block, the left adjacent block, the upper right corner adjacent block, the upper adjacent block, and the upper left corner adjacent block of the current block. For example, when dual prediction is applied to the current block, an L0 MVP candidate list for L0 prediction and an L1 MVP candidate list for L1 prediction can each be configured.
[0219] The decoding device derives the motion information of the current block based on the MVP candidate list (S1230). For example, the decoding device can derive a motion vector predictor candidate for the current block within the motion vector predictor candidate list based on the selection information described above, and can derive the motion information (or motion vector) of the current block based on the derived motion vector predictor candidate. Alternatively, the motion information (or motion vector) of the current block can be derived based on the derived motion vector predictor candidate and the motion vector difference derived based on the information regarding the motion vector difference described above. For example, when dual prediction is applied to the current block, the L0 motion vector predictor for L0 prediction and the L1 motion vector predictor for L1 prediction can be derived from the L0 motion vector predictor candidate list and the L1 motion vector predictor candidate list respectively based on the selection information for L0 prediction and the selection information for L1 prediction. For example, when dual prediction is applied to the current block, the L0 motion vector difference and the L1 motion vector difference can be derived respectively based on the information regarding the motion vector difference. Here, the L0 motion vector difference can be represented by MvdL0, and the L1 motion vector difference can also be represented by MvdL1. Also, the motion vector of the current block can be derived by the L0 motion vector and the L1 motion vector respectively.
[0220] The decoding device generates a predicted sample of the current block based on the motion information (S1240). For example, when dual prediction is applied to the current block, the decoding device can generate an L0 predicted sample for L0 prediction based on the L0 motion vector and an L1 predicted sample for L1 prediction based on the L1 motion vector. Also, the decoding device can generate the predicted sample of the current block based on the L0 predicted sample and the L1 predicted sample.
[0221] The decoding device generates a restored sample based on the prediction sample (S1250). A restored block or a restored picture can be derived based on the restored sample. As described above, an in-loop filtering procedure can be further applied to the restored sample / block / picture.
[0222] The decoding device can generate the restored sample based on the prediction sample and the residual sample. In this case, the decoding device can derive the residual sample based on the residual information. For example, the residual information can represent information used to derive the residual sample, and can include information related to the residual sample, inverse transformation related information, and / or inverse quantization related information. For example, the residual information can include information regarding the quantized transform coefficients.
[0223] For example, the decoding device can decode a bitstream or encoded information to obtain video / video information including all or part of the above-described information (or syntax elements). Also, the bitstream or the encoded information can be stored in a computer-readable storage medium and can cause the above-described decoding method to be performed.
[0224] In the foregoing embodiments, the method is described based on a flowchart in a series of steps or blocks, but the corresponding embodiments are not limited to the order of the steps, and a certain step can occur in a different order or simultaneously with steps different from the foregoing. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, other steps are included, and one or more steps of the flowchart can be deleted without affecting the scope of the embodiments of this document.
[0225] The method according to the embodiments of the foregoing document can be embodied in the form of software, and the encoding device and / or decoding device according to the document can be included in a device that executes video processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0226] In this document, when an embodiment is embodied in software, the foregoing method can be embodied by modules (processes, functions, etc.) that perform the foregoing functions. The modules can be stored in a memory and executed by a processor. The memory can be inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be embodied and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be embodied and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for embodiment (for example, information on instructions) or an algorithm can be stored in a digital storage medium.
[0227] In addition, the decoding device and encoding device to which the example(s) of this document is / are applied can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (augmented reality) device, a videophone video device, a transportation means terminal (e.g., a vehicle (including an autonomous driving vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, etc., and can be used to process a video signal or a data signal. For example, as an OTT video (Over the top video) device, it can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.
[0228] In addition, the processing method to which the example(s) of this document is / are applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Also, multimedia data having a data structure according to the example(s) of this document can be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which data that can be read by a computer is stored. The computer-readable recording medium can include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Also, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (e.g., transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or can be transmitted via a wired or wireless communication network.
[0229] In addition, the example(s) of this document can be embodied in a computer program product by program code, and the program code can be executed by a computer according to the example(s) of this document. The program code can be stored on a carrier readable by a computer.
[0230] FIG. 14 shows an example of a content streaming system to which the example disclosed in this document can be applied.
[0231] As shown in FIG. 14, the content streaming system to which the example of this document is applied can include a large encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.
[0232] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream, and plays the role of transmitting this to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server can be omitted.
[0233] The bitstream can be generated by an encoding method or a bitstream generation method applied to the embodiments of this document, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0234] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server plays the role of a medium to inform the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include a separate control server. In this case, the control server plays the role of controlling commands / responses between each device within the content streaming system.
[0235] The streaming server can receive content from a media repository and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0236] Examples of the user device include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (for example, a smartwatch, a smart glass, an HMD (head mounted display), a digital TV, a desktop computer, and a digital signage.
[0237] Each server in the content streaming system can be operated as a distributed server. In this case, the data received by each server can be distributedly processed.
[0238] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and implemented in a device, and the technical features of the device claims in this specification can be combined and implemented in a method. Also, the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented in a device, and the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented in a method.
Claims
1. A video decoding method performed by a decoding device, comprising: obtaining video information from a bitstream; deriving a motion vector predictor (MVP) candidate list for the current block based on an inter prediction mode and neighboring blocks of the current block; deriving motion information of the current block based on the MVP candidate list; generating a prediction sample for the current block based on the motion information; generating reconstructed samples based on the predicted samples; The motion information includes at least one of an L0 motion vector for L0 prediction or an L1 motion vector for L1 prediction; the L0 motion vector is derived based on an L0 motion vector predictor and an L0 motion vector differential, and the L1 motion vector is derived based on an L1 motion vector predictor and an L1 motion vector differential; The image information includes an L1 motion vector difference zero flag and a symmetric motion vector differences (SMVD) flag for the current block, based on a value of the SMVD flag being equal to one, the L1 motion vector differential is derived from the L0 motion vector differential; Based on the value of the L1 motion vector differential zero flag being equal to 1, the L1 motion vector differential is derived as 0; The L1 motion vector differential is derived as 0 only if the value of the SMVD flag is not equal to 1.
2. The video information includes header information, The video decoding method of claim 1 , wherein the header information includes the L1 motion vector difference zero flag.
3. The video information includes a sequence parameter set (SPS) and a coding unit (CU) syntax for the current block; The SPS includes an SMVD availability flag; The CU syntax includes inter prediction type information and an inter affine flag, the inter prediction type information indicating whether bi-prediction is applied to the current block, The video decoding method of claim 1 , further based on the SMVD available flag, the inter prediction type information, and the inter affine flag, the CU syntax includes the SMVD flag.
4. the absolute value of the L1 motion vector differential is equal to the absolute value of the L0 motion vector differential; The method of claim 1 , wherein a sign of the L1 motion vector differential is different from a sign of the L0 motion vector differential.
5. A video encoding method performed by an encoding device, comprising: deriving a motion vector predictor (MVP) candidate list for the current block based on an inter prediction mode and neighboring blocks of the current block; generating information for representing motion information of the current block based on the MVP candidate list; encoding image information including information for representing motion information of the current block; The motion information includes at least one of an L0 motion vector for L0 prediction or an L1 motion vector for L1 prediction; the L0 motion vector is derived based on an L0 motion vector predictor and an L0 motion vector differential, and the L1 motion vector is derived based on an L1 motion vector predictor and an L1 motion vector differential; The image information includes an L1 motion vector difference zero flag and a symmetric motion vector differences (SMVD) flag for the current block, based on a value of the SMVD flag being equal to one, the L1 motion vector differential is derived from the L0 motion vector differential; Based on the value of the L1 motion vector differential zero flag being equal to 1, the L1 motion vector differential is derived as 0; The L1 motion vector differential is derived as 0 only if the value of the SMVD flag is not equal to 1.
6. The video information includes header information, The video encoding method of claim 5 , wherein the header information includes the L1 motion vector difference zero flag.
7. The video information includes a sequence parameter set (SPS) and a coding unit (CU) syntax for the current block; The SPS includes an SMVD availability flag; The CU syntax includes inter prediction type information and an inter affine flag, the inter prediction type information indicating whether bi-prediction is applied to the current block, The video encoding method of claim 5 , further based on the SMVD available flag, the inter prediction type information, and the inter affine flag, the CU syntax includes the SMVD flag.
8. the absolute value of the L1 motion vector differential is equal to the absolute value of the L0 motion vector differential; The method of claim 5 , wherein a sign of the L1 motion vector differential is different from a sign of the L0 motion vector differential.
9. 1. A method for transmitting data for video information, comprising: obtaining a bitstream for the video information, The bitstream is generated by performing the steps of: deriving a motion vector predictor (MVP) candidate list for the current block based on an inter prediction mode and neighboring blocks of the current block; generating information for representing motion information of the current block based on the MVP candidate list; and encoding image information for generating the bitstream; the image information includes information for representing motion information of the current block; transmitting the data including the bitstream; The motion information includes at least one of an L0 motion vector for L0 prediction or an L1 motion vector for L1 prediction; the L0 motion vector is derived based on an L0 motion vector predictor and an L0 motion vector differential, and the L1 motion vector is derived based on an L1 motion vector predictor and an L1 motion vector differential; The image information includes an L1 motion vector difference zero flag and a symmetric motion vector differences (SMVD) flag for the current block, based on a value of the SMVD flag being equal to one, the L1 motion vector differential is derived from the L0 motion vector differential; Based on the value of the L1 motion vector differential zero flag being equal to 1, the L1 motion vector differential is derived as 0; The method of claim 1, wherein the L1 motion vector differential is derived as 0 only if the value of the SMVD flag is not equal to 1.
Citation Information
Patent Citations
Method and apparatus for motion vector prediction-based image / video coding
JP7644286B2
Symmetric motion vector difference coding
WO2020132272A1
Symmetric motion vector difference coding
WO2020221256A1