Video or video coding based on a sub-block unit time motion vector predictor candidate
By deriving sub-block unit motion vectors from reference sub-blocks in collocated reference pictures and integrating positions across block levels, the method addresses the inefficiencies in high-resolution video coding, improving compression efficiency and reducing complexity.
Patent Information
- Application Number
- JP2024162338
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-13
- Filing Date
- 2024-09-19
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2040-06-15
AI Technical Summary
The increasing demand for high-resolution and high-quality video, including immersive media, has led to higher transmission and storage costs due to increased video data, necessitating a more efficient video coding technology, particularly in sub-block units for temporal motion vector prediction.
The method involves deriving sub-block unit motion vectors based on reference sub-blocks in collocated reference pictures, using base motion vectors for unusable references, and integrating corresponding positions at both sub-coding and coding block levels for improved prediction performance.
This approach enhances video compression efficiency, reduces computational complexity, and simplifies hardware implementation by optimizing inter-prediction and coding efficiency through efficient calculation of sub-block positions.
Smart Images

Figure 0007712447000007 
Figure 0007712447000008 
Figure 0007712447000009
Abstract
Description
Technical Field
[0001] The present technology relates to video or video coding, for example, to video or video coding technology based on a temporal motion vector predictor candidate in sub-block units.
Background Art
[0002] Recently, the demand for high-resolution and high-quality video / video such as 4K or UHD (Ultra High Definition) video / video of 8K or higher has been increasing in various fields. As the video / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to the existing video / video data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband line, or storing (storing) video / video data using an existing storage medium, the transmission cost (cost) and storage cost increase.
[0003] Also, recently, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcasting of video / video having video characteristics different from those of real-world video, such as game video, has been increasing.
[0004] Therefore, in order to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality video / video having various characteristics as described above, a highly efficient video / video compression technology is required.
[0005] Also, in order to improve the video / video coding efficiency, there has been discussion on the temporal motion vector prediction technology in sub-block units. For this purpose, a method for efficiently performing the process of patching the motion vectors in sub-block units by temporal motion vector prediction in sub-block units is required.
Summary of the Invention
Problems to be Solved by the Invention
[0006] The technical problem of this document is to provide a method and apparatus for improving the efficiency of video / video coding.
[0007] Another technical problem of this document is to provide an efficient inter prediction method and apparatus.
[0008] Still another technical problem of the present invention is to provide a method and apparatus for deriving a temporal motion vector based on sub-blocks to improve prediction performance.
[0009] Still another technical problem of the present invention is to provide a method and apparatus for efficiently deriving corresponding positions of sub-blocks for deriving (inducing) a temporal motion vector based on sub-blocks.
[0010] Still another technical problem of the present invention is to provide a method and apparatus for integrating the corresponding position at the sub-coding block level and the corresponding position at the coding block level for deriving a temporal motion vector based on sub-blocks.
Means for Solving the Problems
[0011] According to one embodiment of this document, in subblock-based temporal motion vector prediction (sbTMVP), the sub-block unit motion vector for the current block can be derived based on the reference sub-blocks on the collocated reference picture.
[0012] According to one embodiment of this document, the reference sub-blocks on the collocated reference picture for a plurality of sub-blocks in the current block can be derived based on the center sample position of each sub-block in the current block.
[0013] According to one embodiment of this document, for the sub-block unit motion vector for the unusable reference sub-block in the reference sub-block, the base motion vector can be used.
[0014] According to one embodiment of this document, the base motion vector can be derived from (on) the collocated reference picture based on the center sample position of the current block.
[0015] According to one embodiment of this document, a video / video decoding method executed by a decoding device is provided. The video / video decoding method can have the method disclosed in the embodiment of this document.
[0016] According to one embodiment of this document, a decoding device for executing video / video decoding is provided. The decoding device can execute the method disclosed in the embodiment of this document.
[0017] According to one embodiment of this document, a video / video encoding method executed by an encoding device is provided. The video / video encoding method can have the method disclosed in the embodiment of this document.
[0018] According to one embodiment of this document, an encoding device for executing video / video encoding is provided. The encoding device can execute the method disclosed in the embodiment of this document.
[0019] According to one embodiment of this document, a computer-readable digital storage medium storing encoded video / video information generated according to the video / video encoding method disclosed in at least one of the embodiments of this document is provided.
[0020] According to one embodiment of this document, there is provided an encoded information that causes a decoding device to execute a video / video decoding method disclosed in at least one of the embodiments of this document, or a computer-readable digital storage medium storing the encoded video / video information.
Advantages of the Invention
[0021] This document can achieve various effects. For example, it can improve the overall video / video compression efficiency. Also, it can reduce the computational complexity through efficient inter-prediction and improve the overall coding efficiency. Further, by efficiently calculating the corresponding positions of sub-blocks for deriving sub-block-based temporal motion vectors in sub-block-based temporal motion vector prediction (sbTMVP), the efficiency in terms of complexity and prediction performance can be improved. Additionally, by integrating the method of calculating the corresponding positions at the sub-coding block level and the coding block level for deriving sub-block-based temporal motion vectors, a simplification effect can be achieved in terms of hardware implementation.
[0022] The effects obtained through the specific embodiments of this document are not limited to the effects listed above. For example, there may be various technical effects that can be understood or induced by a person having ordinary skill in the related art from this document. Thus, the specific effects of this document are not limited to those explicitly described in this document and may include various effects that can be understood or induced from the technical features of this document.
Brief Description of the Drawings
[0023]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Mode for Carrying Out the Invention
[0024] This document can be modified in various ways and can have various examples. Specific examples will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific examples. The terms commonly used in this document are merely used to explain specific examples and are not used with the intention of limiting the technical idea of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this document are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude in advance the existence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0025] On the other hand, each configuration in the drawings described in this document is independently illustrated for the convenience of explaining different characteristic functions, and does not mean that each configuration is implemented by separate hardware or separate software. For example, among the configurations, two or more configurations can be combined to form one configuration, and one configuration can be divided into a plurality of configurations. Examples in which each configuration is integrated and / or separated are included in the scope of rights of this document as long as they do not deviate from the essence of this document.
[0026] In this document, "A or B" can mean "only A", "only B", or "both A and B". Also, in this document, "A or B" can be interpreted as "A and / or B". For example, in this document, "A, B or C" can mean "only A", "only B", "only C", or "any combination of A, B and C".
[0027] The slashes ( / ) and commas used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B or C".
[0028] In this document, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, in this document, expressions such as "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".
[0029] Also, in this document, "at least one of A, B and C" can mean "only A", "only B", "only C", or "any combination of A, B and C". Also, "at least one of A, B or C" and "at least one of A, B and / or C" can mean "at least one of A, B and C".
[0030] Also, the parentheses used in this document can mean "for example". Specifically, when it is shown as "prediction (intra prediction)", "intra prediction" is proposed as an example of "prediction". As another expression, "prediction" in this document is not limited to "intra prediction", but "intra prediction" is proposed as an example of "prediction". Also, when it is shown as "prediction (i.e., intra prediction)", "intra prediction" is proposed as an example of "prediction".
[0031] This document relates to video / video coding. For example, the methods / examples disclosed in this document can be applied to the methods disclosed in the VVC (Versatile Video Coding) standard. Also, the methods / examples disclosed in this document can be applied to the methods disclosed in the EVC (Essential Video Coding) standard, the AV1 (AOMedia Video 1) standard, the AVS2 (2nd Generation Of Audio Video Coding Standard), or the next-generation video / video coding standard (e.g., H.267 or H.267).
[0032] This document presents various embodiments related to video / video coding, and unless otherwise noted, the above embodiments can also be executed in combination with each other.
[0033] In this document, video can mean a collection of a series of images over time. A picture generally means a unit representing one image at a specific time period, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile can include one or more CTUs (Coding Tree Units). One picture can be composed of one or more slices / tiles. A tile is a rectangular region of CTUs within a particular tile column and particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width that can be specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture.A tile scan can indicate a specific sequential ordering of CTUs partitioning a picture, where the CTUs can be ordered consecutively in a CTU raster scan within a tile, and tiles in a picture can be ordered consecutively in a raster scan of the tiles of the picture. A slice can include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that may be exclusively contained in a single NAL unit.
[0034] On the other hand, one picture can be divided into two or more sub-pictures. A sub-picture is an rectangular region of one or more slices within a picture.
[0035] A pixel or pel can mean the smallest unit that makes up a picture (or video). Also, the term "sample" can be used as the term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, can indicate only the pixel / pixel value of the luma component, or can indicate only the pixel / pixel value of the chroma component. Alternatively, a sample can mean the pixel value in the spatial domain (domain), and when such a pixel value is converted to the frequency domain (domain), it can also mean the conversion coefficient in the frequency domain.
[0036] A unit can indicate the basic unit of video processing. A unit can include at least one of a specific region of a picture and information related to the region. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as "block" or "area". In a general case, an M×N block can include a set (or array) of samples (or sample array) consisting of M columns and N rows, or a set (or array) of transform coefficients.
[0037] Also, in this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation can be omitted. When quantization / inverse quantization is omitted, the quantized transform coefficient can be called a transform coefficient. When transformation / inverse transformation is omitted, the transform coefficient can also be called a coefficient or a residual coefficient, or, for the sake of uniformity of expression, can still be called a transform coefficient.
[0038] In this document, the quantized transform coefficients and the transform coefficients can each be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information can include information regarding the transform coefficient(s), and the information regarding the transform coefficient(s) can be signaled via a residual coding syntax. The transform coefficients can be derived based on the residual information (or the information regarding the transform coefficient(s)), and the scaled transform coefficients can be derived via an inverse transform (scaling) with respect to the transform coefficients. Based on an inverse transform (transformation) with respect to the scaled transform coefficients, residual samples can be derived. This can be applied / expressed similarly in other parts of this document.
[0039] In this document, the technical features separately described within one drawing can be implemented separately or simultaneously.
[0040] Hereinafter, with reference to the accompanying drawings, preferred embodiments of this document will be described in more detail. Hereinafter, the same reference numerals will be used for the same components on the drawings, and redundant descriptions for the same components can be omitted.
[0041] FIG. 1 schematically shows an example of a video / image coding system that can be applied to an embodiment of this document.
[0042] Referring to FIG. 1, the video / image coding system can include a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data in a file or streaming form to the receiving device via a digital storage medium or a network.
[0043] The above source device can include a video source, an encoding device, and a transmitting unit. The above receiving device can include a receiving unit, a decoding device, and a renderer. The above encoding device can be called a video / video encoding device, and the above decoding device can be called a video / video decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can also include a display unit, and the display unit can also be composed of a separate device or an external component.
[0044] The video source can obtain video / video through processes such as video / video capture, synthesis, or generation. The video source can include a video / video capture device and / or a video / video generation device. The video / video capture device can include, for example, one or more cameras, a video / video archive containing previously captured video / video, etc. The video / video generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / video. For example, virtual video / video can be generated through a computer, etc., in which case the video / video capture process can be replaced by the process of generating related data.
[0045] The encoding device can encode the input video / video. The encoding device can execute a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / video information) can be output in the form of a bitstream.
[0046] The transmitting unit can transmit the encoded video / video information or data output in bitstream form to the receiving unit of the receiving device via a digital storage medium or a network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating media files via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the above bitstream and transmit it to the decoding device.
[0047] The decoding device can decode the video / video by executing a series of procedures such as inverse quantization, inverse transformation, prediction, etc., corresponding to the operation of the encoding device.
[0048] The renderer can render the decoded video / video. The rendered video / video can be displayed via the display unit.
[0049] Figure 2 is a diagram schematically explaining the configuration of a video / video encoding device to which the embodiments of this document can be applied. Hereinafter, the encoding device can include a video encoding device and / or a video encoding device.
[0050] Referring to FIG. 2, the encoding apparatus 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be called a reconstructer or a reconstructed block generator. The aforementioned image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be constituted by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (Decoded Picture Buffer) and can also be constituted by a digital storage medium. The above hardware components can further include the memory 270 as an internal / external component.
[0051] The video segmentation unit 210 can divide the input video (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-Tree Binary-Tree Ternary-Tree) structure. For example, one coding unit can be divided into multiple coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. In this case, for example, the quad-tree structure can be applied first, and the binary-tree structure and / or the ternary-tree structure can be applied thereafter. Alternatively, the binary-tree structure can also be applied first. The coding procedure according to this document can be executed based on the final coding unit that cannot be further divided. In this case, based on the coding efficiency according to video characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth so that the coding unit with an optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, conversion, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can each be divided or partitioned from the aforementioned final coding unit.The above prediction unit is a unit for sample prediction, and the above conversion unit is a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.
[0052] The term "unit" can, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally also represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel in one picture (or video).
[0053] The encoding device 200 can subtract the prediction signal (predicted block, predicted sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from the input video signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input video signal (original block, original sample array) within the encoder 200 can be called the subtraction unit 231. The prediction unit can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a predicted block including the predicted samples for the current block. The prediction unit can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240 as described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.
[0054] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to the current block or remotely located depending on the prediction mode. The prediction modes in intra prediction can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of fineness of the prediction direction. However, this is merely an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent blocks.
[0055] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The above motion information can include a motion vector and a reference picture index. The above motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the above reference block and the reference picture including the above temporal neighboring block may be the same or different. The above temporal neighboring blocks can be called by names such as collocated reference blocks and collocated CUs (colCUs), and the reference picture including the above temporal neighboring blocks can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the case of the Motion Vector Prediction (MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0056] The prediction unit 220 can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for predicting a block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can be based on the Intra Block Copy (IBC) prediction mode for predicting a block, or can be based on the palette mode. The IBC prediction mode or the palette mode can be used for content video / motion video coding such as games, for example, like SCC (Screen Content Coding). IBC basically performs prediction within the current picture, but can be executed in a manner similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on information regarding the palette table and the palette index.
[0057] The prediction signal generated through the above prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when representing the relationship information between pixels with a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to a pixel block having the same size of a square, and can also be applied to blocks of various (variable) sizes that are not square.
[0058] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be referred to as residual information. The quantization unit 233 can reorder the block-form quantized transform coefficients in a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can execute various encoding methods such as exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), CABAC (Context-Adaptive Binary Arithmetic Coding), etc. In addition to the quantized transform coefficients, the entropy encoding unit 240 can also encode, either together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.). The encoded information (e.g., encoded video / video information) can be transmitted or stored in the form of a bitstream in units of NAL (Network Abstraction Layer) units. The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information can further include general constraint information. In this document, the information and / or syntax elements transmitted / signaled from the encoding device to the decoding device can be included in the video / video information. The video / video information can be encoded through the aforementioned encoding procedure and included in the bitstream.The above bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 240 can be configured as internal / external elements of the encoding device 200, or the transmitting unit can also be included in the entropy encoding unit 240.
[0059] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transformation unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 250 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and, as will be described later, can also be used for inter prediction of the next picture after passing through filtering.
[0060] On the other hand, LMCS (Luma Mapping with Chroma Scaling) can also be applied in the picture encoding and / or restoration process.
[0061] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240, as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.
[0062] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve the encoding efficiency.
[0063] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the block where the motion information in the current picture was derived (or encoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.
[0064] FIG. 3 is a diagram schematically illustrating the configuration of a video / video decoding apparatus to which the embodiments of this document can be applied. Hereinafter, the decoding apparatus can include a video decoding apparatus and / or a video decoding apparatus.
[0065] Referring to FIG. 3, the decoding apparatus 300 can include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 can include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 321. The entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filtering unit 350 described above can be configured by one hardware component (for example, a decoder chipset or a processor) according to an embodiment. Further, the memory 360 can include a DPB (Decoded Picture Buffer) and can also be configured by a digital storage medium. The above hardware component can further include the memory 360 as an internal / external component.
[0066] When a bitstream including video / video information is input, the decoding device 300 can restore the video corresponding to the process in which the video / video information was processed by the encoding device of FIG. 2. For example, the decoding device 300 can derive units / blocks based on the block division related information obtained from the above bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding is, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximum coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. Then, the restored video signal decoded and output via the decoding device 300 can be played back via a playback device.
[0067] The decoding device 300 can receive the signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the above bitstream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The above video / video information can further include information regarding various parameter sets such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the above video / video information can further include general constraint information. The decoding device can decode a picture based on the information regarding the above parameter set and / or the above general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the above decoding procedure and obtained from the above bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration, the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the syntax element information (and adjacent to) the decoding target syntax element, the decoding information of the decoding target block, or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin by the determined context model, and executes arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit (inter prediction unit 332 and intra prediction unit 331), and the residual value for which entropy decoding is executed by the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information can be input to the residual processing unit 320. The residual processing unit 320 can derive a residual signal (residual block, residual sample, residual sample array). Also, among the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / video / picture decoding device, and the above decoding device can also be classified (categorized) into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The above information decoder can include the above entropy decoding unit 310, and the above sample decoder can include at least one of the above inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.
[0068] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output the transform coefficients. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the above reordering can be performed based on the coefficient scan order executed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using quantization parameters (for example, quantization step size information) to obtain the transform coefficients.
[0069] In the inverse transform unit 322, the transform coefficients are inverse transformed to obtain a residual signal (residual block, residual sample array).
[0070] The prediction unit can perform prediction on the current block and generate a predicted block including the prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.
[0071] The prediction unit 320 can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for the prediction of one block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can be based on the Intra Block Copy (IBC) prediction mode for the prediction of a block, or can be based on the palette mode. The IBC prediction mode or the palette mode can be used for content video / moving video coding such as games, for example, like SCC (Screen Content Coding). IBC basically performs prediction within the current picture, but can be executed in a way similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index can be included in and signaled in the video / video information.
[0072] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to or away from the current block depending on the prediction mode. The prediction modes in intra prediction can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to an adjacent block.
[0073] The inter prediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The above motion information can include a motion vector and a reference picture index. The above motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 332 can construct a motion information candidate list based on adjacent blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.
[0074] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the prediction signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the processing target block, as in the case where the skip mode is applied, the predicted block can be used as the restored block.
[0075] The addition unit 340 can be called a restoration unit or a restoration block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and as will be described later, it can also be output after filtering, or can be used for inter prediction of the next picture.
[0076] On the other hand, LMCS (Luma Mapping with Chroma Scaling) can also be applied during the picture decoding process.
[0077] The filtering unit 350 can apply filtering to the restoration signal to improve the subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0078] The (modified) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the blocks for which the motion information within the current picture has been derived (or decoded) and / or the motion information of the blocks within the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 for being utilized as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks within the current picture and can transmit them to the intra prediction unit 331.
[0079] In this document, the examples described in the filtering unit 260, the inter prediction unit 221, and the intra prediction unit 222 of the encoding apparatus 200 can be applied in the same or corresponding manner to the filtering unit 350, the inter prediction unit 332, and the intra prediction unit 331 of the decoding apparatus 300, respectively.
[0080] As described above, in performing video coding, prediction is executed to increase the compression efficiency. Through this, a predicted block including prediction samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is also derived in the encoding apparatus and the decoding apparatus, and the encoding apparatus can increase the video coding efficiency by signaling information (residual information) regarding the residual between the original block and the predicted block, which is not the original sample value of the original block, to the decoding apparatus. The decoding apparatus can derive a residual block including residual samples based on the residual information, and combine the residual block and the predicted block to generate a restored block including restored samples, and can generate a restored picture including the restored block.
[0081] The residual information can be generated through the conversion and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, execute a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, and execute a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, thereby signaling the relevant residual information (via a bitstream) to a decoding device. Here, the residual information can include information such as the value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can execute an inverse quantization / inverse conversion process based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse convert the quantized conversion coefficients for reference in the inter prediction of subsequent pictures to derive a residual block and generate a restored picture based on this.
[0082] FIG. 4 shows an example of a schematic video / film encoding method to which the embodiments of this document are applicable.
[0083] The method disclosed in FIG. 4 can be executed by the encoding device 200 in FIG. 2 described above. Specifically, S400 can be executed by the inter prediction unit 221 or the intra prediction unit 222 of the encoding device 200, and S410, S420, S430, S440 can be executed by the subtraction unit 231, the conversion unit 232, the quantization unit 233, and the entropy encoding unit 240 of the encoding device 200, respectively.
[0084] Referring to FIG. 4, the encoding device can derive a prediction sample through prediction for the current block (S400). The encoding device can determine whether to perform inter prediction or intra prediction on the current block, and can determine a specific inter prediction mode or a specific intra prediction mode based on the RD cost. According to the determined mode, the encoding device can derive a prediction sample for the current block.
[0085] The encoding device can compare the original sample and the prediction sample for the current block to derive a residual sample (S410).
[0086] The encoding device can derive conversion coefficients through a conversion procedure for the residual sample (S420), quantize the derived conversion coefficients, and derive the quantized conversion coefficients (S430).
[0087] The encoding device can encode video information including prediction information and residual information, and output the encoded video information in the form of a bitstream (S440). The prediction information is information related to the prediction procedure, and can include prediction mode information and information related to motion information (for example, when inter prediction is applied). The residual information can include information related to the quantized conversion coefficients. The residual information can be entropy encoded.
[0088] The output bitstream can be transmitted to the decoding device via a storage medium or a network.
[0089] FIG. 5 shows an example of a schematic video / video decoding method to which the embodiments of this document are applicable.
[0090] The method disclosed in FIG. 5 can be executed by the decoding device 300 of FIG. 3 described above. Specifically, S500 can be executed by the inter prediction unit 332 or the intra prediction unit 331 of the decoding device 300. In S500, the procedure of decoding the prediction information included in the bitstream to derive the values of relevant syntax elements can be executed by the entropy decoding unit 310 of the decoding device 300. S510, S520, S530, and S540 can be executed by the entropy decoding unit 310, the inverse quantization unit 321, the inverse transform unit 322, and the addition unit 340 of the decoding device 300, respectively.
[0091] Referring to FIG. 5, the decoding device can execute operations corresponding to the operations executed by the encoding device. The decoding device can execute inter prediction or intra prediction for the current block based on the received prediction information to derive a prediction sample (S500).
[0092] The decoding device can derive the quantized transform coefficients for the current block based on the received residual information (S510). The decoding device can derive the quantized transform coefficients from the residual information through entropy decoding.
[0093] The decoding device can inverse-quantize the quantized transform coefficients to derive the transform coefficients (S520).
[0094] The decoding device derives the residual samples through an inverse transform process for the transform coefficients (S530).
[0095] The decoding device can generate a restored sample for the current block based on the prediction sample and the residual sample, and generate a restored picture based on this (S540). As described above, an in-loop filtering procedure can be further applied to the restored picture later.
[0096] On the one hand, as described above, when performing prediction on the current block, intra prediction or inter prediction can be applied. Hereinafter, the case where inter prediction is applied to the current block will be described.
[0097] The prediction unit (more specifically, the inter prediction unit) of the encoding / decoding apparatus can derive a prediction sample by performing inter prediction in block units. Inter prediction can indicate a prediction derived in a way that depends on data elements (such as sample values or motion information) of pictures other than the current picture. When inter prediction is applied to the current block, a predicted block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture pointed to by a reference picture index. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the adjacent block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When inter prediction is applied, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block can be called by names such as a collocated reference block or a collocated CU (colCU), and the reference picture including the temporal neighboring block can also be called a collocated picture (colPic). For example, a motion information candidate list can be configured based on the adjacent blocks of the current block, and a flag or index information indicating which candidate is selected (used) can be signaled in order to derive the motion vector and / or reference picture index of the current block.Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the motion information of the current block can be the same as that of the selected adjacent block. In the case of the skip mode, different from the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of the selected adjacent block is used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.
[0098] The above motion information can include L0 motion information and / or L1 motion information depending on the inter prediction type (such as L0 prediction, L1 prediction, Bi prediction, etc.). The motion vector in the L0 direction may be referred to as the L0 motion vector or MVL0, and the motion vector in the L1 direction may be referred to as the L1 motion vector or MVL1. The prediction based on the L0 motion vector may be called L0 prediction, the prediction based on the L1 motion vector may be called L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector may be called bi-prediction. Here, the L0 motion vector can indicate the motion vector related to the reference picture list L0 (L0), and the L1 motion vector can indicate the motion vector related to the reference picture list L1 (L1). The reference picture list L0 can include, as reference pictures, pictures prior in output order to the current picture, and the reference picture list L1 can include pictures subsequent in output order to the current picture. The prior pictures may be called forward (reference) pictures, and the subsequent pictures may be called backward (reference) pictures. The reference picture list L0 can further include, as reference pictures, pictures subsequent in output order to the current picture. In this case, the prior pictures within the reference picture list L0 can be indexed first, and the subsequent pictures can be indexed next. The reference picture list L1 can further include, as reference pictures, pictures prior in output order to the current picture. In this case, the subsequent pictures within the reference picture list L1 can be indexed first, and the prior pictures can be indexed next. Here, the output order may correspond to the POC (Picture Order Count) order.
[0099] Also, for predicting the current block within a picture, various inter-prediction modes can be used. For example, various modes such as merge mode, skip mode, MVP (Motion Vector Prediction) mode, affine mode, sub-block merge mode, MMVD (Merge with MVD) mode, HMVP (Historical Motion Vector Prediction) mode, etc. can be used. DMVR (Decoder side Motion Vector Refinement) mode, AMVR (Adaptive Motion Vector Resolution) mode, Bi-prediction with CU-level Weight (BCW), Bi-Directional Optical Flow (BDOF), etc. can be further used as additional modes. The affine mode can also be called the affine motion prediction mode. The MVP mode can also be called the AMVP (Advanced Motion Vector Prediction) mode. In this document, some modes and / or motion information candidates derived by some modes can also be included as one of the motion information related candidates of other modes. For example, the HMVP candidate can be added as a merge candidate for the merge / skip mode, or can also be added as an mvp candidate for the MVP mode. When the HMVP candidate is used as a motion information candidate for the merge mode or skip mode, the HMVP candidate can be called the HMVP merge candidate.
[0100] Prediction mode information indicating the inter prediction mode of the current block can be signaled from an encoding device to a decoding device. At this time, the prediction mode information can be included in a bitstream and received by the decoding device. The prediction mode information can include index information indicating one of a number (plural) of candidate modes. Alternatively, the inter prediction mode can also be indicated through hierarchical signaling of flag information. In this case, the prediction mode information can include one or more flags. For example, a skip flag is signaled to indicate whether the skip mode can be applied. When the skip mode is not applied, a merge flag is signaled to indicate whether the merge mode can be applied. When the merge mode is not applied, it can be indicated that the MVP mode is applied, or additional flags for classification (distinction) can also be signaled. The affine mode can be signaled as an independent mode, or can also be signaled as a mode subordinate to the merge mode or the MVP mode, etc. For example, the affine mode can include an affine merge mode and an affine MVP mode.
[0101] In addition, when applying inter prediction to the current block, the motion information of the current block can be utilized. The encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding device can use the original block in the original picture for the current block to search for a highly correlated similar reference block within a predetermined search range in the reference picture in fractional pixel units, and through this, derive the motion information. The similarity of the blocks can be derived based on the difference in sample values based on phase. For example, the similarity of the blocks can be calculated based on the SAD (Sum of Absolute Differences) between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, the motion information can be derived based on the reference block with the smallest SAD in the search space (search area). The derived motion information can be signaled to the decoding device according to various methods based on the inter prediction mode.
[0102] As described above, based on the motion information derived by the inter prediction mode, a predicted block for the current block can be derived. The predicted block can include the prediction samples (prediction sample array) of the current block. When the motion vector (MV) of the current block indicates a fractional sample unit, an interpolation procedure can be executed, through which the prediction samples of the current block can be derived based on the reference samples of the fractional sample unit in the reference picture. When Affine inter prediction is applied to the current block, prediction samples can be generated based on the sample / sub-block unit MV. When bi-prediction is applied, the prediction samples derived through the weighted sum or weighted average (by phase) of the prediction samples derived based on the L0 prediction (i.e., prediction using the reference picture and MVL0 in the reference picture list L0) and the prediction samples derived based on the L1 prediction (i.e., prediction using the reference picture and MVL1 in the reference picture list L1) can be used as the prediction samples of the current block. When bi-prediction is applied, if the reference picture used for the L0 prediction and the reference picture used for the L1 prediction are located in different temporal directions with respect to the current picture (i.e., when it is bi-prediction and at the same time corresponds to bi-directional prediction), this can be called true bi-prediction.
[0103] As described above, it is as previously mentioned that the restored samples and the restored picture can be generated based on the prediction samples derived as above, and then procedures such as in-loop filtering can be executed.
[0104] FIG. 6 shows an example of a video / video encoding method based on schematic inter prediction to which the embodiments of this document are applicable.
[0105] The method disclosed in FIG. 6 can be executed by the encoding apparatus 200 of FIG. 2 described above. Specifically, S600 can be executed by the inter prediction unit 221 of the encoding apparatus 200, S610 can be executed by the subtraction unit 231 of the encoding apparatus 200, and S620 can be executed by the entropy encoding unit 240 of the encoding apparatus 200.
[0106] Referring to FIG. 6, the encoding apparatus can perform inter prediction on the current block (S600). The encoding apparatus can derive the inter prediction mode and motion information of the current block and generate a prediction sample of the current block. Here, the process of determining the inter prediction mode, deriving the motion information, and generating the prediction sample may be executed simultaneously, or one process may be executed prior to the other processes. For example, the inter prediction unit of the encoding apparatus can include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit can determine the prediction mode for the current block, the motion information derivation unit can derive the motion information of the current block, and the prediction sample derivation unit can derive the prediction sample of the current block. For example, the inter prediction unit of the encoding apparatus can search for a block similar to the current block within a predetermined region (search region) of the reference picture through motion estimation, and derive a reference block whose difference from the current block is the smallest or below a predetermined criterion. Based on this, a reference picture index indicating the reference picture in which the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding apparatus can determine the mode to be applied to the current block among various prediction modes. The encoding apparatus can compare the RD costs for various prediction modes and determine the optimal prediction mode for the current block.
[0107] For example, when the skip mode or the merge mode is applied to the current block, the encoding device can construct a merge candidate list, and derive a reference block with the smallest difference from the current block or a difference less than or equal to a predetermined criterion among the reference blocks indicated by the merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block can be selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.
[0108] As another example, when the (A)MVP mode is applied to the current block, the encoding device can construct an (A)MVP candidate list, and use the motion vector of the selected mvp (motion vector predictor) candidate among the mvp candidates included in the (A)MVP candidate list as the mvp of the current block. In this case, for example, the motion vector indicating the reference block derived by the above-described motion estimation can be used as the motion vector of the current block, and the mvp candidate having the motion vector with the smallest difference from the motion vector of the current block among the mvp candidates can be the selected mvp candidate. An MVD (Motion Vector Difference), which is the difference obtained by subtracting the mvp from the motion vector of the current block, can be derived. In this case, information regarding the MVD can be signaled to the decoding device. Also, when the (A)MVP mode is applied, the value of the reference picture index can be composed of reference picture index information and signaled to the decoding device separately.
[0109] The encoding device can derive a residual sample based on a prediction sample (S610). The encoding device can derive a residual sample by comparing the original sample of the current block with the prediction sample.
[0110] The encoding device can encode video information including prediction information and residual information (S620). The encoding device can output the encoded video information in the form of a bitstream. Here, the prediction information can include, as information regarding the prediction procedure, prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and information regarding motion information. The information regarding motion information can include candidate selection information (e.g., merge index, mvp flag, or mvp index) which is information for deriving a motion vector. Also, the information regarding motion information can include information regarding the aforementioned MVD and / or reference picture index information. Also, the information regarding motion information can include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information regarding residual samples. The residual information can include information regarding quantized transform coefficients for the residual samples.
[0111] The output bitstream can be stored in a (digital) storage medium and transmitted to the decoding device, or can also be transmitted to the decoding device via a network.
[0112] On the other hand, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is because the same prediction result as that executed by the decoding device is derived by the encoding device, and through this, the coding efficiency can be improved. Therefore, the encoding device can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in the memory and utilize it as a reference picture for inter prediction. As described above, an in-loop filtering procedure or the like can be further applied to the reconstructed picture.
[0113] FIG. 7 shows an example of a video / video decoding method based on a schematic inter prediction to which the embodiments of this document are applicable.
[0114] The method disclosed in FIG. 7 can be executed by the decoding device 300 of FIG. 3 described above. Specifically, S700 can be executed by the inter prediction unit 332 of the decoding device 300. In S700, the procedure of decoding the prediction information included in the bitstream to derive the value of the related syntax element can be executed by the entropy decoding unit 310 of the decoding device 300. S710 and S720 can be executed by the inter prediction unit 332 of the decoding device 300, S730 can be executed by the residual processing unit 320 of the decoding device 300, and S740 can be executed by the addition unit 340 of the decoding device 300.
[0115] Referring to FIG. 7, the decoding device can execute operations corresponding to the operations executed by the encoding device. The decoding device can perform a prediction on the current block based on the received prediction information and derive a prediction sample.
[0116] Specifically, the decoding device can determine a prediction mode for the current block based on the received prediction information (S700). The decoding device can determine which inter prediction mode is applicable to the current block based on the prediction mode information in the prediction information.
[0117] For example, based on the merge flag, it can be determined whether the merge mode is applicable to the current block, or whether the (A)MVP mode is determined. Alternatively, based on the mode index, one of various inter prediction mode candidates can be selected. The inter prediction mode candidates can include the skip mode, the merge mode, and / or the (A)MVP mode, or can include various inter prediction modes.
[0118] The decoding device can derive the motion information of the current block based on a determined inter-prediction mode (S710). For example, when the skip mode or the merge mode is applied to the current block, the decoding device can construct a merge candidate list and select one merge candidate from the merge candidates included in the merge candidate list. The above selection can be executed based on the aforementioned selection information (merge index). Using the motion information of the selected merge candidate, the motion information of the current block can be derived. The motion information of the selected merge candidate can be used as the motion information of the current block.
[0119] As another example, when the (A)MVP mode is applied to the current block, the decoding device can construct an (A)MVP candidate list and use the motion vector of the mvp (motion vector predictor) candidate selected from among the mvp candidates included in the (A)MVP candidate list as the mvp of the current block. The above selection can be executed based on the aforementioned selection information (mvp flag or mvp index). In this case, the MVD of the current block can be derived based on the information regarding MVD, and the motion vector of the current block can be derived based on the mvp and MVD of the current block. Also, the reference picture index of the current block can be derived based on the reference picture index information. The picture indicated by the reference picture index within the reference picture list regarding the current block can be derived as the reference picture to be referred to for the inter-prediction of the current block.
[0120] On the other hand, without constructing a candidate list, the motion information of the current block can be derived. In this case, the motion information of the current block can be derived according to the procedure disclosed in the prediction mode described later. In this case, the construction of the candidate list as described above can be omitted.
[0121] The decoding device can generate a prediction sample for the current block based on the motion information of the current block (S720). In this case, a reference picture can be derived based on the reference picture index of the current block, and the prediction sample of the current block can be derived using the samples of the reference block indicated by the motion vector of the current block on the reference picture. In this case, as will be described later, in some cases, a prediction sample filtering procedure for all or part of the prediction sample of the current block can be further executed.
[0122] For example, the inter prediction unit of the decoding device can include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. Based on the prediction mode information received by the prediction mode determination unit, it determines the prediction mode for the current block. Based on the information related to the motion information received by the motion information derivation unit, it derives the motion information (such as a motion vector and / or a reference picture index, etc.) of the current block. And the prediction sample derivation unit can derive the prediction sample of the current block.
[0123] The decoding device can generate a residual sample for the current block based on the received residual information (S730). The decoding device can generate a restored sample for the current block based on the prediction sample and the residual sample, and generate a restored picture based on this (S740). As described above, an in-loop filtering procedure or the like can be further applied to the restored picture.
[0124] FIG. 8 exemplarily shows an inter prediction procedure. The inter prediction procedure disclosed in FIG. 8 can be applied to the inter prediction process (when the inter prediction mode is applied) disclosed in FIGS. 6 and 7 described above.
[0125] Referring to FIG. 8, as described above, the inter prediction procedure can include an inter prediction mode determination step, a motion information derivation step according to the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information. The inter prediction procedure can be executed in an encoding device and a decoding device, as described above. In this document, the coding device can include an encoding device and / or a decoding device.
[0126] The coding device can determine an inter prediction mode for the current block (S800). For prediction of the current block within a picture, various inter prediction modes can be used. For example, various modes such as a merge mode, a skip mode, an MVP (Motion Vector Prediction) mode, an affine mode, a sub-block merge mode, an MMVD (Merge with MVD) mode, etc. can be used. A DMVR (Decoder side Motion Vector Refinement) mode, an AMVR (Adaptive Motion Vector Resolution) mode, a Bi-prediction with CU-level Weight (BCW), a Bi-Directional Optical Flow (BDOF), etc. can be used additionally or alternatively as accompanying modes. The affine mode can also be called an affine motion prediction mode. The MVP mode can also be called an AMVP (Advanced Motion Vector Prediction) mode. In this document, some modes and / or motion information candidates derived by some modes can also be included as one of the motion information related candidates of other modes. For example, an HMVP candidate can be added as a merge candidate for the merge / skip mode, or can also be added as an MVP candidate for the MVP mode. When the HMVP candidate is used as a motion information candidate for the merge mode or the skip mode, the HMVP candidate can be called an HMVP merge candidate.
[0127] Prediction mode information indicating the inter prediction mode of the current block can be signaled from an encoding device to a decoding device. The prediction mode information can be included in a bitstream and received by the decoding device. The prediction mode information can include index information indicating one of a number of candidate modes. Alternatively, the inter prediction mode can also be indicated through hierarchical signaling of flag information. In this case, the prediction mode information can include one or more flags. For example, a skip flag is signaled to indicate whether the skip mode can be applied. If the skip mode is not applied, a merge flag is signaled to indicate whether the merge mode can be applied. If the merge mode is not applied, it can be indicated that the MVP mode has been applied, or additional flags for further classification can also be signaled. The affine mode can be signaled as an independent mode, or can also be signaled as a mode subordinate to the merge mode or the MVP mode, etc. For example, the affine mode can include an affine merge mode and an affine MVP mode.
[0128] The coding device can derive motion information for the current block (S810). The motion information derivation can be derived based on the inter prediction mode.
[0129] The coding device can perform inter prediction using the motion information of the current block. The encoding device can derive the optimal motion information for the current block through a motion estimation procedure. For example, the encoding device can search for a highly correlated similar reference block within a predetermined search range in the reference picture in pixel units of a fraction using the original block in the original picture for the current block, and through this, the motion information can be derived. The similarity of the blocks can be derived based on the difference in sample values based on phase. For example, the similarity of the blocks can be calculated based on the SAD between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, the motion information can be derived based on the reference block with the smallest SAD within the search space. The derived motion information can be signaled to the decoding device in various ways based on the inter prediction mode.
[0130] The coding device can perform inter prediction based on the motion information for the current block (S820). The coding device can derive a predicted sample for the current block based on the motion information. The current block including the predicted sample can be called a predicted block.
[0131] On the other hand, when deriving the motion information of the current block, motion information candidates can be derived based on spatially adjacent blocks and temporally adjacent blocks of the current block, and motion information candidates for the current block can be selected based on the derived motion information candidates. At this time, the selected motion information candidate can be used as the motion information of the current block.
[0132] FIG. 9 exemplarily shows the spatially adjacent blocks and temporally adjacent blocks of the current block.
[0133] Referring to FIG. 9, the spatial adjacent block refers to an adjacent block located around the current block 900 that is the target for performing the current inter prediction, and can include an adjacent block located around the left side of the current block 900 or an adjacent block located around the upper side of the current block 900. For example, the spatial adjacent block can include the lower left corner adjacent block, the left adjacent block, the upper right corner adjacent block, the upper adjacent block, and the upper left corner adjacent block of the current block 900. In FIG. 9, the spatial adjacent block is indicated by "S".
[0134] In one embodiment, the encoding device / decoding device can search for the spatial adjacent blocks (e.g., the lower left corner adjacent block, the left adjacent block, the upper right corner adjacent block, the upper adjacent block, the upper left corner adjacent block) of the current block in a specified order to detect available adjacent blocks, and derive the motion information of the detected adjacent blocks as spatial motion information candidates.
[0135] A temporal neighboring block is a block located on a picture different from the current picture (i.e., a reference picture) that includes the current block 900, and refers to a block (a collocated block; col block) at the same position as the current block 900 within the reference picture. Here, the reference picture can be before or after the current picture in terms of POC (Picture Order Count). Also, the reference picture used when deriving the temporal neighboring block can be referred to as a collocated reference picture or a col picture. Also, a collocated block can indicate a block located within the col picture corresponding to the position of the current block 900, and can be referred to as a col block. For example, as shown in FIG. 9, the temporal neighboring block can include a col block (i.e., a col block including the lower-right corner sample) located corresponding to the lower-right corner sample position of the current block 900 within the reference picture (i.e., the col picture) and / or a col block (i.e., a col block including the center lower-right sample) located corresponding to the center lower-right sample position of the current block 900 within the reference picture (i.e., the col picture). In FIG. 9, the temporal neighboring block is indicated by "T".
[0136] In one embodiment, the encoding device / decoding device can search for temporal neighboring blocks of the current block (e.g., a col block including the lower-right corner sample, a col block including the center lower-right sample) in a specified order to detect available blocks, and derive the motion information of the detected blocks as temporal motion information candidates. Thus, the technique of using temporal neighboring blocks can be referred to as TMVP (Temporal Motion Vector Prediction). Also, the temporal motion information candidates can be referred to as TMVP candidates.
[0137] On the other hand, depending on the inter-prediction mode, motion information can also be derived in units of sub-blocks to perform prediction. For example, in the case of the affine mode or the TMVP mode, motion information can be derived in units of sub-blocks. In particular, the method of deriving candidates for temporal motion information in units of sub-blocks can be referred to as sub-block-based Temporal Motion Vector Prediction (sbTMVP).
[0138] sbTMVP is a method that utilizes the motion field in the col picture to improve the motion vector prediction (MVP) and merge mode of the coding unit within the current picture. The col picture of sbTMVP can be the same as the col picture used by TMVP. However, TMVP performs motion prediction at the coding unit (CU) level, while sbTMVP can perform motion prediction at the sub-block level or the sub-coding unit (sub-CU) level. Also, TMVP derives temporal motion information from the col block within the col picture (where the col block corresponds to the lower-right corner sample position of the current block or the center lower-right sample position of the current block), while sbTMVP derives temporal motion information after applying a motion shift to the col picture. Here, the motion shift can include the process of obtaining a motion vector from one of the spatially adjacent blocks of the current block and shifting by the above motion vector.
[0139] FIG. 10 exemplarily shows the spatially adjacent blocks that can be used to derive candidates for sub-block-based temporal motion information (sbTMVP candidates).
[0140] Referring to FIG. 10, the spatially adjacent block can include at least one of the lower left corner adjacent block A0, the left adjacent block A1, the upper right corner adjacent block B0, and the upper adjacent block B1 of the current block. In some cases, the spatially adjacent block may further include other adjacent blocks other than the adjacent blocks shown in FIG. 10, or may not include a specific adjacent block among the adjacent blocks shown in FIG. 10. Also, the spatially adjacent block may include only a specific adjacent block. For example, it may include only the left adjacent block A1 of the current block.
[0141] For example, the encoding device / decoding device detects the motion vector of the earliest available spatially adjacent block while searching for the spatially adjacent blocks in a predetermined search order, and determines the block at the position indicated by the motion vector of the spatially adjacent block in the reference picture as the col block (i.e., the collocated reference block). Here, the motion vector of the spatially adjacent block may be referred to as a temporal motion vector (temporal MV).
[0142] At this time, the usability of the spatially adjacent block can be determined based on the reference picture information, prediction mode information, position information, etc. of the spatially adjacent block. For example, when the reference picture of the spatially adjacent block is the same as the reference picture of the current block, the spatially adjacent block can be determined to be usable. Alternatively, when the spatially adjacent block is coded in the intra prediction mode or the spatially adjacent block is located outside the current picture / tile, the spatially adjacent block can be determined to be unusable.
[0143] Also, the search order of the spatially adjacent blocks can be defined in various ways. For example, it may be in the order of A1, B1, B0, A0. Alternatively, only A1 can be searched to determine whether A1 is usable.
[0144] FIG. 11 is a diagram schematically explaining the process of deriving the sub-block based temporal motion information candidate (sbTMVP candidate).
[0145] Referring to FIG. 11, first, the encoding / decoding device can determine whether the spatially adjacent block of the current block (e.g., A1 block) is available. For example, when the reference picture of the spatially adjacent block (e.g., A1 block) uses the col picture, the spatially adjacent block (e.g., A1 block) can be determined to be available, and the motion vector of the spatially adjacent block (e.g., A1 block) can be derived. At this time, the motion vector of the spatially adjacent block (e.g., A1 block) can be referred to as the temporal MV (tempMV), and this motion vector can be used for motion shift. Alternatively, if it is determined that the spatially adjacent block (e.g., A1 block) is unavailable, the temporal MV (i.e., the motion vector of the spatially adjacent block) can be set to a zero vector. In other words, in this case, the motion shift can be applied with a motion vector set to (0, 0).
[0146] Next, the encoding / decoding device can apply a motion shift based on the motion vector of the spatially adjacent block (e.g., A1 block). For example, the motion shift can be shifted (e.g., to A1`) to the position indicated by the motion vector of the spatially adjacent block (e.g., A1 block). That is, by applying the motion shift, the motion vector of the spatially adjacent block (e.g., A1 block) can be added to the coordinates of the current block.
[0147] Next, the encoding / decoding device can derive collocated subblocks that are motion-shifted on the col picture and obtain motion information (such as motion vectors, reference indices, etc.) for each collocated subblock. For example, the encoding / decoding device can derive each collocated subblock on the col picture corresponding to the position that is motion-shifted at each subblock position within the current block (i.e., the position indicated by the motion vector of the spatially adjacent block (e.g., A1)). Then, the motion information of each collocated subblock can be used as the motion information for each subblock with respect to the current block (i.e., the sbTMVP candidate).
[0148] Also, scaling can be applied to the motion vectors of the collocated subblocks. The above scaling can be performed based on the difference in the temporal distance between the reference picture of the col block and the reference picture of the current block. Therefore, the above scaling can be referred to as temporal motion scaling, through which the reference picture of the current block and the reference picture of the temporal motion vector can be aligned. In this case, the encoding / decoding device can obtain the motion vectors of the scaled collocated subblocks as the motion information for each subblock with respect to the current block.
[0149] Also, when deriving the sbTMVP candidate, there may be cases where there is no motion information in the collocated subblock. In this case, base motion information (or default motion information) can be derived for the collocated subblock without motion information, and this base motion information can be used as the motion information for the subblock with respect to the current block. The base motion information can be derived from the block located at the center of the col block (i.e., the col CU including the collocated subblock). For example, motion information (such as a motion vector) can be derived from the block including the sample located at the lower right end among the four samples located at the center of the col block and used as the base motion information.
[0150] As described above, in the case of the affine mode or the sbTMVP mode that derives motion information in sub-block units, affine merge candidates and sbTMVP candidates can be derived, and a merge candidate list based on sub-blocks can be constructed based on such candidates. At this time, for the affine mode or the sbTMVP mode, flag information indicating whether it is enabled or disabled can be signaled. When the sbTMVP mode is available based on the above flag information, the sbTMVP candidates derived as described above can be added to the firstly-ordered entry of the merge candidate list based on sub-blocks. Then, the affine merge candidates can be added to the next entry of the merge candidate list based on sub-blocks. Here, the maximum number of candidates in the merge candidate list based on sub-blocks may be five.
[0151] Also, in the case of the sbTMVP mode, the size of the sub-block can be fixed. For example, it can be fixed to a size of 8×8. Also, in the sbTMVP mode, it can be applied only to blocks whose width and height are both 8 or more.
[0152] On the other hand, in the current VVC standard, time motion information candidates (sbTMVP candidates) based on sub-blocks can be derived as shown in Table 1 below.
[0153]
Table 1-1
[0154]
Table 1-2
[0155]
Table 1-3
[0156] When deriving sbTMVP candidates according to the method as shown in Table 1 above, default MV and sub-block MV(s) may be considered. Here, the default MV may be referred to as subblock-based temporal merging base motion data or base motion data (base motion information). Referring to Table 1 above, the default MV can correspond to ctrMV (or ctrMVLX) in Table 1. The sub-block MV can correspond to mvSbCol (or mvLXSbcol) in Table 1.
[0157] For example, if a sub-block or sub-block MV is available through the sbTMVP derivation process, the above sub-block MV is assigned to the sub-block, or if the sub-block or sub-block MV is not available, the above default MV can be used as the sub-block MV for the sub-block. Here, the default MV derives motion information from a position corresponding to the center pixel position of the corresponding block (i.e., col CU) on the col picture, and each sub-block MV can derive motion information from the top-left position of the corresponding sub-block (i.e., col sub-block) on the col picture. At this time, the corresponding block (i.e., col CU) can be derived from a position where the motion is shifted based on the motion vector (i.e., temporal MV) of the spatially adjacent block A1 as shown in FIG. 11.
[0158] FIGS. 12 to 15 are diagrams schematically explaining a method of calculating a corresponding position for deriving a default MV and a sub-block MV according to a block size in the sbTMVP derivation process.
[0159] In FIGS. 12 to 15, the pixels (samples) hatched with dotted lines indicate the corresponding positions within each sub-block for deriving each sub-block MV, and the pixels (samples) hatched with solid lines indicate the corresponding positions within the CU for deriving the default MV.
[0160] For example, referring to FIG. 12, when the current block (i.e., the current CU) is 8×8 in size, the motion information of the sub-block can be derived based on the upper left sample position within the 8×8 sized sub-block, and the default motion information of the sub-block can be derived based on the center sample position within the current 8×8 sized block (i.e., the current CU).
[0161] Alternatively, for example, referring to FIG. 13, when the current block (i.e., the current CU) is 16×8 in size, the motion information of each sub-block can be derived based on the upper left sample position within each 8×8 sized sub-block, and the default motion information of each sub-block can be derived based on the center sample position within the current 16×8 sized block (i.e., the current CU).
[0162] Alternatively, for example, referring to FIG. 14, when the current block (i.e., the current CU) is 8×16 in size, the motion information of each sub-block can be derived based on the upper left sample position within each 8×8 sized sub-block, and the default motion information of each sub-block can be derived based on the center sample position within the current 8×16 sized block (i.e., the current CU).
[0163] Alternatively, for example, referring to FIG. 15, when the current block (i.e., the current CU) is 16×16 in size, the motion information of each sub-block can be derived based on the upper left sample position within each 8×8 sized sub-block, and the default motion information of each sub-block can be derived based on the center sample position within the current 16×16 sized block (i.e., the current CU).
[0164] As can be seen from FIGS. 12 to 15 described above, since the motion information of the sub-block is biased towards the upper left pixel position, there is a problem that the sub-block MV is induced at a position far from the position where the default MV indicating the representative motion information of the current CU is derived. As a simple example, in the case of an 8×8 block shown in FIG. 12, one CU includes one sub-block, but there is a contradiction in that the sub-block MV and the default MV are represented by different motion information. Also, since the method of calculating the corresponding positions between the sub-block and the current CU block is different (that is, the corresponding position for deriving the MV of the sub-block is the upper left sample position, and the corresponding position for deriving the default MV is the center sample position), additional modules may be required when implementing the hardware (H / W).
[0165] Therefore, in order to improve the above problems, in this document, a method of integrating the method of deriving the corresponding position of the CU for the default MV and the method of deriving the corresponding position of the sub-block for each sub-block MV in the process of deriving the sbTMVP candidate is proposed. According to an embodiment in this document, it has the effect of integration by using only one module for deriving each corresponding position according to the block size from the perspective of hardware (H / W). For example, the method of calculating the corresponding position when the block size is a 16×16 block and the method of calculating the corresponding position when the block size is an 8×8 block can be implemented identically, so a simplification effect can be achieved in terms of the implementation of the hardware. Here, a 16×16 block can represent a CU, and an 8×8 block can represent each sub-block.
[0166] In one embodiment, when deriving the sbTMVP candidate, the center sample position can be used as the corresponding position for deriving the motion information of the sub-block and as the corresponding position for deriving the default motion information, and it can be implemented as shown in Table 2 below.
[0167] The following Table 2 is a spec showing an example of a method for deriving sub-block motion information and default motion information according to an embodiment in this document.
[0168]
Table 2-1
[0169]
Table 2-2
[0170]
Table 2-3
[0171] Referring to Table 2 above, when deriving the sbTMVP candidate, the position of the current block (i.e., the current CU) including the sub-block can be derived. The left-upper sample position (xCtb, yCtb) of the coding tree block (or coding tree unit) including the current block and the right-lower center sample position (xCtr, yCtr) of the current block can be derived as in the formulas (8-514) to (8-517) of Table 2 above. At this time, the positions (xCtb, yCtb) and (xCtr, yCtr) can be calculated based on the left-upper sample position (xCb, yCb) of the current block with reference to the left-upper sample of the current picture.
[0172] Also, the col block (i.e., col CU) on the col picture corresponding to the current block (i.e., the current CU) including the sub-block can be derived. At this time, the position of the col block can be set to (xColCtrCb, yColCtrCb), and this position can indicate the position of the col block including the position of (xCtr, yCtr) in the col picture with reference to the left-upper sample of the col picture.
[0173] In addition, base motion data (i.e., default motion information) for sbTMVP can be derived. The base motion data can include a default MV (e.g., ctrMvLX). For example, col blocks on a col picture can be derived. At this time, the position of the col block can be derived as (xColCb, yColCb), and this position can be the position where a motion shift (e.g., tempMv) is applied to the above-derived col block position (xColCtrCb, yColCtrCb). The motion shift can be executed by adding a motion vector (e.g., tempMv) derived from a spatially adjacent block (e.g., A1 block) of the current block to the current col block position (xColCtrCb, yColCtrCb) as described above. Next, a default MV (e.g., ctrMvLX) can be derived based on the position (xColCb, yColCb) of the motion-shifted col block. Here, the default MV (e.g., ctrMvLX) can indicate a motion vector derived from a position corresponding to the lower right center sample of the col block.
[0174] In addition, col sub-blocks on the col picture corresponding to sub-blocks (referred to as current sub-blocks) within the current block can be derived. First, the positions of each of the current sub-blocks can be derived. The position of each sub-block can be represented as (xSb, ySb), and this position (xSb, ySb) can indicate the position of the current sub-block with reference to the upper left sample of the current picture. For example, the position (xSb, ySb) of the current sub-block can be calculated as in the mathematical formulas (8-523) to (8-524) in Table 2 above, which can indicate the position of the center sample at the lower right end of the sub-block. Next, the positions of each of the col sub-blocks on the col picture can be derived. The position of each col sub-block can be represented as (xColSb, yColSb), and this position (xColSb, yColSb) can be the position where a motion shift (e.g., tempMv) is applied to the position (xSb, ySb) of the current sub-block. The motion shift can be executed, as described above, by adding a motion vector (e.g., tempMv) derived from a spatially adjacent block (e.g., A1 block) of the current block to the position (xSb, ySb) of the current sub-block. Next, motion information of the col sub-blocks (e.g., motion vector mvLXSbCol, flag availableFlagLXSbCol indicating whether it is available) can be derived based on the positions (xColSb, yColSb) of each of the motion-shifted col sub-blocks.
[0175] At this time, when there are col sub-blocks that are unusable in the col sub-blocks (e.g., when availableFlagLXSbCol is 0), base motion data (i.e., default motion information) can be used for the unusable col sub-blocks. For example, for the motion vector (e.g., mvLXSbCol) of an unusable col sub-block, a default MV (e.g., ctrMvLX) can be used.
[0176] FIGS. 16 to 19 are exemplary diagrams schematically illustrating a method of integrating corresponding positions for deriving a default MV and sub-block MVs according to a block size in the derivation process of sbTMVP.
[0177] In FIGS. 16 to 19, the pixels (samples) hatched with a dotted line indicate the corresponding positions within each sub-block for deriving each sub-block MV, and the pixels (samples) hatched in practice indicate the corresponding positions within the CU for deriving the default MV.
[0178] For example, referring to FIG. 16, when the current block (i.e., the current CU) is of 8×8 size, the motion information of the current sub-block can be derived from the corresponding position of the col sub-block on the col picture based on the right lower center sample position within the 8×8 size sub-block and used. The default motion information of the current sub-block can be derived from the corresponding position of the col block (i.e., the col CU) on the col picture based on the right lower center sample position within the current 8×8 size block (i.e., the current CU) and used. In this case, as shown in FIG. 16, the motion information of the current sub-block and the default motion information can be derived from the same sample position (the same corresponding position).
[0179] Alternatively, for example, referring to FIG. 17, when the current block (i.e., the current CU) is of 16×8 size, the motion information of the current sub-block can be derived from the corresponding position of the col sub-block on the col picture based on the right lower center sample position within the 8×8 size sub-block and used. The default motion information of the current sub-block can be derived from the corresponding position of the col block (i.e., the col CU) on the col picture based on the right lower center sample position within the current 16×8 size block (i.e., the current CU) and used.
[0180] Alternatively, for example, referring to FIG. 18, when the current block (i.e., the current CU) is 8×16 in size, the motion information of the current sub-block can be derived from the corresponding position of the col sub-block on the col picture based on the right-bottom center sample position within the 8×8 size sub-block and used. The default motion information of the current sub-block can be derived from the corresponding position of the col block (i.e., the col CU) on the col picture based on the right-bottom center sample position within the current block (i.e., the current CU) of 8×16 size and used.
[0181] Alternatively, for example, referring to FIG. 19, when the current block (i.e., the current CU) is 16×16 in size or larger, the motion information of the current sub-block can be derived from the corresponding position of the col sub-block on the col picture based on the right-bottom center sample position within the 8×8 size sub-block and used. The default motion information of the current sub-block can be derived from the corresponding position of the col block (i.e., the col CU) on the col picture based on the right-bottom center sample position within the current block (i.e., the current CU) of 16×16 size (or 16×16 size or larger) and used.
[0182] However, the embodiments in the present document described above are merely one example, and the default motion information and the motion information of the current sub-block can be derived based not only on the center position (i.e., the right-bottom sample position) but also on other sample positions. For example, the default motion information can be derived based on the left-top sample position of the current CU, and the motion information of the current sub-block can also be derived based on the left-top sample position of the sub-block.
[0183] When implementing the embodiments in the present document as described above in hardware, since the same H / W module can be used to derive the temporal motion, a pipeline as shown in FIGS. 20 and 21 can be configured.
[0184] FIG. 20 and FIG. 21 are exemplary diagrams schematically showing the configuration of a pipeline capable of integrally calculating corresponding positions for deriving a default MV and a sub-block MV in the sbTMVP derivation process.
[0185] Referring to FIGS. 20 and 21, a corresponding position calculation module can calculate corresponding positions for deriving a default MV and a sub-block MV. For example, as shown in FIGS. 20 and 21, when the position (posX, posY) and block size (blkszX, blkszY) of a block are input to the corresponding position calculation module, the center position of the input block (i.e., the lower right sample position) can be output. When the position and block size of the current CU are input to the corresponding position calculation module, the center position of the col block on the col picture (i.e., the lower right sample position), which is the corresponding position for deriving the default MV, can be output. Alternatively, when the position and block size of the current sub-block are input to the corresponding position calculation module, the center position of the col sub-block on the col picture (i.e., the lower right sample position), which is the corresponding position for deriving the current sub-block MV, can be output.
[0186] Thus, when corresponding positions for deriving a default MV and a sub-block MV are output from the corresponding position calculation module, the motion vectors (i.e., temporal mv) derived from the above corresponding positions can be patched. Then, sub-block-based temporal motion information (i.e., sbTMVP candidate) can be induced based on the patched motion vectors (i.e., temporal mv). For example, as in FIGS. 20 and 21, depending on the implementation of the H / W, sbTMVP candidates can be derived in parallel based on clock cycles, or sbTMVP candidates can be derived sequentially.
[0187] The following drawings are created to illustrate a specific example of this document. Since the names of specific devices, specific terms, and names (such as the names of syntax / syntax elements, etc.) described in the drawings are presented exemplarily, the technical features of this document are not limited to the specific names used in the following drawings.
[0188] FIG. 22 schematically shows an example of a video / video encoding method according to an embodiment in this document.
[0189] The method disclosed in FIG. 22 can be executed by the encoding device 200 disclosed in FIG. 2. Specifically, steps (S2200) to (S2230) in FIG. 22 can be executed by the prediction unit 220 (more specifically, the inter-prediction unit 221) disclosed in FIG. 2, step (S2240) in FIG. 22 can be executed by the residual processing unit 230 disclosed in FIG. 2, and step (S2250) in FIG. 22 can be executed by the entropy encoding unit 240 disclosed in FIG. 2. Also, the method disclosed in FIG. 22 can be executed including the embodiments described above in this document. Therefore, in FIG. 22, for the content overlapping with the above-described embodiments, specific descriptions will be omitted or simplified.
[0190] Referring to FIG. 22, the encoding device can derive a reference sub-block on a collocated reference picture for a sub-block within the current block (S2200).
[0191] Here, the collocated reference picture refers to the reference picture used to derive the temporal motion information (i.e., sbTMVP) as described above, and can indicate the aforementioned col picture. The reference sub-block can indicate the aforementioned col sub-block.
[0192] In one embodiment, the encoding device can derive a reference sub-block on the collocated reference picture based on the position of the sub-blocks within the current block. Here, the current block may be referred to as the current coding unit (CU) or the current coding block (CB), and the sub-blocks included within the current block may also be referred to as the current coding sub-blocks.
[0193] For example, the encoding device can first identify the position of the current block and then identify the position of the sub-blocks within the current block. As described with reference to Table 2 above, the position of the current block can be indicated based on the upper-left sample position (xCtb, yCtb) of the coding tree block and the lower-right center sample position (xCtr, yCtr) of the current block. The position of the sub-blocks within the current block can be indicated by (xSb, ySb) respectively, and this position (xSb, ySb) can indicate the lower-right center sample position of the sub-block. Here, the lower-right center sample position (xSb, ySb) of the sub-block can be calculated based on the upper-left sample position of the sub-block and the sub-block size, and can be calculated as shown in formulas (8 - 523) to (8 - 524) of Table 2 above.
[0194] Then, the encoding device can derive a reference sub-block on the collocated reference picture based on the lower-right center sample position of each sub-block within the current block. As described with reference to Table 2 above, the reference sub-block can be indicated by the position (xColSb, yColSb) on the collocated reference picture, and the position (xColSb, yColSb) can be derived on the collocated reference picture based on the lower-right center sample position (xSb, ySb) of each sub-block within the current block.
[0195] On the one hand, the left-upper sample position used in this document may also be referred to as the upper-left sample position, the left upper-side sample position, etc., and the right-lower center sample position may also be referred to as the lower-right center sample position, the center lower-right sample position, the lower right-side center sample position, the center lower right-side sample position, etc.
[0196] Also, when deriving a reference sub-block, motion shift can be applied. The encoding device can execute motion shift based on the motion vector derived from the spatially adjacent block of the current block. The spatially adjacent block of the current block can be the left adjacent block located on the left side of the current block, and for example, it can be referred to as the A1 block shown in FIGS. 10 and 11. In this case, when the left adjacent block (for example, the A1 block) is available, a motion vector can be derived from the left adjacent block, or when the left adjacent block is not available, a zero vector can be derived. Here, whether the spatially adjacent block is available or not can be determined by the reference picture information, prediction mode information, position information, etc. of the spatially adjacent block. For example, when the reference picture of the spatially adjacent block is the same as that of the current block, the spatially adjacent block can be determined to be available. Alternatively, when the spatially adjacent block is coded in the intra prediction mode or the spatially adjacent block is located outside the current picture / tile, the spatially adjacent block can be determined to be unavailable.
[0197] That is, the encoding device applies a motion shift (i.e., the motion vector of a spatially adjacent block (e.g., the A1 block)) to the right-bottom center sample position (xSb, ySb) of each sub-block within the current block, and can derive a reference sub-block on the collocated reference picture based on the motion-shifted position. At this time, the position (xColSb, yColSb) of the reference sub-block can be indicated by the position obtained by motion-shifting the right-bottom center sample position (xSb, ySb) of each sub-block within the current block by the motion vector of the spatially adjacent block (e.g., the A1 block), and can be calculated as in the mathematical formulas (8-525) to (8-526) of Table 2 above.
[0198] The encoding device can derive a time motion information candidate based on sub-blocks based on the reference sub-block (S2210).
[0199] On the other hand, the time motion information candidate based on sub-blocks in this document is what is referred to as the aforementioned sbTMVp (subblock-based Temporal Motion Vector Prediction) candidate, and can be replaced or used in combination with the time motion vector predictor candidate based on sub-blocks. That is, as described above, when deriving motion information in units of sub-blocks and performing prediction, an sbTMVP candidate can be derived, and motion prediction can be performed at the sub-block level (or at the sub-coding unit (sub-CU) level) based on the sbTMVP candidate.
[0200] The encoding device can derive motion information for the sub-blocks within the current block based on the time motion information candidate based on sub-blocks (S2220).
[0201] The time motion information candidate based on sub-blocks can include a sub-block unit motion vector. At this time, the sub-block unit motion vector can include a motion vector derived based on the reference sub-block.
[0202] In one embodiment, the encoding device can derive the sub-block unit motion vector for the reference sub-block as motion information for the sub-blocks within the current block. For example, the encoding device can derive the sub-block unit motion vector based on whether the reference sub-block is available. For the available reference sub-blocks in the reference sub-block, the encoding device can derive the sub-block unit motion vector for the available reference sub-blocks based on the motion vectors of the available reference sub-blocks. For the unavailable reference sub-blocks in the reference sub-block, the encoding device can use the base motion vector as the sub-block unit motion vector for the unavailable reference sub-blocks.
[0203] The base motion vector can correspond to the aforementioned default motion vector and can be derived on the collocated reference picture based on the position of the current block.
[0204] In one embodiment, the encoding device can identify the position of the reference coding block on the collocated reference picture based on the right-bottom center sample position of the current block, and can derive the base motion vector based on the position of the reference coding block. The reference coding block can be referred to as the col block located on the collocated reference picture corresponding to the current block including sub-blocks. As described with reference to Table 2 above, the position of the reference coding block can be indicated by (xColCtrCb, yColCtrCb), and the position (xColCtrCb, yColCtrCb) can indicate the position of the reference coding block covering the position (xCtr, yCtr) within the collocated reference picture with reference to the left-top sample of the collocated reference picture. The position (xCtr, yCtr) can indicate the right-bottom center sample position of the current block.
[0205] Also, when deriving the base motion vector, a motion shift can be applied to the position (xColCtrCb, yColCtrCb) of the reference coding block. As described above, the motion shift can be performed by adding the motion vector derived from the spatially adjacent block (e.g., A1 block) of the current block to the reference coding block position (xColCtrCb, yColCtrCb) covering the lower right center sample. The encoding device can derive the base motion vector based on the position (xColCb, yColCb) of the motion-shifted reference coding block. That is, the base motion vector can be a motion vector derived from the position motion-shifted on the collocated reference picture based on the lower right center sample position of the current block.
[0206] On the other hand, whether or not the above reference sub-block is available can be determined based on whether it is located outside the collocated reference picture or based on the motion vector. For example, an unavailable reference sub-block can include a reference sub-block located outside (deviating from) the collocated reference picture or a reference sub-block with an unavailable motion vector. For example, when the reference sub-block is based on the intra mode, IBC (Intra Block Copy) mode, or palette mode, the above reference sub-block can be a sub-block with an unavailable motion vector. Alternatively, when the reference coding block covering the corrected position derived based on the position of the reference sub-block is based on the intra mode, IBC mode, or palette mode, the above reference sub-block can be a sub-block with an unavailable motion vector.
[0207] At this time, in one embodiment, the motion vector of the available reference sub-block can be derived based on the motion vector of the block covering the modified location derived based on the upper left corner sample position of the reference sub-block. For example, as shown in Table 2 above, the modified location can be derived based on a mathematical formula such as ((xColSb>>3)<<3, (yColSb>>3)<<3). Here, xColSb and yColSb respectively indicate the x coordinate and y coordinate of the upper left corner sample position of the reference sub-block, >> can indicate arithmetic right shift, and << can indicate arithmetic left shift.
[0208] On the other hand, as described above, when deriving the time motion information candidate based on the sub-block, it can be seen that the motion vector for the reference sub-block is derived based on the position of the sub-block within the current block, and the base motion vector is derived based on the position of the current block. For example, as described with reference to FIGS. 16 to 19, for a current block of 8×8 size, the motion vector for the reference sub-block and the base motion vector can be derived based on the lower right center sample position of the current block. For a current block having a size larger than 8×8, the motion vector for the reference sub-block is derived based on the lower right center sample position of each sub-block within the current block, and the base motion vector can be derived based on the lower right center sample position of the current block.
[0209] The encoding device can generate predicted samples of the current block based on the motion information for the sub-blocks within the current block (S2230).
[0210] The encoding device can select optimal motion information based on the RD (Rate-Distortion) cost and generate a prediction sample based on this. For example, when the motion information derived for each sub-block of the current block (i.e., sbTMVP) is selected as the optimal motion information, the encoding device can generate a prediction sample for the current block based on the motion information for the sub-blocks of the current block derived as described above.
[0211] The encoding device can derive a residual sample based on the prediction sample (S2240) and encode video information including information about the residual sample (S2250).
[0212] That is, the encoding device can derive a residual sample based on the original sample for the current block and the prediction sample of the current block. And the encoding device can generate information about the residual sample. Here, the information about the residual sample can include value information of quantized transform coefficients, position information, transform technique, transform kernel, quantization parameter, etc. derived by performing conversion and quantization on the residual sample.
[0213] The encoding device can encode the information about the residual sample and output it as a bitstream, and can transmit this to the decoding device via a network or a storage medium.
[0214] FIG. 23 schematically shows an example of a video / video decoding method according to an embodiment in this document.
[0215] The method disclosed in FIG. 23 can be executed by the decoding device 300 disclosed in FIG. 3. Specifically, steps (S2300) to (S2330) in FIG. 23 can be executed by the prediction unit 330 (more specifically, the inter prediction unit 332) disclosed in FIG. 3, and step (S2340) in FIG. 23 can be executed by the addition unit 340 disclosed in FIG. 3. Also, the method disclosed in FIG. 23 can be executed including the embodiments described above in this document. Therefore, in FIG. 23, specific descriptions of the content overlapping with the above-described embodiments will be omitted or simplified.
[0216] Referring to FIG. 23, the decoding device can derive a reference sub-block on a collocated reference picture for a sub-block within the current block (S2300).
[0217] Here, the collocated reference picture refers to the reference picture used to derive the temporal motion information (i.e., sbTMVP) as described above, and can indicate the aforementioned col picture. The reference sub-block can indicate the aforementioned col sub-block.
[0218] In one embodiment, the decoding device can derive a reference sub-block on the collocated reference picture based on the position of the sub-block within the current block. Here, the current block can be referred to as the current coding unit (CU) or the current coding block (CB), and the sub-block included within the current block can also be referred to as the current coding sub-block.
[0219] For example, the decoding device can first identify the position of the current block and then identify the position of the sub-blocks within the current block. As described with reference to Table 2 above, the position of the current block can be indicated based on the upper left sample position (xCtb, yCtb) of the coding tree block and the center sample position (xCtr, yCtr) of the lower right side of the current block. The positions of the sub-blocks within the current block can be indicated by (xSb, ySb) respectively, and this position (xSb, ySb) can indicate the center sample position of the lower right side of the sub-block. Here, the center sample position (xSb, ySb) of the lower right side of the sub-block can be calculated based on the upper left sample position of the sub-block and the sub-block size, and can be calculated as in the mathematical formulas (8-523) to (8-524) of Table 2 above.
[0220] Then, the decoding device can derive the reference sub-blocks on the collocated reference picture based on the center sample positions of the lower right sides of the respective sub-blocks within the current block. As described with reference to Table 2 above, the reference sub-blocks can be indicated by the position (xColSb, yColSb) on the collocated reference picture, and the position (xColSb, yColSb) can be derived on the collocated reference picture based on the center sample positions (xSb, ySb) of the lower right sides of the respective sub-blocks within the current block.
[0221] On the other hand, the upper left sample position used in this document may also be referred to as the upper left sample position, the upper left side sample position, etc., and the center sample position of the lower right side may also be referred to as the center sample position of the lower right side, the center lower right side sample position, the lower right side center sample position, the center lower right side sample position, etc.
[0222] Also, when deriving a reference sub-block, a motion shift can be applied. The decoding device can execute a motion shift based on a motion vector derived from a spatially adjacent block of the current block. The spatially adjacent block of the current block can be a left adjacent block located on the left side of the current block, and can be referred to as, for example, the A1 block shown in FIGS. 10 and 11. In this case, when the left adjacent block (e.g., the A1 block) is available, a motion vector can be derived from the left adjacent block, or when the left adjacent block is not available, a zero vector can be derived. Here, whether the spatially adjacent block is available can be determined based on reference picture information, prediction mode information, position information, etc. of the spatially adjacent block. For example, when the reference picture of the spatially adjacent block is the same as the reference picture of the current block, the spatially adjacent block can be determined to be available. Alternatively, when the spatially adjacent block is coded in the intra prediction mode or the spatially adjacent block is located outside the current picture / tile, the spatially adjacent block can be determined to be unavailable.
[0223] That is, the decoding device applies a motion shift (i.e., the motion vector of the spatially adjacent block (e.g., the A1 block)) to the right lower center sample position (xSb, ySb) of each sub-block in the current block, and based on the motion-shifted position, a reference sub-block can be derived on the collocated reference picture. At this time, the position (xColSb, yColSb) of the reference sub-block can be indicated by the position where the motion vector of the spatially adjacent block (e.g., the A1 block) is motion-shifted to the right lower center sample position (xSb, ySb) of each sub-block in the current block, and can be calculated as in the mathematical formulas (8-525) to (8-526) of Table 2 above.
[0224] The decoding device can derive a time motion information candidate based on the sub-block based on the reference sub-block (S2310).
[0225] On the one hand, in this document, the time motion information candidate based on sub-blocks is what is referred to as the aforementioned sbTMVp (Subblock-Based Temporal Motion Vector Prediction) candidate, and can be replaced or used in combination with the time motion vector predictor candidate based on sub-blocks. That is, as described above, when deriving motion information in sub-block units and performing prediction, sbTMVP candidates can be derived, and motion prediction can be performed at the sub-block level (or at the sub-coding unit (sub-CU) level) based on the above sbTMVP candidates.
[0226] The decoding device can derive motion information for the sub-blocks within the current block based on the time motion information candidate based on sub-blocks (S2320).
[0227] The time motion information candidate based on sub-blocks can include sub-block unit motion vectors. At this time, the sub-block unit motion vector can include the motion vector derived based on the reference sub-block.
[0228] In one embodiment, the decoding device can derive the sub-block unit motion vector for the reference sub-block as the motion information for the sub-blocks within the current block. For example, the decoding device can derive the sub-block unit motion vector based on whether the reference sub-block is available. For the available reference sub-blocks in the reference sub-block, the decoding device can derive the sub-block unit motion vector for the available reference sub-blocks based on the motion vectors of the available reference sub-blocks. For the unavailable reference sub-blocks in the reference sub-block, the decoding device can use the base motion vector as the sub-block unit motion vector for the unavailable reference sub-blocks.
[0229] The base motion vector can correspond to the aforementioned default motion vector and can be derived on the collocated reference picture based on the position of the current block.
[0230] In one embodiment, the decoding device can identify the position of the reference coding block on the collocated reference picture based on the right-bottom center sample position of the current block, and can derive the base motion vector based on the position of the reference coding block. The reference coding block may refer to a col block located on the collocated reference picture corresponding to the current block including sub-blocks. As described with reference to Table 2 above, the position of the reference coding block can be indicated by (xColCtrCb, yColCtrCb), and the position (xColCtrCb, yColCtrCb) can indicate the position of the reference coding block covering the position (xCtr, yCtr) within the collocated reference picture with reference to the left-top sample of the collocated reference picture. The position (xCtr, yCtr) can indicate the right-bottom center sample position of the current block.
[0231] Also, when deriving the base motion vector, a motion shift can be applied to the position (xColCtrCb, yColCtrCb) of the reference coding block. The motion shift can be executed by adding the motion vector derived from the spatially adjacent block (e.g., A1 block) of the current block to the reference coding block position (xColCtrCb, yColCtrCb) covering the right-bottom center sample. The decoding device can derive the base motion vector based on the position (xColCb, yColCb) of the motion-shifted reference coding block. That is, the base motion vector can be a motion vector derived from the motion-shifted position on the collocated reference picture based on the right-bottom center sample position of the current block.
[0232] On the other hand, whether the reference sub-block is available can be determined based on whether it is located outside the collocated reference picture or based on the motion vector. For example, an unavailable reference sub-block can include a reference sub-block that deviates outside the collocated reference picture or a reference sub-block for which the motion vector is unavailable. For example, when the reference sub-block is based on the intra mode, IBC (Intra Block Copy) mode, or palette mode, the reference sub-block can be a sub-block for which the motion vector is unavailable. Alternatively, when the reference coding block covering the modified position derived based on the position of the reference sub-block is based on the intra mode, IBC mode, or palette mode, the reference sub-block can be a sub-block for which the motion vector is unavailable.
[0233] At this time, in one embodiment, the motion vector of the available reference sub-block can be derived based on the motion vector of the block covering the modified location derived based on the top-left sample position of the reference sub-block. For example, as shown in Table 2 above, the modified location can be derived based on a formula such as ((xColSb >> 3) << 3, (yColSb >> 3) << 3). Here, xColSb and yColSb respectively indicate the x-coordinate and y-coordinate of the top-left sample position of the reference sub-block, >> can indicate arithmetic right shift, and << can indicate arithmetic left shift.
[0234] On the one hand, as described above, when deriving the time motion information candidates based on sub-blocks, it can be seen that the motion vector for the reference sub-block is derived based on the position of the sub-block within the current block, and the base motion vector is derived based on the position of the current block. For example, as described with reference to FIGS. 16 to 19, for a current block of 8×8 size, the motion vector for the reference sub-block and the base motion vector can be derived based on the right-bottom center sample position of the current block. For a current block having a size larger than 8×8, the motion vector for the reference sub-block is derived based on the right-bottom center sample position of each sub-block within the current block, and the base motion vector can be derived based on the right-bottom center sample position of the current block.
[0235] The decoding device can generate predicted samples of the current block based on the motion information for the sub-blocks within the current block (S2330).
[0236] In one embodiment, for a prediction mode in which the decoding device performs prediction based on sub-block unit motion information (i.e., sbTMVP mode) for the current block, the decoding device can generate predicted samples of the current block based on the motion information for the sub-blocks of the current block derived as described above.
[0237] The decoding device can generate restored samples based on the predicted samples (S2340).
[0238] In one embodiment, the decoding device can immediately use the predicted samples as restored samples in the prediction mode, or can also generate restored samples by adding residual samples to the predicted samples.
[0239] When there are residual samples for the current block, the decoding device can receive information regarding the residual for the current block. The information regarding the residual can include transform coefficients regarding the residual samples. The decoding device can derive residual samples (or, a residual sample array) for the current block based on the residual information. The decoding device can generate restored samples based on the predicted samples and the residual samples, and can derive a restored block or a restored picture based on the above restored samples. Thereafter, the decoding device can, if necessary, apply in-loop filtering procedures such as deblock filtering and / or SAO procedures to the above restored picture in order to improve subjective / objective picture quality, as described above.
[0240] In the foregoing embodiments, the method has been described based on a flowchart in a series of steps or blocks, but the embodiments of this document are not limited to the order of the steps, and a certain step may occur in a different order or simultaneously with steps different from the foregoing. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, other steps may be included, or one or more steps of the flowchart can be deleted without affecting the scope of this document.
[0241] The method according to the foregoing document can be embodied in software form, and the encoding device and / or decoding device according to this document can be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, a display device, and the like.
[0242] When the embodiments in this document are implemented in software, the methods described above can be implemented by modules (processes, functions, etc.) that perform the functions described above. The modules can be stored in memory and executed by a processor. The memory can be internal or external to the processor and can be connected to the processor by various well-known means. The processor can include an ASIC (Application-Specific Integrated Circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (Read-Only Memory), a RAM (Random Access Memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or algorithms can be stored in a digital storage medium.
[0243] In addition, the decoding device and encoding device to which this document is applicable can be included in multimedia broadcast transmission / reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conferencing devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providing devices, OTT video (Over The Top video) devices, Internet streaming service providing devices, three-dimensional (3D) video devices, VR (Virtual Reality) devices, AR (Augmented Reality) devices, picture phone video devices, transportation means terminals (e.g., vehicles (including autonomous driving vehicles), airplane terminals, ship terminals, etc.), and medical video devices, etc., and can be used to process video signals or data signals. For example, OTT video (Over The Top video) devices can include game consoles, Blu-ray players, Internet-connected TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorder), etc.
[0244] In addition, the processing method to which this document is applicable can be produced in the form of a program executed by a computer and can be stored in a computer-readable storage (recording) medium. Also, multimedia data having a data structure according to this document can be stored in a computer-readable storage medium. The above-mentioned computer-readable storage medium includes all types of storage devices and distributed storage devices in which data that can be read by a computer is stored. The above-mentioned computer-readable storage medium can include, for example, Blu-ray Disc (BD), Universal Serial (general-purpose serial) Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Also, the above-mentioned computer-readable storage medium includes a medium embodied in the form of a carrier wave (for example, transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a computer-readable storage medium or transmitted via a wired or wireless communication network.
[0245] In addition, the embodiments of this document can be embodied in a computer program product by program code, and the above program code can be executed by a computer according to the embodiments of this document. The above program code can be stored on a carrier readable by a computer.
[0246] FIG. 24 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.
[0247] Referring to FIG. 24, the content streaming system to which the embodiments of this document are applied can generally include an encoding server, a streaming server, a web server, a media storage device (repository), a user device, and a multimedia input device.
[0248] The above encoding server compresses the content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and plays the role of sending this to the above streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate a bitstream, the above encoding server can be omitted.
[0249] The above bitstream can be generated by an encoding method or a bitstream generation method applied to the embodiments of this document, and the above streaming server can temporarily store the above bitstream in the process of sending or receiving the above bitstream.
[0250] The above streaming server sends multimedia data to a user device based on a user request via a web server, and the above web server plays the role of a medium for informing the user of what services are available. When the user requests a desired service from the above web server, the above web server transmits this to the streaming server, and the above streaming server sends multimedia data to the user. At this time, the above content streaming system can include a separate control server. In this case, the above control server plays the role of controlling commands / responses between each device within the above content streaming system.
[0251] The above streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from the above encoding server, the above content can be received in real time. In this case, in order to provide a smooth streaming service, the above streaming server can store the above bitstream for a certain period of time.
[0252] Examples of the above user devices include mobile phones, smart phones, laptop computers, digital broadcast terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), navigation devices, slate PCs, tablet PCs, ULTRABOOKs (registered trademark), wearable devices (e.g., smartwatches (watch-type terminals), smart glasses (glass-type terminals), HMDs (Head Mounted Displays), digital TVs, desktop computers, and digital signatures (signid).
[0253] Each server in the above content streaming system can be operated as a distributed server. In this case, the data received by each server can be distributedly processed.
[0254] The claims described in this document can be combined in various ways. For example, the technical features of the method claims in this document can be combined and implemented in a device, and the technical features of the device claims in this document can be combined and implemented in a method. Also, the technical features of the method claims in this document and the technical features of the device claims can be combined and implemented in a device, and the technical features of the method claims in this document and the technical features of the device claims can be combined and implemented in a method.
Claims
1. A video decoding method executed by a decoding device, comprising: deriving a collocated sub-block in a collocated picture for a sub-block in a current block; deriving a temporal motion information candidate based on the sub-block from the collocated sub-block; deriving motion information for the sub-block in the current block based on the temporal motion information candidate based on the sub-block; generating a prediction sample for the current block based on the motion information for the sub-block in the current block; generating a restored sample based on the prediction sample; applying deblocking filtering to the restored sample, wherein based on the collocated sub-block being available, each position of the collocated sub-blocks in the collocated picture is derived based on the position of the sub-blocks in the current block; each of the positions of the collocated sub-blocks in the collocated picture is derived by applying a motion shift to the right-bottom center sample position of each of the sub-blocks in the current block; the temporal motion information candidate based on the sub-block includes a sub-block unit motion vector; a base motion vector is used as the sub-block unit motion vector for collocated sub-blocks that are not available among the collocated sub-blocks.
2. A video encoding method executed by an encoding device, comprising: deriving a collocated sub-block in a collocated picture for a sub-block in a current block; deriving a temporal motion information candidate based on the sub-block from the collocated sub-block; deriving motion information for the sub-block in the current block based on the temporal motion information candidate based on the sub-block; generating a prediction sample for the current block based on the motion information for the sub-block in the current block; deriving a residual sample based on the prediction sample; encoding video information including information about the residual sample. Based on the collocate sub-block being available, the respective positions of the collocate sub-blocks in the collocate picture are derived based on the positions of the sub-blocks in the current block, the respective positions of the collocate sub-blocks in the collocate picture are derived by applying a motion shift to the right lower end center sample position of each of the sub-blocks in the current block, the time motion information candidate based on the sub-block includes a sub-block unit motion vector, A method in which a base motion vector is used as a sub-block unit motion vector for a collocate sub-block that is not available among the collocate sub-blocks.
3. A method for transmitting data related to video, a step of obtaining a bitstream related to the video, where the bitstream a step of deriving collocate sub-blocks in a collocate picture for sub-blocks in a current block, a step of deriving a time motion information candidate based on the sub-block based on the collocate sub-block, a step of deriving motion information for the sub-blocks in the current block based on the time motion information candidate based on the sub-block, a step of generating predicted samples of the current block based on the motion information for the sub-blocks in the current block, a step of deriving residual samples based on the predicted samples, a step of encoding video information including information related to the residual samples, and a step generated based on a step of transmitting the data including the bitstream, Based on the collocate sub-block being available, the respective positions of the collocate sub-blocks in the collocate picture are derived based on the positions of the sub-blocks in the current block, the respective positions of the collocate sub-blocks in the collocate picture are derived by applying a motion shift to the right lower end center sample position of each of the sub-blocks in the current block, the time motion information candidate based on the sub-block includes a sub-block unit motion vector, A method in which a base motion vector is used as a sub-block unit motion vector for a collocate sub-block that is unusable among the collocate sub-blocks.
Citation Information
Patent Citations
Moving image decoding device
WO2017195608A1
Method for processing video on basis of inter prediction mode and apparatus therefor
WO2018066927A1
System and method for signaling of motion merge modes in video coding
WO2020142448A1