Temporal motion vector predictor candidate-based image or video coding of sub block unit
By deriving sub-block-based temporal motion vectors using reference sub-blocks on a co-located reference picture, the method addresses inefficiencies in high-resolution image/video coding, enhancing compression efficiency and reducing complexity.
Patent Information
- Application Number
- JP2025116627
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-06-13
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2040-06-15
AI Technical Summary
Existing video/image coding technologies face inefficiencies in handling high-resolution and high-quality images/videos, particularly in immersive media, leading to increased transmission and storage costs, and require improved methods for sub-block-based motion vector prediction to enhance coding efficiency.
The method involves deriving sub-block-based temporal motion vectors using reference sub-blocks on a co-located reference picture, utilizing center sample positions and base motion vectors to improve prediction performance, and integrating corresponding positions at the sub-coding and coding block levels.
This approach enhances overall image/video compression efficiency, reduces computational complexity, and simplifies hardware implementation by efficiently calculating corresponding positions for sub-block-based temporal motion vectors, thereby improving coding efficiency.
Smart Images

Figure 2025137602000001_ABST
Abstract
Description
[Technical Field]
[0001] The present technology relates to video or image coding, for example, to image or image coding techniques based on sub-block-based temporal motion vector predictor candidates. [Background technology]
[0002] Recently, demand for high-resolution, high-quality images / videos, such as 4K or 8K or higher UHD (Ultra High Definition) images / videos, is increasing in various fields. As the resolution and quality of image / video data increases, the amount of information or bits transmitted increases relative to existing image / video data. Therefore, when transmitting image data using existing media such as wired or wireless broadband lines, or storing image / video data using existing storage media, transmission costs and storage costs increase.
[0003] In addition, interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content and holograms has been increasing recently, and the broadcast of images / videos with visual characteristics different from real images, such as game images, is increasing.
[0004] Therefore, a highly efficient image / video compression technology is required to effectively compress, transmit, store, and play back high-resolution, high-quality image / video information having the above-mentioned various characteristics.
[0005] In addition, there is a discussion about sub-block-based temporal motion vector prediction technology to improve image / video coding efficiency. For this, a method is needed to efficiently perform the process of patching sub-block-based motion vectors in sub-block-based temporal motion vector prediction. Summary of the Invention [Problem to be solved by the invention]
[0006] The technical problem of this document is to provide a method and apparatus for increasing the efficiency of video / image coding.
[0007] Another technical problem of this document is to provide an efficient inter-prediction method and apparatus.
[0008] It is still another technical object of the present invention to provide a method and apparatus for deriving sub-block-based temporal motion vectors to improve prediction performance.
[0009] It is still another technical object of the present invention to provide a method and apparatus for efficiently deriving corresponding positions of sub-blocks for deriving (guiding) a sub-block-based temporal motion vector.
[0010] It is still another technical object of the present invention to provide a method and apparatus for integrating corresponding positions at the sub-coding block level and corresponding positions at the coding block level to derive sub-block-based temporal motion vectors. [Means for solving the problem]
[0011] According to one embodiment of this document, in subblock-based Temporal Motion Vector Prediction (sbTMVP), a subblock-wise motion vector for a current block can be derived based on a reference subblock on (within) a collocated reference picture.
[0012] According to one embodiment of this document, reference sub-blocks on a co-located reference picture for multiple sub-blocks in a current block can be derived based on the center sample positions of each of the sub-blocks in the current block.
[0013] According to one embodiment of this document, a base motion vector can be used for a sub-block-wise motion vector for an unavailable reference sub-block in a reference sub-block.
[0014] According to one embodiment of this document, a base motion vector can be derived on (from) a co-located reference picture based on the center sample position of the current block.
[0015] According to an embodiment of the present document, there is provided a video / image decoding method executed by a decoding device, which may include the methods disclosed in the embodiments of the present document.
[0016] According to an embodiment of the present document, there is provided a decoding device for performing video / image decoding, which is capable of performing the methods disclosed in the embodiments of the present document.
[0017] According to an embodiment of the present document, there is provided a video / image encoding method executed by an encoding device, which may include the methods disclosed in the embodiments of the present document.
[0018] According to an embodiment of the present document, there is provided an encoding device for performing video / image encoding, wherein the encoding device is capable of performing the method disclosed in the embodiment of the present document.
[0019] According to one embodiment of the present document, there is provided a computer-readable digital storage medium having encoded video / image information stored thereon, the encoded video / image information being generated according to the video / image encoding method disclosed in at least one of the embodiments of the present document.
[0020] According to one embodiment of the present document, there is provided a computer-readable digital storage medium having stored thereon encoded information or encoded video / image information that causes a decoding device to perform the video / image decoding method disclosed in at least one of the embodiments of the present document. [Effects of the Invention]
[0021] The present disclosure may achieve various advantages. For example, it may improve overall image / video compression efficiency. It may also reduce computational complexity through efficient inter-prediction, thereby improving overall coding efficiency. It may also improve efficiency in terms of complexity and prediction performance by efficiently calculating corresponding positions of sub-blocks for deriving sub-block-based temporal motion vectors in sub-block-based temporal motion vector prediction (sbTMVP). It may also simplify hardware implementation by integrating methods for calculating corresponding positions at the sub-coding block level and the coding block level for deriving sub-block-based temporal motion vectors.
[0022] The effects obtained through the specific embodiments of this document are not limited to the effects listed above. For example, there may be various technical effects that a person having ordinary skill in the related art can understand or derive from this document. Therefore, the specific effects of this document are not limited to those explicitly described in this document, but may include various effects that can be understood or derive from the technical features of this document. [Brief explanation of the drawings]
[0023] [Figure 1] FIG. 1 illustrates a schematic diagram of an example video / image coding system to which embodiments of the present document can be applied. [Figure 2] 1 is a diagram illustrating the configuration of a video / image encoding device to which an embodiment of the present document can be applied; [Figure 3] 1 is a diagram illustrating the configuration of a video / image decoding device to which an embodiment of the present document can be applied. [Figure 4] 1 illustrates an example of a general video / image encoding method to which embodiments of the present document may be applied. [Figure 5] FIG. 1 illustrates an example of a general video / image decoding method to which embodiments of the present document may be applied. [Figure 6] 1 illustrates an example of a general inter-prediction based video / image encoding method to which embodiments of the present document can be applied. [Figure 7] 1 illustrates an example of a general inter-prediction based video / image decoding method to which embodiments of the present document can be applied. [Figure 8] FIG. 10 is a diagram illustrating an example of an inter-prediction procedure. [Figure 9] FIG. 2 is a diagram illustrating exemplary spatial and temporal neighboring blocks of a current block. [Figure 10] FIG. 10 is a diagram illustrating exemplary spatially adjacent blocks that can be used to derive sub-block-based temporal motion information candidates (sbTMVP candidates). [Figure 11] FIG. 1 is a diagram for explaining the process of deriving sub-block-based temporal motion information candidates (sbTMVP candidates). [Figure 12] 10 is a diagram illustrating a method for calculating corresponding positions for deriving default MVs and sub-block MVs according to block size during the sbTMVP derivation process. [Figure 13] 10 is a diagram illustrating a method for calculating corresponding positions for deriving default MVs and sub-block MVs according to block size during the sbTMVP derivation process. [Figure 14]10 is a diagram illustrating a method for calculating corresponding positions for deriving default MVs and sub-block MVs according to block size during the sbTMVP derivation process. [Figure 15] 10 is a diagram illustrating a method for calculating corresponding positions for deriving default MVs and sub-block MVs according to block size during the sbTMVP derivation process. [Figure 16] FIG. 10 is an exemplary diagram illustrating a method of integrating corresponding positions for deriving default MVs and sub-block MVs according to block size in the process of deriving sbTMVP. [Figure 17] FIG. 10 is an exemplary diagram illustrating a method of integrating corresponding positions for deriving default MVs and sub-block MVs according to block size in the process of deriving sbTMVP. [Figure 18] FIG. 10 is an exemplary diagram illustrating a method of integrating corresponding positions for deriving default MVs and sub-block MVs according to block size in the process of deriving sbTMVP. [Figure 19] FIG. 10 is an exemplary diagram illustrating a method of integrating corresponding positions for deriving default MVs and sub-block MVs according to block size in the process of deriving sbTMVP. [Figure 20] 10 is an exemplary diagram illustrating a pipeline configuration that can integrate and calculate corresponding positions for deriving default MVs and sub-block MVs in the sbTMVP derivation process. [Figure 21] 10 is an exemplary diagram illustrating a pipeline configuration that can integrate and calculate corresponding positions for deriving default MVs and sub-block MVs in the sbTMVP derivation process. [Figure 22]1 is a diagram illustrating a schematic diagram of an example video / image encoding method according to an embodiment of the present document; [Figure 23] 1 is a diagram illustrating a schematic diagram of an example of a video / image decoding method according to an embodiment of the present document; [Figure 24] FIG. 1 illustrates an example of a content streaming system to which embodiments disclosed herein can be applied. DETAILED DESCRIPTION OF THE INVENTION
[0024] This document may be modified in various ways and may have various embodiments. Specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to the specific embodiments. Common terms used in this document are used merely to describe specific embodiments and are not intended to limit the technical ideas of this document. Singular expressions include plural expressions unless the context clearly dictates otherwise. In this document, terms such as "comprise" or "have" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0025] Meanwhile, each component in the drawings described in this document is illustrated independently for the convenience of explaining the different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Implementations in which each component is integrated and / or separated are also within the scope of this document as long as they do not deviate from the essence of this document.
[0026] In this document, "A or B" can mean "only A," "only B," or "both A and B." Also, in this document, "A or B" can be interpreted as "A and / or B." For example, in this document, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B, and C."
[0027] A slash ( / ) or a comma (comma) used in this document can mean "and / or." For example, "A / B" can mean "A and / or B." Therefore, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0028] In this document, "at least one of A and B" can mean "only A," "only B," or "both A and B." Also, in this document, the expressions "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B."
[0029] Also, in this document, "at least one of A, B and C" can mean "only A," "only B," "only C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" can mean "at least one of A, B and C."
[0030] Furthermore, parentheses used in this document may mean "for example." Specifically, when "prediction (intra prediction)" is used, "intra prediction" is proposed as an example of "prediction." In other words, "prediction" in this document is not limited to "intra prediction," and "intra prediction" is proposed as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is used, "intra prediction" is proposed as an example of "prediction."
[0031] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document may be applied to methods disclosed in the Versatile Video Coding (VVC) standard. The methods / embodiments disclosed in this document may also be applied to methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the 2nd Generation Of Audio Video Coding Standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.267).
[0032] This document presents various embodiments relating to video / image coding, and unless otherwise stated, the embodiments may also be implemented in combination with each other.
[0033] In this document, video may refer to a collection of a series of images over time. A picture generally refers to a unit that represents one image at a specific time period, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may include one or more Coding Tree Units (CTUs). A picture may consist of one or more slices / tiles. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs, and the rectangular region has the same height as the picture, and the width can be specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture.A tile scan may indicate a specific sequential ordering of CTUs partitioning a picture, in which the CTUs are ordered consecutively in a CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice includes an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that may be exclusively contained in a single NAL unit.
[0034] On the other hand, a picture can be divided into two or more sub-pictures, each of which is a rectangular region of one or more slices within a picture.
[0035] A pixel or a pel may refer to the smallest unit constituting a picture (or image). A "sample" may also be used as a term corresponding to a pixel. A sample may generally refer to a pixel or a pixel value, and may refer to only a pixel / pixel value of a luma component, or may refer to only a pixel / pixel value of a chroma component. Alternatively, a sample may refer to a pixel value in the spatial domain, or, when such a pixel value is transformed into the frequency domain, may refer to a transform coefficient in the frequency domain.
[0036] A unit may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area. In general, an M×N block may include samples (or a sample array) consisting of M columns and N rows, or a set (or array) of transform coefficients.
[0037] Also, in this document, at least one of quantization / dequantization and / or transform / inverse transform can be omitted. When quantization / dequantization is omitted, the quantized transform coefficients can be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients can also be referred to as coefficients or residual coefficients, or can still be referred to as transform coefficients for the sake of uniformity of expression.
[0038] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficient(s), and the information about the transform coefficient(s) may be signaled via a residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficient(s)), and scaled transform coefficients may be derived through an inverse transform (scaling) on the transform coefficients. Residual samples may be derived based on an inverse transform (transform) on the scaled transform coefficients. This may be similarly applied / expressed in other parts of this document.
[0039] In this document, technical features described separately in one drawing may be embodied separately or simultaneously.
[0040] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Hereinafter, the same reference numerals will be used for the same components in the drawings, and duplicated descriptions of the same components may be omitted.
[0041] FIG. 1 illustrates schematically an example of a video / image coding system that can be applied to embodiments of this document.
[0042] 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or a network.
[0043] The source device may include a video source, an encoding device, and a sending unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / video encoding device, and the decoding device may be referred to as a video / video decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may also include a display unit, which may be a separate device or an external component.
[0044] A video source can acquire video / images through a video / image capture, synthesis, or generation process. A video source can include a video / image capture device and / or a video / image generation device. A video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated via a computer, etc., in which case the video / image capture process can be replaced by a process in which related data is generated.
[0045] An encoding device can encode input video / images. The encoding device can perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0046] The transmitter can transmit the encoded video / image information or data output in the form of a bitstream to a receiver in the receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter can include elements for generating a media file in a predetermined file format and elements for transmission via a broadcasting / communication network. The receiver can receive / extract the bitstream and transmit it to a decoding device.
[0047] The decoding device can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, and prediction that correspond to the operations of the encoding device.
[0048] The renderer can render the decoded video / image, and the rendered video / image can be displayed via a display unit.
[0049] 2 is a diagram illustrating a configuration of a video / image encoding device to which an embodiment of this document can be applied. Hereinafter, the encoding device may include an image encoding device and / or a video encoding device.
[0050] Referring to FIG. 2, the encoding apparatus 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter predictor 221 and an intra predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image dividing unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. Furthermore, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0051] The image division unit 210 may divide an input image (or picture or frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided into coding tree units (CTUs) or largest coding units (LCUs) using a quad-tree, binary-tree, and ternary-tree (QTBTTT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quad-tree structure may be applied first, and then the binary tree structure and / or the ternary tree structure may be applied. Alternatively, the binary tree structure may be applied first. The coding procedure according to this document may be performed based on the final coding unit that is not further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as needed, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be divided or partitioned from the final coding unit.The prediction unit is a unit of sample prediction, and the transform unit is a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0052] The term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block can refer to a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally refer to a pixel or pixel value, or can refer to only a pixel / pixel value of a luma component, or only a pixel / pixel value of a chroma component. A sample can also be used as a term corresponding to a pixel or pel of one picture (or image).
[0053] The encoding apparatus 200 may generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input video signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, a unit in the encoder 200 that subtracts the prediction signal (predicted block, prediction sample array) from the input video signal (original block, original sample array) may be referred to as the subtraction unit 231. The prediction unit may perform prediction on a current block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is to be applied on a current block or CU basis. The prediction unit may generate various information related to prediction, such as prediction mode information, and transmit the information to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The prediction information can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0054] The intra prediction unit 222 may predict the current block by referring to samples in the current picture. The referenced samples may be located adjacent to or distant from the current block depending on the prediction mode. Prediction modes in intra prediction may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0055] The inter prediction unit 221 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (e.g., L0 prediction, L1 prediction, or Bi prediction). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), or the like, and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 221 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of a skip mode or a merge mode, the inter predictor 221 may use motion information of neighboring blocks as motion information for the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the case of the Motion Vector Prediction (MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0056] The prediction unit 220 may generate a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The prediction unit may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / moving picture coding, such as games, such as Screen Content Coding (SCC). IBC basically performs prediction within a current picture, but may be performed similarly to inter prediction in deriving a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described herein. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture may be signaled based on information related to a palette table and a palette index.
[0057] The prediction signal generated by the prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) may be used to generate a reconstructed signal or a residual signal. The transform unit 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph representing inter-pixel relationship information. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size or to blocks of various (variable) sizes that are not square.
[0058] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy coding unit 240. The entropy coding unit 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 233 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy coding unit 240 may perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy coding unit 240 may encode information required for video / image restoration (e.g., values of syntax elements) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / video information) may be transmitted or stored in the form of a bitstream in Network Abstraction Layer (NAL) units. The video / video information may further include information on various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). The video / video information may also include general constraint information. In this document, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / video information. The video / video information may be encoded through the above-described encoding procedure and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoder 240 can be configured as an internal / external element of the encoding device 200, or the transmitter can be included in the entropy encoder 240.
[0059] The quantized transform coefficients output from the quantizer 233 may be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) may be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantizer 234 and the inverse transformer 235. The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter predictor 221 or the intra predictor 222. When there is no residual for the current block, such as when a skip mode is applied, a predicted block may be used as the reconstructed block. The adder 250 may be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next current block in the current picture, or may be used for inter prediction of the next picture after filtering, as described below.
[0060] Meanwhile, Luma Mapping with Chroma Scaling (LMCS) can be applied during picture encoding and / or restoration.
[0061] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods may include, for example, deblock filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 260 may generate various information related to filtering and transmit it to the entropy encoder 240, as will be described later in connection with each filtering method. The filtering information may be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0062] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding apparatus can avoid prediction mismatch between the encoding apparatus 100 and the decoding apparatus, and can also improve encoding efficiency.
[0063] The DPB of the memory 270 may store the modified reconstructed picture to be used as a reference picture in the inter predictor 221. The memory 270 may store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 221 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 270 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 222.
[0064] 3 is a diagram illustrating the configuration of a video / image decoding device to which the embodiments of this document can be applied. Hereinafter, the decoding device may include an image decoding device and / or a video decoding device.
[0065] Referring to FIG. 3, the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter predictor 331 and an intra predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. Depending on the embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured as a single hardware component (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) or may be configured as a digital storage medium. The above hardware components may further include a memory 360 as an internal / external component.
[0066] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image corresponding to the process by which the video / image information was processed by the encoding apparatus of FIG. 2. For example, the decoding apparatus 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding apparatus 300 can perform decoding using processing units applied by the encoding apparatus. Accordingly, the processing unit for decoding is, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximal coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding apparatus 300 can be played back via a playback device.
[0067] The decoding device 300 may receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal may be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 may parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The video / video information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / video information may also include general constraint information. The decoding device may decode pictures based on the information on the parameter sets and / or the general constraint information. Signaled / received information and / or syntax elements, which will be described later in this document, may be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using syntax element information (and neighboring elements) to be decoded, decoding information of the block to be decoded, or symbol / bin information decoded in a previous step, predicts the occurrence probability of bins according to the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element.In this case, after determining a context model, the CABAC entropy decoding method may update the context model using information on the decoded symbol / bin for the context model of the next symbol / bin. Prediction-related information from the information decoded by the entropy decoder 310 may be provided to a prediction unit (inter predictor 332 and intra predictor 331), and residual values entropy-decoded by the entropy decoder 310, i.e., quantized transform coefficients and related parameter information, may be input to the residual processor 320. The residual processor 320 may derive a residual signal (residual block, residual sample, residual sample array). Furthermore, filtering-related information from the information decoded by the entropy decoder 310 may be provided to the filtering unit 350. Meanwhile, a receiver (not shown) for receiving a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiver may be a component of the entropy decoder 310. Meanwhile, the decoding device according to this document may be called a video / image / picture decoding device, and the decoding device may be divided (classified) into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, the inverse transform unit 322, the addition unit 340, the filtering unit 350, the memory 360, the inter prediction unit 332, and the intra prediction unit 331.
[0068] The inverse quantization unit 321 may inverse quantize the quantized transform coefficients to output transform coefficients. The inverse quantization unit 321 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit 321 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.
[0069] The inverse transform unit 322 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0070] The prediction unit may perform prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on the prediction information output from the entropy decoding unit 310, and may determine a specific intra / inter prediction mode.
[0071] The predictor 320 may generate a prediction signal based on various prediction methods, which will be described later. For example, the predictor may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The predictor may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / moving picture coding, such as games, such as Screen Content Coding (SCC). IBC basically performs prediction within a current picture, but may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described herein. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, information regarding a palette table and a palette index may be included in the video / picture information and signaled.
[0072] The intra prediction unit 331 may predict the current block by referring to samples in the current picture. The referenced samples may be located adjacent to or distant from the current block depending on the prediction mode. Prediction modes in intra prediction may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0073] The inter prediction unit 332 may derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter prediction unit 332 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.
[0074] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the current block, such as when skip mode is applied, the predicted block can be used as the reconstructed block.
[0075] The adder 340 may be referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of the next block to be processed in the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture.
[0076] Meanwhile, Luma Mapping with Chroma Scaling (LMCS) can be applied in the picture decoding process.
[0077] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may transmit the modified reconstructed picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods may include, for example, deblock filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0078] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter predictor 332. The memory 360 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information can be transmitted to the inter predictor 260 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 360 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 331.
[0079] In this document, the embodiments described for the filtering unit 260, inter prediction unit 221, and intra prediction unit 222 of the encoding device 200 can also be applied identically or correspondingly to the filtering unit 350, inter prediction unit 332, and intra prediction unit 331 of the decoding device 300, respectively.
[0080] As described above, prediction is performed to improve compression efficiency when performing video coding. Through this, a predicted block including predicted samples for a current block, which is a block to be coded, can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived in the same way by an encoding device and a decoding device. The encoding device can improve video coding efficiency by signaling to the decoding device information (residual information) regarding the residual between an original block and a predicted block, rather than the original sample values of the original block themselves. The decoding device can derive a residual block including residual samples based on the residual information, combine the residual block with the predicted block to generate a reconstructed block including reconstructed samples, and generate a reconstructed picture including the reconstructed block.
[0081] The residual information may be generated through a transform and quantization procedure. For example, an encoding device may derive a residual block between an original block and a predicted block, perform a transform procedure on residual samples (residual sample array) included in the residual block to derive transform coefficients, and perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, and then signal the related residual information to a decoding device (via a bitstream). Here, the residual information may include information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding device may derive residual samples (or residual block) by performing an inverse quantization / inverse transform process based on the residual information. The decoding device may generate a reconstructed picture based on the predicted block and the residual block. The encoding device may also derive a residual block by inverse quantizing / inverse transforming the quantized transform coefficients for reference for inter-prediction of a subsequent picture, and generate a reconstructed picture based on the residual block.
[0082] FIG. 4 shows an example of a general video / image encoding method to which embodiments of this document can be applied.
[0083] The method disclosed in Fig. 4 may be performed by the encoding device 200 of Fig. 2. Specifically, S400 may be performed by the inter prediction unit 221 or the intra prediction unit 222 of the encoding device 200, and S410, S420, S430, and S440 may be performed by the subtraction unit 231, the transformation unit 232, the quantization unit 233, and the entropy coding unit 240 of the encoding device 200, respectively.
[0084] 4, the encoding device may derive a prediction sample through prediction for a current block (S400). The encoding device may determine whether to perform inter prediction or intra prediction for the current block, and may determine a specific inter prediction mode or a specific intra prediction mode based on an RD cost. Depending on the determined mode, the encoding device may derive a prediction sample for the current block.
[0085] The encoding device can derive residual samples by comparing the original samples and predicted samples for the current block (S410).
[0086] The encoding device may derive transform coefficients through a transform procedure on the residual samples (S420), quantize the derived transform coefficients, and derive quantized transform coefficients (S430).
[0087] The encoding device may encode video information including prediction information and residual information and output the encoded video information in the form of a bitstream (S440). The prediction information is information related to a prediction procedure, and may include prediction mode information and information about motion information (e.g., when inter-prediction is applied), etc. The residual information may include information about quantized transform coefficients. The residual information may be entropy coded.
[0088] The output bitstream can be transmitted to a decoding device via a storage medium or a network.
[0089] FIG. 5 shows an example of a general video / image decoding method to which the embodiments of this document can be applied.
[0090] The method disclosed in Fig. 5 may be performed by the decoding device 300 of Fig. 3 described above. Specifically, S500 may be performed by the inter prediction unit 332 or the intra prediction unit 331 of the decoding device 300. In S500, the procedure of decoding prediction information included in the bitstream and deriving values of associated syntax elements may be performed by the entropy decoding unit 310 of the decoding device 300. S510, S520, S530, and S540 may be performed by the entropy decoding unit 310, the inverse quantization unit 321, the inverse transform unit 322, and the adder 340 of the decoding device 300, respectively.
[0091] 5, the decoding device may perform operations corresponding to those performed by the encoding device. The decoding device may perform inter prediction or intra prediction on the current block based on received prediction information to derive prediction samples (S500).
[0092] The decoding device may derive quantized transform coefficients for the current block based on the received residual information (S510). The decoding device may derive the quantized transform coefficients from the residual information through entropy decoding.
[0093] The decoding device can derive the transform coefficients by dequantizing the quantized transform coefficients (S520).
[0094] The decoding device derives residual samples through an inverse transform process on the transform coefficients (S530).
[0095] The decoding apparatus generates reconstructed samples for the current block based on the predicted samples and the residual samples, and generates a reconstructed picture based on the reconstructed samples (S540). Thereafter, as described above, an in-loop filtering procedure can be further applied to the reconstructed picture.
[0096] Meanwhile, as described above, intra prediction or inter prediction can be applied to perform prediction on the current block. Hereinafter, a case where inter prediction is applied to the current block will be described.
[0097] A prediction unit (more specifically, an inter prediction unit) of an encoding / decoding device may perform inter prediction on a block-by-block basis to derive predicted samples. Inter prediction may refer to a prediction derived in a manner dependent on data elements (e.g., sample values or motion information) of a picture other than the current picture. When inter prediction is applied to a current block, a predicted block (prediction sample array) for the current block may be derived based on a reference block (reference sample array) identified by a motion vector in a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, motion information of the current block may be predicted on a block, sub-block, or sample-by-block basis based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction type information (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). When inter prediction is applied, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), or the like, and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, a motion information candidate list may be constructed based on the neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) to derive a motion vector and / or a reference picture index for the current block may be signaled.Inter prediction may be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the motion information of the current block may be the same as that of a selected neighboring block. In the case of skip mode, unlike in the merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of a selected neighboring block may be used as a motion vector predictor, and a motion vector difference may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference.
[0098] The motion information may include L0 motion information and / or L1 motion information depending on the inter-prediction type (such as L0 prediction, L1 prediction, or Bi-prediction). A motion vector in the L0 direction may be referred to as an L0 motion vector or MVL0, and a motion vector in the L1 direction may be referred to as an L1 motion vector or MVL1. Prediction based on an L0 motion vector may be referred to as L0 prediction, prediction based on an L1 motion vector may be referred to as L1 prediction, and prediction based on both an L0 motion vector and an L1 motion vector may be referred to as bi-prediction. Here, the L0 motion vector may indicate a motion vector associated with a reference picture list L0 (L0), and the L1 motion vector may indicate a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include, as reference pictures, pictures that are earlier in output order than the current picture, and the reference picture list L1 may include pictures that are later in output order than the current picture. Previous pictures may be called forward (reference) pictures, and subsequent pictures may be called backward (reference) pictures. The reference picture list L0 may further include, as reference pictures, pictures that are later in output order than the current picture. In this case, the previous picture may be indexed first in the reference picture list L0, and the subsequent picture may be indexed next. The reference picture list L1 may further include, as reference pictures, pictures that are earlier in output order than the current picture. In this case, the subsequent picture may be indexed first in the reference picture list L1, and the previous picture may be indexed next. Here, the output order may correspond to a POC (Picture Order Count) order.
[0099] In addition, various inter-prediction modes can be used to predict the current block in a picture. For example, various modes can be used, such as merge mode, skip mode, Motion Vector Prediction (MVP) mode, affine mode, sub-block merge mode, merge with MVD (MMVD) mode, and Historical Motion Vector Prediction (HMVP) mode. Additional modes can also be used, such as Decoder-side Motion Vector Refinement (DMVR) mode, Adaptive Motion Vector Resolution (AMVR) mode, Bi-prediction with CU-level Weight (BCW), and Bi-Directional Optical Flow (BDOF). Affine mode may also be referred to as affine motion prediction mode. MVP mode may also be referred to as Advanced Motion Vector Prediction (AMVP) mode. In this document, some modes and / or motion information candidates derived by some modes may be included as motion information-related candidates of other modes. For example, an HMVP candidate may be added as a merge candidate in a merge / skip mode, or may be added as an MVP candidate in an MVP mode. When an HMVP candidate is used as a motion information candidate in a merge mode or a skip mode, the HMVP candidate may be referred to as an HMVP merge candidate.
[0100] Prediction mode information indicating the inter prediction mode of the current block may be signaled from the encoding device to the decoding device. In this case, the prediction mode information may be included in a bitstream and received by the decoding device. The prediction mode information may include index information indicating one of a number of candidate modes. Alternatively, the inter prediction mode may be indicated through hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether the skip mode is applicable, and if the skip mode is not applicable, a merge flag may be signaled to indicate whether the merge mode is applicable, and if the merge mode is not applicable, an MVP mode may be indicated to be applied, or an additional distinguishing flag may be signaled. The affine mode may be signaled as an independent mode or as a mode subordinate to the merge mode or MVP mode. For example, the affine mode may include an affine merge mode and an affine MVP mode.
[0101] Furthermore, motion information of the current block may be used when applying inter prediction to the current block. The encoding device may derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding device may use an original block in an original picture for the current block to search for a highly correlated similar reference block in a predetermined search range in the reference picture in fractional pixel units, thereby deriving motion information. Block similarity may be derived based on a difference in sample values based on phase. For example, block similarity may be calculated based on the sum of absolute differences (SAD) between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, motion information may be derived based on the reference block with the smallest SAD within a search space. The derived motion information may be signaled to a decoding device according to various methods based on the inter prediction mode.
[0102] As described above, a predicted block for a current block may be derived based on motion information derived by the inter prediction mode. The predicted block may include prediction samples (prediction sample array) of the current block. If the motion vector (MV) of the current block indicates a fractional sample unit, an interpolation procedure may be performed, through which prediction samples of the current block may be derived based on reference samples of fractional sample units in a reference picture. If affine inter prediction is applied to the current block, prediction samples may be generated based on sample / sub-block unit MVs. If bi-prediction is applied, prediction samples derived through a weighted sum or weighted average (by phase) of prediction samples derived based on L0 prediction (i.e., prediction using a reference picture in reference picture list L0 and MVL0) and prediction samples derived based on L1 prediction (i.e., prediction using a reference picture in reference picture list L1 and MVL1) may be used as prediction samples of the current block. When bi-prediction is applied, if the reference picture used for L0 prediction and the reference picture used for L1 prediction are located in different time directions relative to the current picture (i.e., if it is bi-predictive and supports both directions of prediction), this can be called true bi-prediction.
[0103] As mentioned above, reconstructed samples and reconstructed pictures can be generated based on the predicted samples derived as described above, after which procedures such as in-loop filtering can be performed.
[0104] FIG. 6 illustrates an example of a general inter-prediction based video / image encoding method to which embodiments of this document can be applied.
[0105] The method disclosed in Fig. 6 may be performed by the encoding device 200 of Fig. 2. Specifically, S600 may be performed by the inter prediction unit 221 of the encoding device 200, S610 may be performed by the subtraction unit 231 of the encoding device 200, and S620 may be performed by the entropy coding unit 240 of the encoding device 200.
[0106] Referring to FIG. 6, an encoding apparatus may perform inter prediction on a current block (S600). The encoding apparatus may derive an inter prediction mode and motion information for the current block and generate a predicted sample for the current block. Here, the processes of determining the inter prediction mode, deriving the motion information, and generating the predicted sample may be performed simultaneously, or one process may be performed before the other processes. For example, the inter prediction unit of the encoding apparatus may include a prediction mode determination unit, a motion information derivation unit, and a predicted sample derivation unit, where the prediction mode determination unit may determine a prediction mode for the current block, the motion information derivation unit may derive motion information for the current block, and the predicted sample derivation unit may derive a predicted sample for the current block. For example, the inter prediction unit of the encoding apparatus may search for a block similar to the current block within a predetermined region (search region) of a reference picture through motion estimation and derive a reference block whose difference from the current block is minimal or equal to or less than a predetermined standard. Based on this, a reference picture index indicating the reference picture in which the reference block is located may be derived, and a motion vector may be derived based on the positional difference between the reference block and the current block. The encoding device can determine a mode to be applied to the current block from among various prediction modes, and can compare RD costs for various prediction modes to determine the optimal prediction mode for the current block.
[0107] For example, when a skip mode or a merge mode is applied to a current block, the encoding device may construct a merge candidate list and derive a reference block whose difference from the current block is minimum or equal to or less than a predetermined criterion among reference blocks indicated by merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to a decoding device. Motion information of the current block may be derived using motion information of the selected merge candidate.
[0108] As another example, when the (A)MVP mode is applied to the current block, the encoding apparatus may construct an (A)MVP candidate list and use a motion vector of an MVP (motion vector predictor) candidate selected from the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, a motion vector indicating a reference block derived by the above-described motion estimation may be used as the motion vector of the current block, and the MVP candidate having the smallest difference from the motion vector of the current block among the MVP candidates may be the selected MVP candidate. A motion vector difference (MVD), which is the difference obtained by subtracting the MVP from the motion vector of the current block, may be derived. In this case, information regarding the MVD may be signaled to the decoding apparatus. Furthermore, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and separately signaled to the decoding apparatus.
[0109] The encoding apparatus may derive residual samples based on the predicted samples (S610) by comparing the original samples of the current block with the predicted samples.
[0110] The encoding apparatus may encode video information including prediction information and residual information (S620). The encoding apparatus may output the encoded video information in the form of a bitstream. Here, the prediction information may include prediction mode information (e.g., a skip flag, a merge flag, or a mode index) and information about motion information as information about a prediction procedure. The information about the motion information may include candidate selection information (e.g., a merge index, an MVP flag, or an MVP index) that is information for deriving a motion vector. The information about the motion information may also include information about the above-mentioned MVD and / or reference picture index information. The information about the motion information may also include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information about residual samples. The residual information may include information about quantized transform coefficients for the residual samples.
[0111] The output bitstream can be stored on a (digital) storage medium and transmitted to the decoding device, or can be transmitted to the decoding device via a network.
[0112] Meanwhile, as described above, the encoding apparatus can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is because the encoding apparatus derives the same prediction result as that performed by the decoding apparatus, thereby improving coding efficiency. Therefore, the encoding apparatus can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in a memory and use it as a reference picture for inter prediction. As described above, an in-loop filtering procedure can be further applied to the reconstructed picture.
[0113] FIG. 7 illustrates an example of a general inter-prediction based video / picture decoding method to which the embodiments of this document can be applied.
[0114] The method disclosed in Fig. 7 may be performed by the decoding device 300 of Fig. 3 described above. Specifically, S700 may be performed by the inter prediction unit 332 of the decoding device 300. The procedure of decoding prediction information included in the bitstream and deriving values of associated syntax elements in S700 may be performed by the entropy decoding unit 310 of the decoding device 300. S710 and S720 may be performed by the inter prediction unit 332 of the decoding device 300, S730 may be performed by the residual processing unit 320 of the decoding device 300, and S740 may be performed by the adder 340 of the decoding device 300.
[0115] 7, the decoding device may perform operations corresponding to those performed by the encoding device. The decoding device may perform prediction on the current block based on the received prediction information to derive predicted samples.
[0116] Specifically, the decoding device may determine a prediction mode for a current block based on received prediction information (S700). The decoding device may determine which inter-prediction mode is applied to the current block based on prediction mode information in the prediction information.
[0117] For example, a merge flag may be used to determine whether a merge mode is applied to the current block, or whether an (A)MVP mode is selected. Alternatively, a mode index may be used to select one of various inter-prediction mode candidates. The inter-prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include various inter-prediction modes.
[0118] The decoding device may derive motion information of the current block based on the determined inter-prediction mode (S710). For example, when a skip mode or a merge mode is applied to the current block, the decoding device may construct a merge candidate list and select one of the merge candidates included in the merge candidate list. The selection may be performed based on the selection information (merge index). Motion information of the selected merge candidate may be used to derive motion information of the current block. The motion information of the selected merge candidate may be used as motion information of the current block.
[0119] As another example, when the (A)MVP mode is applied to the current block, the decoding device may construct an (A)MVP candidate list and use a motion vector of an MVP (motion vector predictor) candidate selected from among the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. The selection may be performed based on the selection information (MVP flag or MVP index). In this case, the MVD of the current block may be derived based on information related to the MVD, and the motion vector of the current block may be derived based on the MVP and MVD of the current block. Also, the reference picture index of the current block may be derived based on reference picture index information. A picture indicated by a reference picture index in the reference picture list for the current block may be derived as a reference picture referenced for inter-prediction of the current block.
[0120] On the other hand, the motion information of the current block may be derived without constructing a candidate list, and in this case, the motion information of the current block may be derived according to a procedure disclosed in a prediction mode described below. In this case, the construction of the candidate list as described above may be omitted.
[0121] The decoding device may generate predictive samples for the current block based on the motion information of the current block (S720). In this case, a reference picture may be derived based on the reference picture index of the current block, and the predictive samples of the current block may be derived using samples of the reference block indicated by the motion vector of the current block in the reference picture. In this case, as described below, a predictive sample filtering procedure may be further performed on all or some of the predictive samples of the current block, depending on the case.
[0122] For example, the inter-prediction unit of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit, and may determine a prediction mode for a current block based on the prediction mode information received by the prediction mode determination unit, derive motion information (such as a motion vector and / or a reference picture index) for the current block based on information regarding the motion information received by the motion information derivation unit, and derive a prediction sample for the current block by the prediction sample derivation unit.
[0123] The decoding device may generate residual samples for the current block based on the received residual information (S730). The decoding device may generate reconstructed samples for the current block based on the predicted samples and the residual samples, and generate a reconstructed picture based on the reconstructed samples (S740). Thereafter, as described above, an in-loop filtering procedure or the like may be further applied to the reconstructed picture.
[0124] 8 exemplarily illustrates an inter prediction procedure. The inter prediction procedure disclosed in FIG. 8 may be applied to the inter prediction process (when an inter prediction mode is applied) disclosed in FIG. 6 and FIG. 7.
[0125] 8, as described above, the inter prediction procedure may include an inter prediction mode determination step, a motion information derivation step according to the determined prediction mode, and a prediction (prediction sample generation) step based on the derived motion information. As described above, the inter prediction procedure may be performed by an encoding device and a decoding device. In this document, a coding device may include an encoding device and / or a decoding device.
[0126] The coding apparatus may determine an inter prediction mode for a current block (S800). Various inter prediction modes may be used for predicting the current block in a picture. For example, various modes may be used, such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, and merge with MVD (MMVD) mode. Decoder-side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, bi-prediction with CU-level weight (BCW), bi-directional optical flow (BDOF), etc. may be used additionally or alternatively as additional modes. Affine mode may also be referred to as affine motion prediction mode. MVP mode may also be referred to as advanced motion vector prediction (AMVP) mode. In this document, some modes and / or motion information candidates derived by some modes may be included as one of the motion information-related candidates of other modes. For example, an HMVP candidate may be added as a merge candidate in a merge / skip mode, or as an MVP candidate in an MVP mode. When an HMVP candidate is used as a motion information candidate in a merge mode or a skip mode, the HMVP candidate may be referred to as an HMVP merge candidate.
[0127] Prediction mode information indicating the inter prediction mode of the current block may be signaled from the encoding device to the decoding device. The prediction mode information may be included in a bitstream and received by the decoding device. The prediction mode information may include index information indicating one of multiple candidate modes. Alternatively, the inter prediction mode may be indicated through hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether the skip mode is applicable, and if the skip mode is not applicable, a merge flag may be signaled to indicate whether the merge mode is applicable, and if the merge mode is not applicable, an MVP mode may be signaled to indicate that the MVP mode is applied, or a flag for additional distinction may be further signaled. The affine mode may be signaled as an independent mode or as a mode dependent on the merge mode, MVP mode, etc. For example, the affine mode may include an affine merge mode and an affine MVP mode.
[0128] The coding apparatus may derive motion information for the current block (S810). The motion information may be derived based on the inter prediction mode.
[0129] A coding apparatus may perform inter-prediction using motion information of a current block. An encoding apparatus may derive optimal motion information for a current block through a motion estimation procedure. For example, the encoding apparatus may use an original block in an original picture for the current block to search for a highly correlated similar reference block in a predetermined search range in the reference picture in units of fractional pixels, thereby deriving motion information. Block similarity may be derived based on a difference in sample values based on phase. For example, block similarity may be calculated based on the SAD between the current block (or a template of the current block) and a reference block (or a template of the reference block). In this case, motion information may be derived based on the reference block with the smallest SAD in the search space. The derived motion information may be signaled to a decoding apparatus in various ways based on the inter-prediction mode.
[0130] The coding device may perform inter prediction based on motion information for the current block (S820). The coding device may derive predictive samples for the current block based on the motion information. The current block including the predictive samples may be referred to as a predicted block.
[0131] Meanwhile, when deriving motion information of the current block, motion information candidates may be derived based on spatially and temporally neighboring blocks, and the motion information candidates for the current block may be selected based on the derived motion information candidates. In this case, the selected motion information candidates may be used as motion information of the current block.
[0132] FIG. 9 exemplarily shows spatial and temporal neighboring blocks of the current block.
[0133] 9, spatial neighboring blocks refer to neighboring blocks located around a current block 900 on which inter prediction is currently being performed, and may include neighboring blocks located around the left side of the current block 900 or neighboring blocks located around the top side of the current block 900. For example, spatial neighboring blocks may include a lower left corner neighboring block, a left neighboring block, a top right corner neighboring block, an upper left corner neighboring block, and an upper left corner neighboring block of the current block 900. In FIG. 9, spatial neighboring blocks are indicated by "S."
[0134] In one embodiment, the encoding device / decoding device can search spatial neighboring blocks of the current block (e.g., the lower left corner neighboring block, the left side neighboring block, the upper right corner neighboring block, the upper neighboring block, and the upper left corner neighboring block) in a specified order to detect usable neighboring blocks, and derive motion information of the detected neighboring blocks as spatial motion information candidates.
[0135] The temporal neighboring block is a block located in a picture (i.e., a reference picture) different from the current picture containing the current block 900, and refers to a block (collocated block; col block) that is co-located with the current block 900 in the reference picture. Here, the reference picture may be earlier or later than the current picture in the Picture Order Count (POC). Furthermore, the reference picture used in deriving the temporal neighboring block may be referred to as a co-located reference picture or col picture (collocated picture; col picture). Furthermore, the co-located block (collocated block) may refer to a block located in a col picture that corresponds to the position of the current block 900, and may be referred to as a col block. For example, the temporal neighboring blocks may include a col block (i.e., a col block including a lower right corner sample) located in a reference picture (i.e., a col picture) corresponding to the lower right corner sample position of the current block 900 and / or a col block (i.e., a col block including a center lower right sample) located in a reference picture (i.e., a col picture) corresponding to the center lower right sample position of the current block 900, as shown in Figure 9. In Figure 9, the temporal neighboring blocks are indicated by "T".
[0136] In one embodiment, the encoding device / decoding device may search for a usable block among temporally adjacent blocks of a current block (e.g., a col block including a lower right corner sample, a col block including a lower right center sample) in a specified order, and derive motion information of the detected block as a temporal motion information candidate. A technique using temporally adjacent blocks in this manner may be referred to as TMVP (Temporal Motion Vector Prediction). The temporal motion information candidate may also be referred to as a TMVP candidate.
[0137] On the other hand, depending on the inter prediction mode, it is also possible to derive motion information in subblock units and perform prediction. For example, in the case of affine mode or TMVP mode, motion information can be derived in subblock units. In particular, a method of deriving temporal motion information candidates in subblock units can be referred to as subblock-based TMVP (sbTMVP).
[0138] sbTMVP is a method that uses motion fields in a col picture to improve the motion vector prediction (MVP) and merge mode of coding units within a current picture. The col picture of sbTMVP may be the same as the col picture used by TMVP. However, while TMVP performs motion prediction at the coding unit (CU) level, sbTMVP can perform motion prediction at the sub-block or sub-coding unit (sub-CU) level. TMVP also derives temporal motion information from a col block in a col picture (where the col block is the col block corresponding to the lower right corner sample position or the center lower right sample position of the current block). sbTMVP derives temporal motion information after applying a motion shift from the col picture. Here, the motion shift can include a process of obtaining a motion vector from one of the spatially neighboring blocks of the current block and shifting the motion vector.
[0139] FIG. 10 exemplarily illustrates spatially neighboring blocks that can be used to derive sub-block based temporal motion information candidates (sbTMVP candidates).
[0140] 10, the spatial neighboring blocks may include at least one of the lower left corner neighboring block A0, the left neighboring block A1, the upper right corner neighboring block B0, and the upper neighboring block B1 of the current block. In some cases, the spatial neighboring blocks may further include neighboring blocks other than the neighboring blocks shown in FIG. 10, or may not include a specific neighboring block among the neighboring blocks shown in FIG. 10. Alternatively, the spatial neighboring blocks may include only a specific neighboring block, for example, only the left neighboring block A1 of the current block.
[0141] For example, the encoding / decoding device may search for spatially neighboring blocks according to a predetermined search order, detect the motion vector of the spatially neighboring block that can be used first, and determine the block located at the position indicated by the motion vector of the spatially neighboring block in the reference picture as a col block (i.e., a co-located reference block). Here, the motion vector of the spatially neighboring block may be referred to as a temporal motion vector (temporal MV).
[0142] In this case, the usability of a spatial neighboring block may be determined based on its reference picture information, prediction mode information, position information, etc. For example, if the reference picture of the spatial neighboring block is the same as the reference picture of the current block, the spatial neighboring block may be determined to be usable. Alternatively, if the spatial neighboring block is coded in intra prediction mode or is located outside the current picture / tile, the spatial neighboring block may be determined to be unusable.
[0143] The search order of the spatially adjacent blocks may be variously defined, for example, A1, B1, B0, A0, or it may be determined whether A1 is available by searching only A1.
[0144] FIG. 11 is a diagram for explaining in outline the process of deriving sub-block-based temporal motion information candidates (sbTMVP candidates).
[0145] Referring to FIG. 11, first, the encoding / decoding apparatus may determine whether a spatial neighboring block (e.g., block A1) of a current block is usable. For example, if the reference picture of the spatial neighboring block (e.g., block A1) uses the col picture, the spatial neighboring block (e.g., block A1) may be determined to be usable, and a motion vector of the spatial neighboring block (e.g., block A1) may be derived. In this case, the motion vector of the spatial neighboring block (e.g., block A1) may be referred to as a temporal MV (tempMV), and this motion vector may be used for motion shifting. Alternatively, if it is determined that the spatial neighboring block (e.g., block A1) is unusable, the temporal MV (i.e., the motion vector of the spatial neighboring block) may be set to a zero vector. In other words, in this case, a motion vector set to (0,0) may be applied for motion shifting.
[0146] Next, the encoding / decoding device can apply a motion shift based on the motion vector of the spatially adjacent block (e.g., block A1). For example, the motion shift can be shifted (e.g., A1') to the position indicated by the motion vector of the spatially adjacent block (e.g., block A1). That is, by applying the motion shift, the motion vector of the spatially adjacent block (e.g., block A1) can be added to the coordinates of the current block.
[0147] Next, the encoding / decoding device can derive motion-shifted col subblocks (collocated subblocks) on the col picture and obtain motion information (motion vectors, reference indexes, etc.) for each col subblock. For example, the encoding / decoding device can derive each col subblock on the col picture that corresponds to a motion-shifted position (i.e., a position indicated by the motion vector of a spatially neighboring block (e.g., A1)) for each subblock position in the current block. Then, the motion information of each col subblock can be used as the motion information of each subblock relative to the current block (i.e., sbTMVP candidate).
[0148] In addition, scaling may be applied to the motion vector of the col sub-block. The scaling may be performed based on the difference in temporal distance between the reference picture of the col block and the reference picture of the current block. Therefore, the scaling may be referred to as temporal motion scaling, through which the reference picture of the current block and the reference picture of the temporal motion vector may be aligned. In this case, the encoding / decoding apparatus may obtain the motion vector of the scaled col sub-block as motion information of each sub-block relative to the current block.
[0149] Furthermore, when deriving sbTMVP candidates, there may be cases where motion information does not exist for the col sub-block. In this case, base motion information (or default motion information) can be derived for the col sub-block for which no motion information exists, and this base motion information can be used as the motion information of the sub-block for the current block. The base motion information can be derived from a block located at the center of the col block (i.e., the col CU including the col sub-block). For example, motion information (e.g., a motion vector) can be derived from a block including the sample located at the bottom right of the four samples located at the center of the col block, and this can be used as the base motion information.
[0150] As described above, in the case of affine mode or sbTMVP mode, which derive motion information on a sub-block basis, affine merge candidates and sbTMVP candidates can be derived, and a sub-block-based merge candidate list can be constructed based on these candidates. In this case, flag information indicating whether the affine mode or sbTMVP mode is enabled or disabled can be signaled. If the sbTMVP mode is enabled based on the flag information, the sbTMVP candidate derived as described above can be added to the first-ordered entry of the sub-block-based merge candidate list. Then, the affine merge candidate can be added to the next entry of the sub-block-based merge candidate list. Here, the maximum number of candidates in the sub-block-based merge candidate list can be five.
[0151] In addition, in the sbTMVP mode, the size of the sub-blocks may be fixed, for example, 8 × 8. In addition, the sbTMVP mode can be applied only to blocks whose width and height are both 8 or more.
[0152] On the other hand, in the current VVC standard, it is possible to derive sub-block based temporal motion information candidates (sbTMVP candidates) as shown in Table 1 below.
[0153] [Table 1-1]
[0154] [Table 1-2]
[0155] [Table 1-3]
[0156] In deriving sbTMVP candidates according to the method shown in Table 1 above, a default MV and sub-block MV(s) may be considered. Here, the default MV may be referred to as subblock-based temporal merging base motion data or base motion data (base motion information). Referring to Table 1 above, the default MV may correspond to ctrMV (or ctrMVLX) in Table 1. The sub-block MV may correspond to mvSbCol (or mvLXSbcol) in Table 1.
[0157] For example, if a sub-block or sub-block MV is available during the sbTMVP derivation process, the sub-block MV is assigned to the sub-block. Alternatively, if a sub-block or sub-block MV is unavailable, the default MV may be used as the sub-block MV for the sub-block. Here, the default MV derives motion information from a position corresponding to the center pixel position of the corresponding block (i.e., col CU) on the col picture, and each sub-block MV may derive motion information from the top-left position of the corresponding sub-block (i.e., col sub-block) on the col picture. In this case, the corresponding block (i.e., col CU) may be derived from a position motion-shifted based on the motion vector (i.e., temporal MV) of spatial neighboring block A1, as shown in FIG. 11.
[0158] 12 to 15 are diagrams for explaining a method for calculating corresponding positions for deriving default MVs and sub-block MVs according to block sizes in the sbTMVP derivation process.
[0159] In Figures 12 to 15, pixels (samples) shaded with dotted lines indicate corresponding positions within each sub-block for deriving each sub-block MV, and pixels (samples) shaded with solid lines indicate corresponding positions within a CU for deriving a default MV.
[0160] For example, referring to FIG. 12, if the current block (i.e., the current CU) is 8x8 in size, the motion information of the sub-block can be derived based on the top left sample position within the 8x8 sized sub-block, and the default motion information of the sub-block can be derived based on the center sample position within the 8x8 sized current block (i.e., the current CU).
[0161] Alternatively, for example, referring to Figure 13, if the current block (i.e., the current CU) is 16x8 in size, the motion information of each sub-block can be derived based on the top left sample position within each 8x8 size sub-block, and the default motion information of each sub-block can be derived based on the center sample position within the 16x8 size current block (i.e., the current CU).
[0162] Alternatively, for example, referring to Figure 14, if the current block (i.e., the current CU) is 8x16 size, the motion information of each sub-block can be derived based on the top left sample position within each 8x8 size sub-block, and the default motion information of each sub-block can be derived based on the center sample position within the 8x16 size current block (i.e., the current CU).
[0163] Alternatively, for example, referring to Figure 15, if the current block (i.e., the current CU) is 16x16 size, the motion information of each sub-block can be derived based on the top left sample position within each 8x8 size sub-block, and the default motion information of each sub-block can be derived based on the center sample position within the 16x16 size current block (i.e., the current CU).
[0164] As can be seen from Figures 12 to 15, since the motion information of the sub-block is biased toward the top left pixel position, there is a problem that the sub-block MV is derived at a position far removed from the position where the default MV indicating the representative motion information of the current CU is derived. As a typical example, in the case of an 8x8 block shown in Figure 12, one CU includes one sub-block, but there is a contradiction in that the sub-block MV and the default MV are represented by different motion information. Furthermore, since the methods for calculating the corresponding position between the sub-block and the current CU block are different (i.e., the corresponding position for deriving the sub-block MV is the top left sample position, and the corresponding position for deriving the default MV is the center sample position), an additional module may be required when implementing hardware (H / W).
[0165] To address the above issues, this document proposes a unification method for deriving the corresponding positions of CUs for a default MV and the corresponding positions of sub-blocks for each sub-block MV in the process of deriving sbTMVP candidates. According to one embodiment of this document, the unification is achieved by using only one module for deriving each corresponding position according to the block size in terms of hardware (H / W). For example, the method for calculating the corresponding positions when the block size is 16x16 blocks and the method for calculating the corresponding positions when the block size is 8x8 blocks can be implemented in the same way, thereby simplifying the hardware implementation. Here, the 16x16 blocks represent CUs, and the 8x8 blocks represent each sub-block.
[0166] In one embodiment, when deriving sbTMVP candidates, the center sample position can be used as the corresponding position for deriving motion information of the sub-block and the corresponding position for deriving default motion information, which can be implemented as shown in Table 2 below.
[0167] Table 2 below is a spec showing an example of how to derive sub-block motion information and default motion information according to one embodiment of this document.
[0168] [Table 2-1]
[0169] [Table 2-2]
[0170] [Table 2-3]
[0171] Referring to Table 2 above, when deriving sbTMVP candidates, the position of the current block (i.e., current CU) including the sub-block can be derived. The top left sample position (xCtb, yCtb) of the coding tree block (or coding tree unit) including the current block and the bottom right center sample position (xCtr, yCtr) of the current block can be derived according to equations (8-514) to (8-517) in Table 2 above. In this case, the positions (xCtb, yCtb) and (xCtr, yCtr) can be calculated based on the top left sample position (xCb, yCb) of the current block with respect to the top left sample of the current picture.
[0172] Also, a col block (i.e., col CU) on the col picture located corresponding to the current block (i.e., current CU) including the sub-block can be derived. In this case, the position of the col block can be set to (xColCtrCb, yColCtrCb), and this position can indicate the position of the col block including the position (xCtr, yCtr) in the col picture based on the top left sample of the col picture.
[0173] Furthermore, base motion data (i.e., default motion information) for sbTMVP can be derived. The base motion data can include a default motion vector (e.g., ctrMvLX). For example, a col block on a col picture can be derived. In this case, the position of the col block can be derived as (xColCb, yColCb), which can be a position obtained by applying a motion shift (e.g., tempMv) to the derived col block position (xColCtrCb, yColCtrCb). The motion shift can be performed by adding a motion vector (e.g., tempMv) derived from a spatially neighboring block (e.g., A1 block) of the current block to the current col block position (xColCtrCb, yColCtrCb), as described above. Next, a default motion vector (e.g., ctrMvLX) can be derived based on the motion-shifted col block position (xColCb, yColCb). Here, the default MV (for example, ctrMvLX) may indicate a motion vector derived from a position corresponding to the center sample at the bottom right of the col block.
[0174] Furthermore, a col sub-block on the col picture corresponding to a sub-block (referred to as a current sub-block) in the current block can be derived. First, the position of each current sub-block can be derived. The position of each sub-block can be represented by (xSb, ySb), which can indicate the position of the current sub-block based on the top left sample of the current picture. For example, the position (xSb, ySb) of the current sub-block can be calculated using equations (8-523) to (8-524) in Table 2 above, which can indicate the position of the center sample at the bottom right of the sub-block. Next, the position of each col sub-block on the col picture can be derived. The position of each col sub-block can be represented by (xColSb, yColSb), which can be a position obtained by applying a motion shift (e.g., tempMv) to the position (xSb, ySb) of the current sub-block. As described above, motion shifting can be performed by adding a motion vector (e.g., tempMv) derived from a spatially neighboring block (e.g., block A1) of the current block to the position (xSb, ySb) of the current sub-block. Next, motion information (e.g., motion vector mvLXSbCol, flag indicating whether it is available or not) of the col sub-block can be derived based on the position (xColSb, yColSb) of each of the motion-shifted col sub-blocks.
[0175] In this case, if there is an unavailable col sub-block in the col sub-block (for example, if availableFlagLXSbCol is 0), the base motion data (i.e., default motion information) can be used for the unavailable col sub-block. For example, the motion vector (e.g., mvLXSbCol) for the unavailable col sub-block can use the default MV (e.g., ctrMvLX).
[0176] 16 to 19 are exemplary diagrams for explaining a method of integrating corresponding positions for deriving default MVs and sub-block MVs according to block sizes in the process of deriving sbTMVP.
[0177] In Figures 16 to 19, pixels (samples) shaded with dotted lines indicate corresponding positions within each sub-block for deriving each sub-block MV, and pixels (samples) shaded with solid lines indicate corresponding positions within a CU for deriving a default MV.
[0178] For example, referring to FIG. 16, if the current block (i.e., the current CU) is 8x8 in size, the motion information of the current sub-block can be derived from the col sub-block at a corresponding position on the col picture based on the center sample position of the bottom right edge of the 8x8 sub-block. The default motion information of the current sub-block can be derived from the col block (i.e., col CU) at a corresponding position on the col picture based on the center sample position of the bottom right edge of the 8x8 current block (i.e., the current CU). In this case, as shown in FIG. 16, the motion information of the current sub-block and the default motion information can be derived from the same sample position (the same corresponding position).
[0179] 17, for example, when the current block (i.e., the current CU) is 16x8 in size, the motion information of the current sub-block can be derived from the col sub-block at a corresponding position on the col picture based on the center sample position of the bottom right edge in the 8x8 sub-block. The default motion information of the current sub-block can be derived from the col block (i.e., col CU) at a corresponding position on the col picture based on the center sample position of the bottom right edge in the 16x8 current block (i.e., the current CU).
[0180] 18, for example, when the current block (i.e., the current CU) is 8x16 size, the motion information of the current sub-block can be derived from the col sub-block at a corresponding position on the col picture based on the center sample position of the bottom right edge in the 8x8 size sub-block. The default motion information of the current sub-block can be derived from the col block (i.e., col CU) at a corresponding position on the col picture based on the center sample position of the bottom right edge in the 8x16 size current block (i.e., the current CU).
[0181] 19, for example, when the current block (i.e., the current CU) is 16×16 or larger, the motion information of the current sub-block can be derived from the col sub-block at a corresponding position on the col picture based on the center sample position of the bottom right edge of the 8×8 sub-block.The default motion information of the current sub-block can be derived from the col block (i.e., col CU) at a corresponding position on the col picture based on the center sample position of the bottom right edge of the 16×16 (or larger) current block (i.e., the current CU).
[0182] However, the above-described embodiment in this document is merely an example, and the default motion information and the motion information of the current sub-block can be derived based on other sample positions in addition to the center position (i.e., the bottom right sample position). For example, the default motion information can be derived based on the top left sample position of the current CU, and the motion information of the current sub-block can be derived based on the top left sample position of the sub-block.
[0183] When implementing the above-described embodiments in this document in hardware, the same H / W module can be used to derive motion information (temporal motion), so a pipeline such as that shown in Figures 20 and 21 can be configured.
[0184] 20 and 21 are diagrams illustrating an example of a pipeline configuration that can integrate and calculate corresponding positions for deriving default MVs and sub-block MVs in the sbTMVP derivation process.
[0185] 20 and 21, a corresponding position calculation module may calculate corresponding positions for deriving a default MV and a sub-block MV. For example, as shown in FIGS. 20 and 21, when a block position (posX, posY) and a block size (blkszX, blkszY) are input to the corresponding position calculation module, the center position of the input block (i.e., the bottom right sample position) may be output. When the position of the current CU and the block size are input to the corresponding position calculation module, the center position of the col block on the col picture (i.e., the bottom right sample position), which is the corresponding position for deriving the default MV, may be output. Alternatively, when the position and block size of the current sub-block are input to the corresponding position calculation module, the center position of the col sub-block on the col picture (i.e., the bottom right sample position), which is the corresponding position for deriving the current sub-block MV, may be output.
[0186] In this way, when the corresponding positions for deriving the default MV and the sub-block MV are output from the corresponding position calculation module, the motion vector (i.e., temporal MV) derived from the corresponding positions can be patched. Then, sub-block-based temporal motion information (i.e., sbTMVP candidates) can be derived based on the patched motion vector (i.e., temporal MV). For example, as shown in Figures 20 and 21, depending on the hardware implementation, sbTMVP candidates can be derived in parallel or sequentially based on the clock cycle.
[0187] The following drawings are created to illustrate a specific example of the present document. The names of specific devices and specific terms and names (e.g., names of syntax / syntax elements) shown in the drawings are provided for illustrative purposes only, and the technical features of the present document are not limited to the specific names used in the following drawings.
[0188] FIG. 22 illustrates a schematic diagram of an example video / image encoding method according to an embodiment herein.
[0189] The method disclosed in FIG. 22 may be performed by the encoding device 200 disclosed in FIG. 2. Specifically, steps (S2200) to (S2230) of FIG. 22 may be performed by the prediction unit 220 (more specifically, the inter prediction unit 221) disclosed in FIG. 2, step (S2240) of FIG. 22 may be performed by the residual processing unit 230 disclosed in FIG. 2, and step (S2250) of FIG. 22 may be performed by the entropy encoding unit 240 disclosed in FIG. 2. In addition, the method disclosed in FIG. 22 may be performed including the embodiments described above in this document. Therefore, in FIG. 22, detailed descriptions of content that overlaps with the embodiments described above will be omitted or simplified.
[0190] Referring to FIG. 22, the encoding apparatus can derive a reference sub-block on a collocated reference picture for a sub-block in a current block (S2200).
[0191] Here, the co-located reference picture refers to a reference picture used to derive temporal motion information (i.e., sbTMVP), as described above, and may refer to the col picture described above. The reference sub-block may refer to the col sub-block described above.
[0192] In one embodiment, the encoding device may derive a reference sub-block on a co-located reference picture based on the position of the sub-block within a current block, where the current block may be referred to as a current coding unit (CU) or a current coding block (CB), and the sub-block included in the current block may also be referred to as a current coding sub-block.
[0193] For example, the encoding apparatus may first determine the position of a current block, and then determine the positions of sub-blocks within the current block. As described with reference to Table 2 above, the position of the current block may be represented based on the left-side top sample position (xCtb, yCtb) of the coding tree block and the right-side bottom center sample position (xCtr, yCtr) of the current block. The positions of the sub-blocks within the current block may be represented by (xSb, ySb), and these positions (xSb, ySb) may indicate the right-side bottom center sample position of the sub-block. Here, the right-side bottom center sample position (xSb, ySb) of the sub-block may be calculated based on the left-side top sample position of the sub-block and the sub-block size, and may be calculated as shown in Equations (8-523) to (8-524) of Table 2 above.
[0194] Then, the encoding apparatus can derive a reference sub-block on the co-located reference picture based on the right-bottom center sample position of each sub-block in the current block. As described above with reference to Table 2, the reference sub-block can be indicated by a position (xColSb, yColSb) on the co-located reference picture, and the position (xColSb, yColSb) can be derived on the co-located reference picture based on the right-bottom center sample position (xSb, ySb) of each sub-block in the current block.
[0195] On the other hand, the upper left sample position used in this document may also be referred to as the upper left sample position or the upper left sample position, and the lower right center sample position may also be referred to as the lower right center sample position, the center lower right sample position, the lower right center sample position, or the center lower right sample position.
[0196] Furthermore, motion shifting may be applied when deriving the reference sub-block. The encoding apparatus may perform motion shifting based on a motion vector derived from a spatial neighboring block of the current block. The spatial neighboring block of the current block may be a left neighboring block located to the left of the current block, e.g., the A1 block shown in FIGS. 10 and 11. In this case, if the left neighboring block (e.g., the A1 block) is available, a motion vector may be derived from the left neighboring block. Alternatively, if the left neighboring block is unavailable, a zero vector may be derived. Whether the spatial neighboring block is available may be determined based on the reference picture information, prediction mode information, position information, etc., of the spatial neighboring block. For example, if the reference picture of the spatial neighboring block is the same as the reference picture of the current block, the spatial neighboring block may be determined to be available. Alternatively, if the spatial neighboring block is coded in an intra prediction mode or is located outside the current picture / tile, the spatial neighboring block may be determined to be unavailable.
[0197] That is, the encoding apparatus may apply a motion shift (i.e., a motion vector of a spatially neighboring block (e.g., block A1)) to the right-bottom center sample position (xSb, ySb) of each sub-block in the current block, and derive a reference sub-block on the co-located reference picture based on the motion-shifted position. In this case, the position (xColSb, yColSb) of the reference sub-block may be indicated by a position motion-shifted from the right-bottom center sample position (xSb, ySb) of each sub-block in the current block to a position indicated by the motion vector of the spatially neighboring block (e.g., block A1), and may be calculated as shown in Equations (8-525) to (8-526) in Table 2 above.
[0198] The encoding apparatus may derive sub-block-based temporal motion information candidates based on the reference sub-blocks (S2210).
[0199] Meanwhile, in this document, subblock-based temporal motion information candidates are referred to as the above-mentioned sbTMVp (subblock-based Temporal Motion Vector Prediction) candidates, and can be substituted for or mixed with subblock-based temporal motion vector predictor candidates. That is, when motion information is derived and prediction is performed on a subblock basis as described above, sbTMVP candidates can be derived, and motion prediction can be performed at the subblock level (or sub-coding unit (sub-CU) level) based on the sbTMVP candidates.
[0200] The encoding apparatus may derive motion information for sub-blocks within the current block based on the sub-block-based temporal motion information candidates (S2220).
[0201] The sub-block-based temporal motion information candidate may include a sub-block-based motion vector, where the sub-block-based motion vector may include a motion vector derived based on a reference sub-block.
[0202] In one embodiment, the encoding device may derive a subblock-wise motion vector for a reference subblock as motion information for a subblock in the current block. For example, the encoding device may derive a subblock-wise motion vector based on whether the reference subblock is available. For a available reference subblock in the reference subblock, the encoding device may derive a subblock-wise motion vector for the available reference subblock based on the motion vector of the available reference subblock. For an unavailable reference subblock in the reference subblock, the encoding device may use a base motion vector as the subblock-wise motion vector for the unavailable reference subblock.
[0203] The base motion vector may correspond to the default motion vector mentioned above and may be derived on the co-located reference picture based on the position of the current block.
[0204] In one embodiment, the encoding apparatus may identify the position of a reference coding block on a co-located reference picture based on the position of a center sample of the bottom right edge of the current block, and derive a base motion vector based on the position of the reference coding block. The reference coding block may refer to a col block located on a co-located reference picture corresponding to the current block including a sub-block. As described with reference to Table 2 above, the position of the reference coding block may be represented by (xColCtrCb, yColCtrCb), where the position (xColCtrCb, yColCtrCb) may indicate the position of the reference coding block covering the position (xCtr, yCtr) in the co-located reference picture based on the top left sample of the co-located reference picture. The position (xCtr, yCtr) may indicate the position of the center sample of the bottom right edge of the current block.
[0205] Furthermore, when deriving a base motion vector, a motion shift can be applied to the position (xColCtrCb, yColCtrCb) of the reference coding block. As described above, the motion shift can be performed by adding a motion vector derived from a spatially neighboring block (e.g., block A1) of the current block to the reference coding block position (xColCtrCb, yColCtrCb) covering the center sample at the bottom right edge. The encoding apparatus can derive the base motion vector based on the position (xColCb, yColCb) of the motion-shifted reference coding block. That is, the base motion vector can be a motion vector derived from a motion-shifted position on the co-located reference picture based on the center sample position at the bottom right edge of the current block.
[0206] Meanwhile, whether the reference sub-block is usable can be determined based on whether it is located outside the co-located reference picture or on a motion vector. For example, the unusable reference sub-block may include a reference sub-block located outside (or deviating from) the co-located reference picture or a reference sub-block for which a motion vector is unavailable. For example, if the reference sub-block is based on intra mode, IBC (Intra Block Copy) mode, or palette mode, the reference sub-block may be a sub-block for which a motion vector is unavailable. Alternatively, if the reference coding block covering the modified position derived based on the position of the reference sub-block is based on intra mode, IBC mode, or palette mode, the reference sub-block may be a sub-block for which a motion vector is unavailable.
[0207] In this case, in one embodiment, the motion vector of the available reference sub-block may be derived based on the motion vector of the block covering the modified location derived based on the upper left sample location of the reference sub-block. For example, as shown in Table 2 above, the modified location may be derived based on the formula ((xColSb>>3)<<3, (yColSb>>3)<<3). Here, xColSb and yColSb represent the x and y coordinates of the upper left sample location of the reference sub-block, respectively, and >> represents an arithmetic right shift and << represents an arithmetic left shift.
[0208] Meanwhile, as described above, when deriving temporal motion information candidates based on sub-blocks, it can be seen that the motion vector for the reference sub-block is derived based on the position of the sub-block within the current block, and the base motion vector is derived based on the position of the current block. For example, as described with reference to Figures 16 to 19, for a current block having an 8x8 size, the motion vector for the reference sub-block and the base motion vector can be derived based on the center sample position of the bottom right edge of the current block. For a current block having a size larger than 8x8, the motion vector for the reference sub-block can be derived based on the center sample position of the bottom right edge of each sub-block within the current block, and the base motion vector can be derived based on the center sample position of the bottom right edge of the current block.
[0209] The encoding apparatus may generate a predicted sample of the current block based on motion information for a sub-block within the current block (S2230).
[0210] The encoding device may select optimal motion information based on a rate-distortion (RD) cost and generate a predicted sample based on the optimal motion information. For example, if motion information derived for a sub-block of a current block (i.e., sbTMVP) is selected as the optimal motion information, the encoding device may generate a predicted sample of the current block based on the motion information for the sub-block of the current block derived as described above.
[0211] The encoding apparatus may derive residual samples based on the prediction samples (S2240) and encode video information including information about the residual samples (S2250).
[0212] That is, the encoding apparatus may derive residual samples based on original samples for a current block and predicted samples of the current block, and may generate information about the residual samples, where the information about the residual samples may include information such as value information, position information, transform technique, transform kernel, and quantization parameter of quantized transform coefficients derived by performing transform and quantization on the residual samples.
[0213] The encoding device encodes information about the residual samples and outputs it in a bitstream, which can be transmitted to the decoding device via a network or a storage medium.
[0214] FIG. 23 illustrates a schematic diagram of an example video / image decoding method according to an embodiment of the present document.
[0215] The method disclosed in FIG. 23 may be performed by the decoding device 300 disclosed in FIG. 3. Specifically, steps (S2300) to (S2330) of FIG. 23 may be performed by the prediction unit 330 (more specifically, the inter prediction unit 332) disclosed in FIG. 3, and step (S2340) of FIG. 23 may be performed by the addition unit 340 disclosed in FIG. 3. In addition, the method disclosed in FIG. 23 may be performed including the embodiments described above in this document. Therefore, in FIG. 23, detailed descriptions of content that overlaps with the embodiments described above will be omitted or simplified.
[0216] Referring to FIG. 23, the decoding apparatus can derive a reference sub-block on a collocated reference picture for a sub-block in a current block (S2300).
[0217] Here, the co-located reference picture refers to a reference picture used to derive temporal motion information (i.e., sbTMVP), as described above, and may refer to the col picture described above. The reference sub-block may refer to the col sub-block described above.
[0218] In one embodiment, the decoding device can derive a reference sub-block on a co-located reference picture based on the position of the sub-block within a current block, where the current block may be referred to as a current coding unit (CU) or a current coding block (CB), and the sub-block included in the current block may also be referred to as a current coding sub-block.
[0219] For example, the decoding device may first determine the position of the current block, and then determine the position of the sub-block within the current block. As described with reference to Table 2 above, the position of the current block may be represented based on the left-side uppermost sample position (xCtb, yCtb) of the coding tree block and the right-side lowermost center sample position (xCtr, yCtr) of the current block. The positions of the sub-blocks within the current block may be represented by (xSb, ySb), and this position (xSb, ySb) may indicate the right-side lowermost center sample position of the sub-block. Here, the right-side lowermost center sample position (xSb, ySb) of the sub-block may be calculated based on the left-side uppermost sample position of the sub-block and the sub-block size, and may be calculated as shown in Equations (8-523) to (8-524) of Table 2 above.
[0220] Then, the decoding apparatus can derive a reference sub-block on the co-located reference picture based on the right-bottom center sample position of each sub-block in the current block. As described with reference to Table 2 above, the reference sub-block can be indicated by a position (xColSb, yColSb) on the co-located reference picture, and the position (xColSb, yColSb) can be derived on the co-located reference picture based on the right-bottom center sample position (xSb, ySb) of each sub-block in the current block.
[0221] On the other hand, the upper left sample position used in this document may also be referred to as the upper left sample position or the upper left sample position, and the lower right center sample position may also be referred to as the lower right center sample position, the center lower right sample position, the lower right center sample position, or the center lower right sample position.
[0222] Furthermore, a motion shift may be applied when deriving the reference sub-block. The decoding device may perform the motion shift based on a motion vector derived from a spatial neighboring block of the current block. The spatial neighboring block of the current block may be a left neighboring block located to the left of the current block, e.g., the A1 block shown in FIGS. 10 and 11. In this case, if the left neighboring block (e.g., the A1 block) is available, a motion vector may be derived from the left neighboring block. Alternatively, if the left neighboring block is unavailable, a zero vector may be derived. Whether the spatial neighboring block is available may be determined based on the reference picture information, prediction mode information, position information, etc. of the spatial neighboring block. For example, if the reference picture of the spatial neighboring block is the same as the reference picture of the current block, the spatial neighboring block may be determined to be available. Alternatively, if the spatial neighboring block is coded in an intra prediction mode or is located outside the current picture / tile, the spatial neighboring block may be determined to be unavailable.
[0223] That is, the decoding apparatus may apply a motion shift (i.e., a motion vector of a spatially neighboring block (e.g., block A1)) to the right-bottom center sample position (xSb, ySb) of each sub-block in the current block, and derive a reference sub-block on the co-located reference picture based on the motion-shifted position. In this case, the position (xColSb, yColSb) of the reference sub-block may be indicated by a position motion-shifted from the right-bottom center sample position (xSb, ySb) of each sub-block in the current block to a position indicated by the motion vector of the spatially neighboring block (e.g., block A1), and may be calculated as shown in Equations (8-525) to (8-526) in Table 2 above.
[0224] The decoding apparatus can derive sub-block-based temporal motion information candidates based on the reference sub-blocks (S2310).
[0225] Meanwhile, in this document, subblock-based temporal motion information candidates are referred to as the above-mentioned sbTMVp (Subblock-Based Temporal Motion Vector Prediction) candidates, and can be substituted for or mixed with subblock-based temporal motion vector predictor candidates. That is, as described above, when motion information is derived and prediction is performed on a subblock basis, sbTMVP candidates can be derived, and motion prediction can be performed at the subblock level (or sub-coding unit (sub-CU) level) based on the sbTMVP candidates.
[0226] The decoding apparatus can derive motion information for the sub-blocks within the current block based on the sub-block-based temporal motion information candidates (S2320).
[0227] The sub-block-based temporal motion information candidate may include a sub-block-based motion vector, where the sub-block-based motion vector may include a motion vector derived based on a reference sub-block.
[0228] In one embodiment, the decoding device may derive a sub-block-wise motion vector for a reference sub-block as motion information for a sub-block in a current block. For example, the decoding device may derive a sub-block-wise motion vector based on whether the reference sub-block is available. For available reference sub-blocks in the reference sub-block, the decoding device may derive a sub-block-wise motion vector for the available reference sub-block based on the motion vector of the available reference sub-block. For unavailable reference sub-blocks in the reference sub-block, the decoding device may use a base motion vector as the sub-block-wise motion vector for the unavailable reference sub-block.
[0229] The base motion vector may correspond to the default motion vector mentioned above and may be derived on the co-located reference picture based on the position of the current block.
[0230] In one embodiment, the decoding apparatus may identify the position of a reference coding block on a co-located reference picture based on a right-bottom center sample position of the current block, and derive a base motion vector based on the position of the reference coding block. The reference coding block may refer to a col block located on a co-located reference picture corresponding to the current block including a sub-block. As described with reference to Table 2 above, the position of the reference coding block may be represented by (xColCtrCb, yColCtrCb), where the position (xColCtrCb, yColCtrCb) may indicate the position of the reference coding block covering the position (xCtr, yCtr) in the co-located reference picture based on the left-top sample of the co-located reference picture. The position (xCtr, yCtr) may indicate the right-bottom center sample position of the current block.
[0231] Furthermore, when deriving a base motion vector, a motion shift can be applied to the position (xColCtrCb, yColCtrCb) of the reference coding block. As described above, the motion shift can be performed by adding a motion vector derived from a spatially neighboring block (e.g., block A1) of the current block to the reference coding block position (xColCtrCb, yColCtrCb) covering the center sample at the bottom right edge. The decoding apparatus can derive the base motion vector based on the position (xColCb, yColCb) of the motion-shifted reference coding block. That is, the base motion vector can be a motion vector derived from a motion-shifted position on the co-located reference picture based on the center sample position at the bottom right edge of the current block.
[0232] Meanwhile, whether the reference sub-block is usable may be determined based on whether it is located outside the co-located reference picture or on a motion vector. For example, an unusable reference sub-block may include a reference sub-block that deviates from the outside of the co-located reference picture or a reference sub-block for which a motion vector is unavailable. For example, if the reference sub-block is based on an intra mode, an Intra Block Copy (IBC) mode, or a palette mode, the reference sub-block may be a sub-block for which a motion vector is unavailable. Alternatively, if a reference coding block covering a modified position derived based on the position of the reference sub-block is based on an intra mode, an IBC mode, or a palette mode, the reference sub-block may be a sub-block for which a motion vector is unavailable.
[0233] In this case, in one embodiment, the motion vector of the available reference sub-block may be derived based on the motion vector of the block covering the modified location derived based on the upper left sample location of the reference sub-block. For example, as shown in Table 2 above, the modified location may be derived based on a formula such as ((xColSb>>3)<<3, (yColSb>>3)<<3). Here, xColSb and yColSb represent the x and y coordinates of the upper left sample location of the reference sub-block, respectively, and >> represents an arithmetic right shift and << represents an arithmetic left shift.
[0234] Meanwhile, as described above, when deriving temporal motion information candidates based on sub-blocks, it can be seen that the motion vector for the reference sub-block is derived based on the position of the sub-block within the current block, and the base motion vector is derived based on the position of the current block. For example, as described with reference to Figures 16 to 19, for a current block having an 8x8 size, the motion vector for the reference sub-block and the base motion vector can be derived based on the center sample position of the bottom right edge of the current block. For a current block having a size larger than 8x8, the motion vector for the reference sub-block can be derived based on the center sample position of the bottom right edge of each sub-block within the current block, and the base motion vector can be derived based on the center sample position of the bottom right edge of the current block.
[0235] The decoding apparatus can generate predicted samples of the current block based on motion information for sub-blocks within the current block (S2330).
[0236] In one embodiment, for a prediction mode in which prediction is performed based on sub-block-level motion information for the current block (i.e., sbTMVP mode), the decoding device can generate predicted samples for the current block based on motion information for sub-blocks of the current block derived as described above.
[0237] The decoding device can generate reconstructed samples based on the predicted samples (S2340).
[0238] In one embodiment, the decoding device can use the predicted samples directly as reconstructed samples according to the prediction mode, or can generate reconstructed samples by adding residual samples to the predicted samples.
[0239] If residual samples for the current block exist, the decoding device may receive information about the residual for the current block. The information about the residual may include transform coefficients for the residual samples. The decoding device may derive residual samples (or a residual sample array) for the current block based on the residual information. The decoding device may generate reconstructed samples based on the prediction samples and the residual samples, and may derive a reconstructed block or picture based on the reconstructed samples. As described above, the decoding device may then apply an in-loop filtering procedure, such as a deblocking filtering and / or an SAO procedure, to the reconstructed picture, as needed, to improve subjective / objective image quality.
[0240] In the above-described embodiments, the method is described based on a flow chart with a series of steps or blocks, but the embodiments of this document are not limited to the order of steps, and certain steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flow chart are not exclusive, and other steps may be included, or one or more steps in the flow chart may be deleted without affecting the scope of this document.
[0241] The method according to the present document described above can be implemented in software form, and the encoding device and / or decoding device according to the present document can be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, or a display device.
[0242] When the embodiments herein are implemented in software, the methods described above may be implemented with modules (processes, functions, etc.) that perform the functions described above. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described herein may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the figures may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information on instructions) or algorithms for implementation may be stored on a digital storage medium.
[0243] In addition, decoding devices and encoding devices to which this document applies may be included in multimedia broadcast transmitting / receiving devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video interaction devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providing devices, Over-the-Top (OTT) video devices, Internet streaming service providing devices, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, image telephone video devices, transportation terminals (e.g., vehicles (including autonomous vehicles), airplane terminals, ship terminals, etc.), medical video devices, etc., and may be used to process video or data signals. For example, Over-the-Top (OTT) video devices may include game consoles, Blu-ray players, Internet-connected TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.
[0244] In addition, the processing method to which this document is applied can be produced in the form of a computer-executable program and stored in a computer-readable storage medium. Multimedia data having a data structure according to this document can also be stored in a computer-readable storage medium. The computer-readable storage medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. Examples of the computer-readable storage medium include Blu-ray Discs (BDs), Universal Serial Buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable storage medium also includes media embodied in the form of carrier waves (e.g., transmission via the Internet). A bitstream generated by the encoding method can be stored in a computer-readable storage medium or transmitted via a wired or wireless communication network.
[0245] Furthermore, the embodiments of the present document may be embodied in a computer program product by program code, which may be executed by a computer in accordance with the embodiments of the present document. The program code may be stored on a computer-readable carrier.
[0246] FIG. 24 illustrates an example of a content streaming system in which the embodiments disclosed herein can be applied.
[0247] Referring to FIG. 24, the content streaming system applied to the embodiments of this document can broadly include an encoding server, a streaming server, a web server, a media storage device (storage facility), a user device, and a multimedia input device.
[0248] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0249] The bitstream can be generated by an encoding method or a bitstream generation method applied to the embodiments of this document, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0250] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary that informs the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.
[0251] The streaming server may receive content from a media storage device and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a certain period of time to provide a smooth streaming service.
[0252] Examples of the user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), navigation systems, slate PCs, tablet PCs, ULTRABOOK (registered trademark), wearable devices such as smart watches (watch-type terminals), smart glasses (glass-type terminals), HMDs (Head Mounted Displays), digital TVs, desktop computers, and digital signatures.
[0253] Each server in the content streaming system can be operated as a distributed server, and in this case, data received by each server can be processed in a distributed manner.
[0254] The claims herein may be combined in various ways. For example, the technical features of the method claims herein may be combined and embodied in an apparatus, the technical features of the device claims herein may be combined and embodied in a method, the technical features of the method claims herein and the technical features of the device claims herein may be combined and embodied in an apparatus, and the technical features of the method claims herein and the technical features of the device claims herein may be combined and embodied in a method.
Claims
1. A video decoding method executed by a decoding device, comprising: deriving a co-located sub-block in a co-located picture for a sub-block in a current block; deriving sub-block based temporal motion information candidates based on the collocated sub-blocks; deriving motion information for the sub-blocks within the current block based on the sub-block-based candidate temporal motion information; generating predicted samples of the current block based on the motion information for the sub-blocks within the current block; generating reconstructed samples based on the predicted samples; applying deblock filtering to the reconstructed samples; based on the availability of the co-located sub-blocks, the positions of each of the co-located sub-blocks in the co-located picture are derived based on the positions of the sub-blocks in the current block; the position of each of the co-located sub-blocks within the co-located picture is derived by applying a motion shift to a bottom right-hand center sample position of each of the sub-blocks within the current block; the sub-block-based temporal motion information candidate includes a sub-block-based motion vector; A method wherein a base motion vector is used as a sub-block-wise motion vector for unavailable co-located sub-blocks among the co-located sub-blocks.
2. A video encoding method executed by an encoding device, comprising: deriving a co-located sub-block in a co-located picture for a sub-block in a current block; deriving sub-block based temporal motion information candidates based on the collocated sub-blocks; deriving motion information for the sub-blocks within the current block based on the sub-block-based candidate temporal motion information; generating predicted samples of the current block based on the motion information for the sub-blocks within the current block; deriving residual samples based on the prediction samples; encoding video information including information about the residual samples; based on the availability of the co-located sub-blocks, the positions of each of the co-located sub-blocks in the co-located picture are derived based on the positions of the sub-blocks in the current block; the position of each of the co-located sub-blocks within the co-located picture is derived by applying a motion shift to a bottom right-hand center sample position of each of the sub-blocks within the current block; the sub-block-based temporal motion information candidate includes a sub-block-based motion vector; A method wherein a base motion vector is used as a sub-block-wise motion vector for unavailable co-located sub-blocks among the co-located sub-blocks.
3. A method for transmitting data relating to video, comprising: obtaining a bitstream relating to the video, the bitstream comprising: deriving a co-located sub-block in a co-located picture for a sub-block in a current block; deriving sub-block based temporal motion information candidates based on the collocated sub-blocks; deriving motion information for the sub-blocks within the current block based on the sub-block-based candidate temporal motion information; generating predicted samples of the current block based on the motion information for the sub-blocks within the current block; deriving residual samples based on the prediction samples; encoding video information including information about the residual samples; transmitting the data including the bitstream; based on the availability of the co-located sub-blocks, the positions of each of the co-located sub-blocks in the co-located picture are derived based on the positions of the sub-blocks in the current block; the position of each of the co-located sub-blocks within the co-located picture is derived by applying a motion shift to a bottom right-hand center sample position of each of the sub-blocks within the current block; the sub-block-based temporal motion information candidate includes a sub-block-based motion vector; A method wherein a base motion vector is used as a sub-block-wise motion vector for unavailable co-located sub-blocks among the co-located sub-blocks.
Citation Information
Patent Citations
Moving image decoding device
WO2017195608A1
Method for processing video on basis of inter prediction mode and apparatus therefor
WO2018066927A1
Methods and apparatus of video coding using subblock-based temporal motion vector prediction
WO2020047289A1
Coordination method for sub-block based inter prediction
WO2020103940A1
System and method for signaling of motion merge modes in video coding
WO2020142448A1