Inter-frame prediction-based image or video compilation using SBTMVP
Through sbTMVP technology in video and image compilation, the time motion vector prediction of sub-blocks is used to solve the efficient compression and transmission problems of high-resolution images and videos, achieving more efficient compilation efficiency and simplified hardware implementation.
Patent Information
- Application Number
- CN202080050139.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-13
- Filing Date
- 2020-06-15
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2040-06-15
AI Technical Summary
The prior art is difficult to effectively compress and transmit high-resolution, high-quality video and image data, especially in immersive media such as virtual reality, artificial reality and holograms, resulting in increased transmission and storage costs.
采用基于子块的时间运动矢量预测(sbTMVP)技术,通过在当前子块的中心处定位四个样本之中的右下样本位置,导出参考子块的运动矢量,并在不可用参考子块的情况下使用基本运动矢量,提高帧间预测的效率。
Improves video and image compilation efficiency, reduces computational complexity, simplifies hardware implementation, and improves prediction performance.
Smart Images

Figure CN114080812B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to video or image coding, for example, an image or video coding technique based on inter-frame prediction using sub-block based temporal motion vector prediction (sbTMVP). Background Art
[0002] Recently, there has been an increasing demand for high-resolution and high-quality images and videos such as ultra-high-definition (HUD) images and 4K or 8K or larger videos in various fields. As image and video data becomes higher in resolution and higher in quality, the amount of information or the number of bits transmitted increases compared to existing image and video data. Therefore, if a medium such as an existing wired or wireless broadband line is used to transmit image data or an existing storage medium is used to store image and video data, the transmission cost and storage cost increase.
[0003] Furthermore, interest and demand for immersive media such as virtual reality (VR), artificial reality (AR) content, or holograms have recently increased, and broadcasts of images and videos having image characteristics different from those of real images, such as game images, have increased.
[0004] Therefore, in order to efficiently compress and transmit or store and play back information of high-resolution and high-quality images and videos having such various characteristics, efficient image and video compression technology is required.
[0005] In order to improve the efficiency of image / video coding, a sub-block-based temporal motion vector prediction technology is discussed. To this end, a solution is needed for efficiently performing the process of patching the motion vector of the sub-block unit in the sub-block-based temporal motion vector prediction. Summary of the Invention
[0006] Technical issues
[0007] The purpose of this document is to provide a method and apparatus for improving video / image coding efficiency.
[0008] Another object of this document is to provide a method and apparatus for efficient inter-frame prediction.
[0009] Yet another object of this document is to provide a method and apparatus for improving prediction performance by deriving sub-block based temporal motion vectors.
[0010] Another object of this document is to provide a method and apparatus for efficiently deriving corresponding positions of sub-blocks to derive sub-block-based temporal motion vectors.
[0011] Another object of this document is to provide a method and apparatus for unifying corresponding positions at a sub-coding block level and corresponding positions at a coding block level to derive a sub-block-based temporal motion vector.
[0012] Technical Solution
[0013] According to an embodiment of the present disclosure, a reference subblock for a current subblock may be derived based on a position of a sample located at the bottom right among four samples located at the center of the current subblock in subblock temporal motion vector prediction (sbTMVP).
[0014] According to an embodiment of the present disclosure, sbTMVP candidates can be derived based on the availability of reference subblocks for the current subblock; for available reference subblocks, the motion vectors of the available reference subblocks are derived as sbTMVP candidates, and for unavailable reference subblocks, the basic motion vectors are derived as sbTMVP candidates.
[0015] According to an embodiment of the present disclosure, an unavailable reference subblock may include a reference subblock located outside a reference picture or a reference subblock whose motion vector is unavailable, and for a reference subblock as an intra mode, an IBC (Intra Block Copy) mode, or a palette mode, the reference subblock may be a subblock whose motion vector is unavailable.
[0016] According to an embodiment of the present disclosure, a video / image decoding method performed by a decoding device is provided. The video / image decoding method may include the method disclosed in the embodiment of the present disclosure.
[0017] According to an embodiment of the present disclosure, a decoding device for performing video / image decoding is provided. The decoding device can execute the method disclosed in the embodiment of the present disclosure.
[0018] According to an embodiment of the present disclosure, a video / image encoding method performed by an encoding device is provided. The video / image encoding method may include the method disclosed in the embodiment of the present disclosure.
[0019] According to an embodiment of the present disclosure, a coding apparatus for performing video / image coding is provided. The coding apparatus can execute the method disclosed in the embodiment of the present disclosure.
[0020] According to an embodiment of the present disclosure, a computer-readable digital storage medium is provided, which stores encoded video / image information generated according to the video / image encoding method disclosed in at least one embodiment of the present disclosure.
[0021] According to an embodiment of the present disclosure, a computer-readable digital storage medium stores encoding information or encoded video / image information that enables a decoding device to perform the video / image decoding method disclosed in at least one embodiment of the present disclosure.
[0022] Beneficial effects
[0023] This document can have various effects. For example, it can improve overall image / video compression efficiency. In addition, through efficient inter-frame prediction, computational complexity can be reduced and overall coding efficiency can be improved. In addition, efficiency in terms of complexity and prediction performance can be improved because the corresponding positions of subblocks used to derive subblock-based temporal motion vectors are efficiently calculated in subblock-based temporal motion vector prediction (sbTMVP). In addition, because the methods of calculating corresponding positions at the subcoding block level and corresponding positions at the coding block level to derive subblock-based temporal motion vectors are unified, a simplified effect can be achieved in terms of hardware implementation.
[0024] The effects that can be achieved through the detailed embodiments of this document are not limited to the effects listed. For example, there may be various technical effects that a person of ordinary skill in the art can understand or derive from this document. Therefore, the detailed effects of this document are not limited to the effects explicitly described in this document, and may include various effects that can be understood or derived from the technical features of this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 Schematically illustrates an example of a video / image coding system to which embodiments of this document can be applied.
[0026] Figure 2 is a schematic diagram illustrating the configuration of a video / image encoding device to which the embodiments of this document can be applied.
[0027] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which the embodiments of this document can be applied.
[0028] Figure 4 An example of a video / image encoding method based on inter-frame prediction is shown, and Figure 5 The diagram schematically shows an example of an inter-frame prediction unit in the encoding device.
[0029] Figure 6 An example of a video / image decoding method based on inter-frame prediction is shown, and Figure 7 The diagram schematically shows an example of an inter-frame prediction unit in a decoding device.
[0030] Figure 8 The spatial neighboring blocks and the temporal neighboring blocks of the current block are exemplarily illustrated.
[0031] Figure 9 Temporal neighboring blocks used to derive sub-block-based temporal motion information candidates (sbTMVP candidates) are exemplarily illustrated.
[0032] Figure 10 is a diagram illustrating a process for deriving sub-block based temporal motion information candidates (sbTMVP candidates).
[0033] Figure 11 is a schematic diagram illustrating a method for calculating corresponding positions for deriving a default MV and a sub-block MV according to a block size during sbTMVP derivation.
[0034] Figure 12 is an exemplary view schematically illustrating a method for integrating corresponding positions for deriving a default MV and a sub-block MV according to a block size in an sbTMVP derivation process.
[0035] Figure 13 and Figure 14 is an exemplary view schematically illustrating a configuration of a pipeline for calculating corresponding positions for deriving a default MV and a sub-block MV according to a block size in an sbTMVP derivation process.
[0036] Figure 15 and Figure 16 An example of a video / image encoding method and related components according to one or more embodiments of the present disclosure is schematically shown.
[0037] Figure 17 and Figure 18 An example of a video / image decoding method and related components according to one or more embodiments of the present disclosure is schematically shown.
[0038] Figure 19 An example of a content streaming system to which the embodiments disclosed in this document can be applied is illustrated. DETAILED DESCRIPTION
[0039] This document can be modified in various ways and can have various embodiments, and specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are used to describe specific embodiments rather than to limit the technical spirit of this document. Unless otherwise clearly indicated in the context, singular expressions include plural expressions. Terms such as "including" or "having" in this specification should be understood to indicate the presence of characteristics, numbers, steps, operations, elements, components, or combinations thereof described in this specification, without excluding the possibility of the presence or addition of one or more characteristics, numbers, steps, operations, elements, components, or combinations thereof.
[0040] At the same time, in order to facilitate the description of different feature functions, the elements in the drawings described in this document are illustrated independently. This does not mean that each element is implemented as separate hardware or separate software. For example, at least two elements can be combined to form a single element, or a single element can be divided into multiple elements. Implementations in which elements are combined and / or separated are also included in the scope of the rights of this document unless it deviates from the essence of this document.
[0041] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to methods disclosed in Versatile Video Coding (VVC). In addition, the methods / embodiments disclosed in this document can be applied to methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second-generation Audio Video Coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).
[0042] This document proposes various embodiments of video / image coding, and unless mentioned to the contrary, these embodiments may be performed in combination with each other.
[0043] In this document, video may mean a collection of a series of images according to the passage of time. A picture generally means a unit that represents an image of a specific time period, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A tile is a rectangular area of a CTU within a specific tile column and a specific tile row in a picture. A tile column is a rectangular area where the height of the CTU is equal to the height of the picture and the width is specified by the syntax elements in the picture parameter set. A tile row is a rectangular area where the height of the CTU is specified by the syntax elements in the picture parameter set and the width is equal to the width of the picture. Tile scanning is a specific ordering of the CTUs of the following partitioned pictures: CTUs can be sorted continuously in a tile by a CTU raster scan, while tiles in a picture can be sorted continuously by a raster scan of the tiles of the picture. A slice includes an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that can be exclusively contained in a single NAL unit.
[0044] At the same time, a picture can be divided into two or more sub-pictures. A sub-picture can be a rectangular area of one or more slices within the picture.
[0045] A pixel or picture element (pel) may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a value of a pixel, and may represent only a pixel / pixel value of a luminance component, or only a pixel / pixel value of a chrominance component. Alternatively, a sample may refer to a pixel value in a spatial domain, or may refer to a transform coefficient in a frequency domain when a pixel value is transformed into a frequency domain.
[0046] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. A unit may include a luminance block and two chrominance (e.g., CB, CR) blocks. Depending on the situation, terms such as unit and block, region, etc. may be used interchangeably. In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0047] In addition, in this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or, for uniformity of expression, may still be referred to as a transform coefficient.
[0048] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled via residual coding syntax. The transform coefficients may be derived based on the residual information (or information about the transform coefficients), and the scaled transform coefficients may be derived by inversely transforming (scaling) the transform coefficients. Residual samples may be derived based on the inverse transform of the scaled transform coefficients. This may also be applied / expressed in other parts of this document.
[0049] In this document, the term "A or B" may mean "only A," "only B," or "both A and B." In other words, in this document, the term "A or B" may be interpreted to mean "A and / or B." For example, in this document, the term "A, B, or C" may mean "only A," "only B," "only C," or "any combination of A, B, and C."
[0050] As used in this document, a slash " / " or a comma can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0051] In this document, "at least one of A and B" may mean "only A", "only B", or "both A and B". In addition, in this document, the expression "at least one of A or B" or "at least one of A and / or B" may be interpreted as "at least one of A and B".
[0052] Furthermore, in this document, “at least one of A, B, and C” may mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.” Furthermore, “at least one of A, B, or C” or “at least one of A, B, and / or C” may mean “at least one of A, B, and C.”
[0053] Furthermore, brackets used in this document may mean "for example." Specifically, when the term "prediction (intra-frame prediction)" is used, it may indicate that "intra-frame prediction" is presented as an example of "prediction." In other words, the term "prediction" in this document is not limited to "intra-frame prediction" and may indicate that "intra-frame prediction" is presented as an example of "prediction." Furthermore, even when the term "prediction (i.e., intra-frame prediction)" is used, it may indicate that "intra-frame prediction" is presented as an example of "prediction."
[0054] In this document, technical features explained separately in one drawing may be implemented separately or may be implemented simultaneously.
[0055] Hereinafter, the preferred embodiment of this document will be described in more detail with reference to the accompanying drawings. Hereinafter, in the accompanying drawings, the same reference numerals are used in the same elements, and repeated description of the same elements may be omitted.
[0056] Figure 1 An example of a video / image coding system to which embodiments of this document can be applied is schematically illustrated.
[0057] refer to Figure 1 The video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit the encoded video / image information or data to the receiving device in the form of a file or stream transmission via a digital storage medium or a network.
[0058] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0059] The video source can obtain the video / image by capturing, synthesizing or generating the video / image process. The video source may include a video / image capturing device and / or a video / image generating device. The video / image capturing device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generating device may include, for example, a computer, a tablet computer and a smart phone, and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer or the like. In this case, the video / image capturing process may be replaced by a process that generates relevant data.
[0060] The encoding device can encode the input video / image. The encoding device can perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0061] The transmitter can transmit the encoded video / image information or data, output as a bitstream, to a receiver in a receiving device via a digital storage medium or network in the form of a file or streaming. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include components for generating a media file in a predetermined file format and can also include components for transmitting via a broadcast / communication network. The receiver can receive / extract the bitstream and transmit the received / extracted bitstream to a decoding device.
[0062] The decoding device can decode the video / image by performing a series of processes such as dequantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0063] The renderer can render the decoded video / image, and the rendered video / image can be displayed on a display.
[0064] Figure 2 is a diagram schematically illustrating a configuration of a video / image encoding device to which this document can be applied. Hereinafter, the encoding device may include an image encoding device and / or a video encoding device.
[0065] refer to Figure 2, the encoding device 200 may include an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 described above may be composed of one or more hardware components (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) and may be composed of a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.
[0066] The image splitter 210 may split the input image (or picture, or frame) input to the encoding device 200 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a largest coding unit (LCU), the coding units may be recursively split according to a quadtree, binary tree, ternary tree (QTBTTT) structure. For example, a coding unit may be divided into multiple coding units of increasing depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure may be applied first, followed by the binary tree structure and / or the ternary tree structure. Alternatively, the binary tree structure may be applied first. The coding process according to this document may be performed based on the final coding unit that has not been further split. In this case, based on coding efficiency according to image characteristics, the largest coding unit may be directly used as the final coding unit. Alternatively, the coding unit may be recursively split into coding units of increasing depth as needed, so that the optimally sized coding unit may be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or partitioned from the final coding unit described above. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal based on the transform coefficient.
[0067] Depending on the situation, terms such as unit and block, region, etc. may be used interchangeably. In general, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component. A sample may be used as a term corresponding to a pixel or a picture element (pel) of a picture (or image).
[0068] In the encoding device 200, the prediction signal (prediction block, prediction sample array) output from the inter predictor 221 or the intra predictor 222 is subtracted from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown, the unit in the encoder 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as a subtractor 231. The predictor can perform prediction on a processing target block (hereinafter referred to as a "current block") and can generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra prediction or inter prediction based on the current block or CU. As discussed later in the description of each prediction mode, the predictor can generate various information related to prediction, such as prediction mode information, and send the generated information to the entropy encoder 240. The information about the prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0069] The intra-frame predictor 222 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the reference sample can be located near the current block or separated from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The non-directional mode can include, for example, a DC mode and a planar mode. Depending on the level of detail of the prediction direction, the directional mode can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes can be used depending on the settings. The intra-frame predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0070] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information can include a motion vector and a reference picture index. It can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in a reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the neighboring block as the motion information of the current block. In skip mode, unlike merge mode, the residual signal cannot be sent. In the case of motion information prediction (motion vector prediction, MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0071] The predictor 220 can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra prediction or inter prediction to predict a block, and can also apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can perform prediction on the block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as games such as screen content coding (SCC). Although IBC basically performs prediction in the current picture, its execution is similar to inter prediction in that it derives a reference block in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be considered an example of intra coding or intra prediction. When the palette mode is applied, the sample values in the picture can be signaled based on information about the palette index and the palette table.
[0072] The prediction signal generated by the predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) can be used to generate a reconstruction signal or to generate a residual signal. The transformer 232 can generate a transform coefficient by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT means a transform obtained from a curve graph when the relationship information between pixels is represented by a curve graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to square pixel blocks of the same size, or can be applied to blocks of variable size rather than square blocks.
[0073] The quantizer 233 can quantize the transform coefficients and send them to the entropy encoder 240. The entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal in a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the quantized transform coefficients of the block type into a one-dimensional vector form based on the coefficient scanning order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 240 can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 240 can encode information required for video / image reconstruction in addition to the quantized transform coefficients (e.g., syntax element values, etc.) together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream on a unit basis of the network abstraction layer (NAL). The video / image information may also include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), and the like. Furthermore, the video / image information may also include general constraint information. In this document, information and / or syntax elements transmitted from the encoding device to the decoding device using a signal may be included in the video / image information. The video / image information may be encoded using the above-described encoding process and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. Here, the network may include a broadcast network, a communication network, and / or the like, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, and the like. A transmitter (not shown) that transmits the signal output from the entropy encoder 240 or a memory (not shown) that stores the signal may be configured as an internal / external component of the encoding device 200, or the transmitter may be included in the entropy encoder 240.
[0074] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transformation to the quantized transform coefficients via the dequantizer 234 and the inverse transformer 235, a residual signal (residual block or residual sample) can be reconstructed. The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222, so that a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) can be generated. When there is no residual for the processing target block, as in the case of applying skip mode, the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current picture, and as described later, can be used for inter-frame prediction of the next picture performed by filtering.
[0075] Furthermore, during the picture encoding and / or reconstruction process, luma mapping and chroma scaling (LMCS) may be applied.
[0076] The filter 260 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and the modified reconstructed picture can be stored in the memory 270, especially in the DPB of the memory 270. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. As discussed later in the description of each filtering method, the filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240. The information about filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0077] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-frame predictor 221. Accordingly, the encoding apparatus can avoid prediction mismatch in the encoding apparatus 100 and the decoding apparatus when applying inter-frame prediction, and can also improve encoding efficiency.
[0078] The memory 270DPB can store the modified reconstructed picture so that it can be used as a reference picture in the inter-frame predictor 221. The memory 270 can store the motion information of the blocks in the current picture from which the motion information has been derived (or encoded) and / or the motion information of the blocks in the reconstructed picture. The stored motion information can be sent to the inter-frame predictor 221 to be used as the motion information of the neighboring blocks or the motion information of the temporally neighboring blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-frame predictor 222.
[0079] Figure 3is a diagram schematically illustrating a configuration of a video / image decoding device to which this document can be applied. Hereinafter, a decoding device may include an image decoding device and / or a video decoding device.
[0080] refer to Figure 3 , the video decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be composed of one or more hardware components (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) and may be composed of a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0081] When a bit stream including video / image information is input, the decoding apparatus 300 can be used to decode the bit stream having been decoded. Figure 2 The image is reconstructed accordingly by processing the video / image information in the encoding device. For example, the decoding device 300 can derive the unit / block based on the information related to the block segmentation obtained from the bitstream. The decoding device 300 can perform decoding by using the processing unit applied in the encoding device. Therefore, the processing unit of decoding can be, for example, a coding unit, which can be divided from the coding tree unit or the maximum coding unit along the quadtree structure, the binary tree structure and / or the ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.
[0082] The decoding device 300 can receive the data from the Figure 2The received signal is output by the encoding device, and the entropy decoder 310 can decode the received signal. For example, the entropy decoder 310 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as the Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or general constraint information. In this document, the information and / or syntax elements transmitted / received using signals and / or the syntax elements described later can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on a coding method such as Exponential Golomb coding, CAVLC, or CABAC, and can output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residual. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, use the decoding target syntax element information and the decoding information of the neighboring and decoding target blocks or the information of the symbol / bin decoded in the previous step to determine the context model, predict the bin generation probability based on the determined context model, and perform arithmetic decoding on the bin to generate the symbol corresponding to each syntax element value. Here, after determining the context model, the CABAC entropy decoding method can update the context model using the symbol / bin information decoded by the context model for the next symbol / bin. The information about prediction among the information decoded in the entropy decoder 310 can be provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantized transform coefficients) and associated parameter information for which entropy decoding has been performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). In addition, the information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. In addition, a receiver (not shown) that receives a signal output from the encoding device may also constitute the decoding device 300 as an internal / external element, and the receiver may be a component of the entropy decoder 310. In addition, the decoding device according to this document may be referred to as a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of the dequantizer 321, the inverse transformer 322, the adder 340, the filter 350, the memory 360, the inter-frame predictor 332, and the intra-frame predictor 331.
[0083] The dequantizer 321 can output the transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into a two-dimensional block. In this case, the rearrangement can be performed based on the order of coefficient scanning performed in the encoding device. The dequantizer 321 can dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0084] The inverse transformer 322 obtains a residual signal (residual block, residual sample array) by performing inverse transformation on the transformation coefficients.
[0085] The predictor may perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310, and may determine a specific intra / inter prediction mode.
[0086] The predictor 320 can generate prediction signals based on various prediction methods. For example, the predictor can apply not only intra prediction or inter prediction to predict a block, but also apply both intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can perform prediction on a block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as games such as screen content coding (SCC). Although IBC essentially performs prediction in the current picture, its execution is similar to inter prediction in that it derives a reference block in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be considered an example of intra coding or intra prediction. When the palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.
[0087] The intra-frame predictor 331 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the reference samples can be located near the current block or separated from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 can determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.
[0088] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index for the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information about the prediction can include information indicating the mode used for inter-frame prediction of the current block.
[0089] The adder 340 adds the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (inter-frame predictor 332 or intra-frame predictor 331), so that a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) can be generated. When there is no residual for the processing target block as in the case of applying skip mode, the prediction block can be used as the reconstructed block.
[0090] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, may be output through filtering as described below, or may be used for inter prediction of the next picture.
[0091] In addition, luma mapping and chroma scaling (LMCS) can be applied to the picture decoding process.
[0092] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and store the modified reconstructed image in the memory 360, specifically, in the DPB of the memory 360. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0093] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block from which the motion information in the current picture is derived (or decoded) and / or the motion information of the block in the reconstructed picture. The stored motion information can be sent to the inter-frame predictor 332 to be used as the motion information of the spatially adjacent block or the motion information of the temporally adjacent block. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and transmit the reconstructed samples to the intra-frame predictor 331.
[0094] In the present disclosure, the embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 200 may be the same as or respectively correspond to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300. This can also be applied to the unit 332 and the intra-frame predictor 331.
[0095] As described above, when performing video coding, prediction is performed to improve compression efficiency. A prediction block including prediction samples for a current block, i.e., a target coding block, can be generated by prediction. In this case, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived similarly in the encoding device and the decoding device. The encoding device can improve image coding efficiency by signaling information (residual information) about the residual between the original block, rather than the original sample values of the original block themselves, and the prediction block, to the decoding device. The decoding device can derive a residual block including residual samples based on the residual information, can generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and can generate a reconstructed picture including the reconstructed block.
[0096] Residual information can be generated through a transformation and quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transformation process on the residual samples (residual sample array) included in the residual block, derive quantized transform coefficients by performing a quantization process on the transform coefficients, and can signal the relevant residual information to the decoding device (via a bitstream). In this case, the residual information may include information such as value information, position information, a transformation scheme, a transform kernel, and a quantization parameter for the quantized transform coefficients. The decoding device can perform a dequantization / inverse transformation process based on the residual information and derive residual samples (or residual blocks). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. In addition, the encoding device can derive a residual block for inter-frame prediction reference of a subsequent picture by dequantizing / inverse transforming the quantized transform coefficients, and can generate a reconstructed picture.
[0097] Meanwhile, as described above, when prediction is performed on the current block, intra prediction or inter prediction may be applied. Hereinafter, a case where inter prediction is applied to the current block will be described.
[0098] A predictor (more specifically, an inter-frame predictor) in an encoding / decoding device can derive prediction samples by performing inter-frame prediction on a block-by-block basis. Inter-frame prediction may refer to predictions derived using a method that relies on data elements (e.g., sample values or motion information) from pictures other than the current picture. When inter-frame prediction is applied to the current block, a prediction block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector in a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information for the current block can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter-frame prediction type (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). When inter-frame prediction is applied, neighboring blocks may include spatially neighboring blocks in the current picture and temporally neighboring blocks in reference pictures. The reference picture comprising the reference block and the reference picture comprising the temporally neighboring block may be the same or different. Temporally neighboring blocks may be referred to by names such as collocated reference blocks, collocated CUs (colCUs), and the reference picture including the temporally neighboring blocks may be referred to as collocated pictures (colPics). For example, a motion information candidate list may be configured based on neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled. Inter-frame prediction may be performed based on various prediction modes, and for example, in skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected neighboring block may be used as a motion vector predictor and the motion vector difference may be signaled. In this case, the motion vector of the current block may be derived by using the sum of the motion vector predictor and the motion vector difference.
[0099] The motion information may further include L0 motion information and / or L1 motion information according to the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). The L0-direction motion vector may be referred to as the L0 motion vector or MVL0, and the L1-direction motion vector may be referred to as the L1 motion vector or MVL1. Prediction based on the L0 motion vector may be referred to as L0 prediction, prediction based on the L1 motion vector may be referred to as L1 prediction, and prediction based on both the L0 motion vector and the L1 motion vector may be referred to as bi-prediction. Here, the L0 motion vector may indicate a motion vector associated with reference picture list L0, and the L1 motion vector may indicate a motion vector associated with reference picture list L1. Reference picture list L0 may include pictures preceding the current picture in output order, and reference picture list L1 may include pictures following the current picture in output order as reference pictures. The preceding picture may be referred to as a forward (reference) picture, and the subsequent picture may be referred to as a backward (reference) picture. Reference picture list L0 may further include pictures following the current picture in output order as reference pictures. In this case, the previous picture can be indexed first in the reference picture list L0, and then the subsequent picture can be indexed. The reference picture list L1 can further include pictures that precede the current picture in the output order as reference pictures. In this case, the subsequent picture can be indexed first in the reference picture list L1, and then the previous picture can be indexed. Here, the output order can correspond to the picture order count (POC) order.
[0100] In addition, various inter-frame prediction modes can be used to predict the current block in the picture. For example, various modes such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, merge with MVD (MMVD) mode and historical motion vector prediction (HMVP) mode can be used. Decoder-side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, dual prediction with CU-level weights (BCW), bidirectional optical flow (BDOF), etc. can be further used as additional modes. Affine mode can also be referred to as affine motion prediction mode. MVP mode can also be referred to as advanced motion vector prediction (AMVP) mode. In this document, some modes and / or motion information candidates derived from some modes can also be included in one of the motion information-related candidates in other modes. For example, an HMVP candidate can be added to the merge candidate of the merge / skip mode, or can also be added to the MVP candidate of the MVP mode. If the HMVP candidate is used as a motion information candidate for the merge mode or skip mode, the HMVP candidate can be referred to as an HMVP merge candidate.
[0101] Prediction mode information indicating the inter-frame prediction mode of the current block can be sent from the encoding device to the decoding device by signaling. In this case, the prediction mode information can be included in the bitstream and received by the decoding device. The prediction mode information may include index information indicating one of a plurality of candidate modes. Alternatively, the inter-frame prediction mode may be indicated by hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, whether the skip mode is applied may be indicated by signaling a skip flag, whether the merge mode is applied may be indicated by signaling a merge flag when the skip mode is not applied, and whether the MVP mode is applied or when the merge mode is not applied may be further signaled for additional distinction. The affine mode may be signaled as an independent mode or as a subordinate mode with respect to the merge mode or the MVP mode. For example, the affine mode may include an affine merge mode and an affine MVP mode.
[0102] In addition, when inter-frame prediction is applied to the current block, the motion information of the current block can be used. The encoding device can derive the best motion information for the current block through a motion estimation process. For example, the encoding device can search for a similar reference block with high correlation in units of fractional pixels within a predetermined search range in the reference picture using the original block in the original picture for the current block, and derive motion information through the searched reference block. The similarity of the blocks can be derived based on the difference in sample values based on the phase. For example, the similarity of the blocks can be calculated based on the sum of absolute differences (SAD) between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, the motion information can be derived based on the reference block with the smallest SAD in the search area. The derived motion information can be signaled to the decoding device according to various methods based on the inter-frame prediction mode.
[0103] A prediction block for the current block can be derived based on motion information derived according to an inter-prediction mode. The prediction block may include prediction samples (an array of prediction samples) for the current block. When the motion vector (MV) of the current block indicates fractional sample units, an interpolation process may be performed, and prediction samples for the current block may be derived based on reference samples in fractional sample units in a reference picture through interpolation. When affine inter-prediction is applied to the current block, prediction samples may be generated based on sample / sub-block units of MV. When bi-prediction is used, prediction samples derived by weighted summing or weighted averaging of prediction samples derived based on L0 prediction (i.e., prediction using reference pictures in reference picture list L0 and MVL0) and (depending on the phase) prediction based on L1 prediction (i.e., prediction using reference pictures in reference picture list L1 and MVL1) may be used as prediction samples for the current block. When bi-prediction is used, if the reference pictures used for L0 prediction and L1 prediction are located in different temporal directions based on the current picture (i.e., if the prediction corresponds to bi-prediction or bi-directional prediction), this may be referred to as true bi-prediction.
[0104] Reconstructed samples and reconstructed pictures may be generated based on the derived prediction samples, and thereafter, processes such as in-loop filtering may be performed as described above.
[0105] Figure 4 An example of a video / image encoding method based on inter-frame prediction is shown, and Figure 5 The diagram schematically illustrates one example of an inter prediction unit in the encoding apparatus. Figure 5 The inter-frame prediction unit in the encoding device can also be applied as Figure 2 The inter-frame prediction unit 221 of the encoding device 200 is the same as or corresponds to it.
[0106] refer to Figure 4 and Figure 5 , the encoding device performs inter-frame prediction on the current block (S400). The encoding device can derive the inter-frame prediction mode and motion information of the current block and generate prediction samples for the current block. Here, the inter-frame prediction mode determination process, motion information derivation process, and prediction sample generation process can be performed simultaneously, and any one process can be performed earlier than the other process.
[0107] For example, the inter-frame prediction unit 221 of the encoding device may include a prediction mode determination unit 221-1, a motion information derivation unit 221-2, and a prediction sample derivation unit 221-3. The prediction mode determination unit 221-1 may determine a prediction mode for the current block, the motion information derivation unit 221-2 may derive motion information for the current block, and the prediction sample derivation unit 221-3 may derive prediction samples for the current block. For example, the inter-frame prediction unit 221 of the encoding device may search for a block similar to the current block in a predetermined area (search area) of a reference picture through motion estimation, and derive a reference block whose difference with the current block is minimal or equal to or less than a predetermined standard. Based on this, a reference picture index indicating the reference picture in which the reference block is located may be derived, and a motion vector may be derived based on the position difference between the reference block and the current block. The encoding device may determine a mode to be applied to the current block from among various prediction modes. The encoding device may compare the RD costs of the various prediction modes and determine the optimal prediction mode for the current block.
[0108] For example, when skip mode or merge mode is applied to the current block, the encoding device may configure a merge candidate list (described below) and derive a reference block whose difference with the current block is the smallest or equal to or less than a predetermined standard among the reference blocks indicated by the merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to the decoding device. Motion information of the current block may be derived using the motion information of the selected merge candidate.
[0109] As another example, when the (A)MVP mode is applied to the current block, the encoding device may configure an (A)MVP candidate list and use the motion vector of a selected MVP candidate among the motion vector predictor (MVP) candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, a motion vector indicating a reference block derived by motion estimation may be used as the motion vector of the current block, and the MVP candidate with the motion vector having the smallest difference with the motion vector of the current block among the MVP candidates may become the selected MVP candidate. A motion vector difference (MVD) may be derived, which is the difference obtained by subtracting the MVP from the motion vector of the current block. In this case, information about the MVD may be signaled to the decoding device. In addition, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and separately signaled to the decoding device.
[0110] The encoding apparatus may derive residual samples based on the prediction samples (S410). The encoding apparatus may derive residual samples by comparing the initial samples with the prediction samples of the current block.
[0111] The encoding device encodes the image information including prediction information and residual information (S420). The encoding device can output the encoded image information in the form of a bitstream. The prediction information may include information about prediction mode information (e.g., a skip flag, a merge flag, or a mode index, etc.) and information about motion information as information related to the prediction process. The information about the motion information may include candidate selection information (e.g., a merge index, an mvp flag, or an mvp index), which is information for deriving a motion vector. In addition, the information about the motion information may include information about MVD and / or reference picture index information. In addition, the information about the motion information may include information indicating whether L0 prediction, L1 prediction, or dual prediction is applied. The residual information is information about the residual sample. The residual information may include information about the quantized transform coefficients used for the residual sample.
[0112] The output bitstream may be stored in a (digital) storage medium and transmitted to a decoding device or transmitted to a decoding device via a network.
[0113] At the same time, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is to derive the same prediction result as the prediction result performed by the decoding device, and thus, the coding efficiency can be improved. Therefore, the encoding device can store the reconstructed picture (or reconstructed sample or reconstructed block) in a memory and use the reconstructed picture as a reference picture. As described above, the in-loop filtering process can be further applied to the reconstructed picture.
[0114] Figure 6 An example of a video / image decoding method based on inter-frame prediction is shown, and Figure 7 The diagram schematically illustrates one example of an inter prediction unit in a decoding device. Figure 7 The inter-frame prediction unit in the decoding device can also be applied as Figure 3 The inter-frame prediction unit 332 of the decoding device 300 is the same as or corresponds to it.
[0115] refer to Figure 6 and Figure 7 The decoding device may perform an operation corresponding to the operation performed by the encoding device. The decoding device may perform prediction on the current block and derive a prediction sample based on the received prediction information.
[0116] Specifically, the decoding apparatus may determine a prediction mode of the current block based on the received prediction information (S600).The decoding apparatus may determine which inter prediction mode to apply to the current block based on prediction mode information in the prediction information.
[0117] For example, whether merge mode or (A)MVP mode is applied to the current block may be determined based on a merge flag. Alternatively, one of various inter-frame prediction mode candidates may be selected based on a mode index. Inter-frame prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include the various inter-frame prediction modes described above.
[0118] The decoding device derives the motion information of the current block based on the determined inter-frame prediction mode (S610). For example, when the skip mode or merge mode is applied to the current block, the decoding device may configure a merge candidate list and select a merge candidate from among the merge candidates included in the merge candidate list. Here, the selection may be performed based on the selection information (merge index). The motion information of the current block may be derived by using the motion information of the selected merge candidate. The motion information of the selected merge candidate may be used as the motion information of the current block.
[0119] As another example, when the (A)MVP mode is applied to the current block, the decoding device may configure an (A)MVP candidate list and use the motion vector of a selected motion vector predictor (MVP) candidate among the motion vector predictor (MVP) candidates included in the (A)MVP candidate list as the MVP of the current block. Here, the selection may be performed based on selection information (MVP flag or MVP index). In this case, the MVD of the current block may be derived based on the information about the MVD, and the motion vector of the current block may be derived based on the MVP of the current block and the MVD. In addition, the reference picture index of the current block may be derived based on the reference picture index information. The picture indicated by the reference picture index in the reference picture list of the current block may be derived as the reference picture referenced by the inter-frame prediction of the current block.
[0120] At the same time, the motion information of the current block can be derived without the candidate list configuration, and in this case, the motion information of the current block can be derived according to the process disclosed in the prediction mode. In this case, the candidate list configuration can be omitted.
[0121] The decoding apparatus may generate prediction samples for the current block based on the motion information of the current block (S620). In this case, a reference picture may be derived based on a reference picture index of the current block, and the prediction samples of the current block may be derived by using samples of the reference block indicated by the motion vector of the current block on the reference picture. In this case, in some cases, a prediction sample filtering process may be further performed for all or some of the prediction samples of the current block.
[0122] For example, the inter-frame prediction unit 332 of the decoding device may include a prediction mode determination unit 332-1, a motion information derivation unit 332-2 and a prediction sample derivation unit 332-3, and the prediction mode determination unit 332-1 can determine the prediction mode for the current block based on the received prediction mode information, the motion information derivation unit 332-2 can derive the motion information (motion vector and / or reference picture index) of the current block based on the information about the received motion information, and the prediction sample derivation unit 332-3 can derive the prediction sample of the current block.
[0123] The decoding device generates residual samples of the current block based on the received residual information (S630). The decoding device may generate reconstructed samples of the current block based on the predicted samples and the residual samples, and generate a reconstructed picture based on the generated reconstructed samples (S640). Thereafter, as described above, the in-loop filtering process may be further applied to the reconstructed picture.
[0124] As described above, the inter-frame prediction process may include an inter-frame prediction mode determination step, a motion information derivation step depending on the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information. The inter-frame prediction process may be performed by the encoding device and decoding device as described above.
[0125] Figure 8 The spatial neighboring blocks and the temporal neighboring blocks of the current block are exemplarily illustrated.
[0126] refer to Figure 8 The spatial neighboring blocks refer to neighboring blocks located around the current block 800 (which is a target for performing inter-frame prediction currently), and may include neighboring blocks located around the left side of the current block 800 or neighboring blocks located around the top of the current block 800. For example, the spatial neighboring blocks may include a lower left neighboring block, a left neighboring block, an upper right neighboring block, an upper neighboring block, and an upper left neighboring block of the current block 800. Figure 8 The spatially neighboring blocks are illustrated as "S".
[0127] According to an exemplary embodiment, the encoding apparatus / decoding apparatus may detect available neighboring blocks by searching for spatial neighboring blocks of a current block (e.g., a lower left neighboring block, a left neighboring block, an upper right neighboring block, an upper neighboring block, and an upper left neighboring block) according to a predetermined order, and derive motion information of the detected neighboring blocks as spatial motion information candidates.
[0128] A temporally neighboring block is a block located on a picture (i.e., a reference picture) other than the current picture including the current block 800, and refers to a collocated block of the current block 800 in the reference picture. Here, the reference picture may be before or after the current picture in terms of the picture order count (POC). In addition, a reference picture used to derive a temporally neighboring block may be referred to as a collocated reference picture or a col picture (collocated picture). In addition, a collocated block may refer to a block located at a position corresponding to the position of the current block 800 in the col picture, and is referred to as a col block. For example, Figure 8 As shown, the temporally neighboring blocks may include a col block positioned at a position corresponding to the lower right corner sample of the current block 800 in the reference picture (i.e., the col picture) (i.e., the col block including the lower right corner sample) and / or a col block positioned at a position corresponding to the center lower right sample of the current block 800 in the reference picture (i.e., the col picture) (i.e., the col block including the center lower right sample). Figure 8 A temporally adjacent block is illustrated as "T".
[0129] According to an exemplary embodiment, the encoding device / decoding device can detect available blocks by searching for temporally neighboring blocks of the current block (e.g., a col block including a lower right corner sample and a col block including a center lower right sample) in a predetermined order, and derive motion information of the detected blocks as temporal motion information candidates. As described above, the technique of using temporally neighboring blocks may be referred to as temporal motion vector prediction (TMVP). Furthermore, the temporal motion information candidates may be referred to as TMVP candidates.
[0130] At the same time, prediction can also be performed by deriving motion information in units of sub-blocks according to the inter-frame prediction mode. For example, in the affine mode or TMVP mode, motion information can be derived in units of sub-blocks. Specifically, the method for deriving temporal motion information candidates in units of sub-blocks can be referred to as sub-block-based temporal motion vector prediction (sbTMVP).
[0131] sbTMVP is a method of using the motion field within the col picture to improve the motion vector prediction (MVP) and merge mode of the coding unit within the current picture. The col picture of sbTMVP can be the same as the col picture used by TMVP. However, in TMVP, motion prediction is performed at the coding unit (CU) level. In contrast, in sbTMVP, motion prediction can be performed at the sub-block level or the sub-coding unit (sub-CU) level. In addition, in TMVP, temporal motion information is derived from the col block within the col picture (in this case, the col block is the col block corresponding to the lower right sample position of the current block or the center lower right sample position of the current block). In sbTMVP, temporal motion information is derived after applying motion shifting from the col picture. In this case, motion shifting can include the process of obtaining a motion vector from one of the spatially neighboring blocks of the current block and shifting by the motion vector.
[0132] Figure 9 Spatial neighboring blocks that may be used to derive sub-block-based temporal motion information candidates (sbTMVP candidates) are exemplarily illustrated.
[0133] refer to Figure 9 , the spatial neighboring block may include at least one of the lower left neighboring block A0, the left neighboring block A1, the upper right neighboring block B0 and the upper neighboring block B1 of the current block. In some cases, the spatial neighboring block may further include Figure 9 Another neighboring block other than the one shown, or may not include Figure 9 Furthermore, the spatially neighboring blocks may include only specific neighboring blocks, and for example, may include only the left neighboring block A1 of the current block.
[0134] For example, the encoding device / decoding device may first detect the motion vector of an available spatially neighboring block while searching for the spatially neighboring blocks in a predetermined search order, and may determine the block at the position indicated by the motion vector of the spatially neighboring block in the reference picture as the col block (i.e., the collocated reference block). In this case, the motion vector of the spatially neighboring block may be expressed as a temporal motion vector (temporal MV).
[0135] In this case, whether the spatially neighboring block is available can be determined based on its reference picture information, prediction mode information, position information, etc. For example, if the reference picture of the spatially neighboring block is the same as the reference picture of the current block, it can be determined that the corresponding spatially neighboring block is available. Alternatively, if the spatially neighboring block is coded in intra-frame prediction mode or is located outside the current picture / block, it can be determined that the corresponding spatially neighboring block is unavailable.
[0136] Furthermore, the search order of spatially neighboring blocks may be defined differently and may be in the order of, for example, A1, B1, B0, and A0. Alternatively, whether A1 is available may be determined by searching only A1.
[0137] Figure 10 is a diagram for schematically describing a process of deriving a sub-block-based temporal motion information candidate (sbTMVP candidate).
[0138] refer to Figure 10 , first, the encoding / decoding device can determine whether the spatial neighboring block (e.g., A1 block) of the current block is available. For example, if the reference picture of the spatial neighboring block (e.g., A1 block) uses the col picture, it can be determined that the spatial neighboring block (e.g., A1 block) is available, and the motion vector of the spatial neighboring block (e.g., A1 block) can be derived. In this case, the motion vector of the spatial neighboring block (e.g., A1 block) can be expressed as a temporal MV (tempMV), and the motion vector can be used in motion shifting. Alternatively, if it is determined that the spatial neighboring block (e.g., A1 block) is not available, the temporal MV (i.e., the motion vector of the spatial neighboring block) can be set to a zero vector. In other words, in this case, the motion vector set to (0,0) can be applied to motion shifting.
[0139] Next, the encoding / decoding apparatus may apply motion shifting based on the motion vector of the spatially adjacent block (e.g., block A1). For example, the motion shift may be shifted (e.g., A1') to the position indicated by the motion vector of the spatially adjacent block (e.g., block A1). That is, by applying motion shifting, the motion vector of the spatially adjacent block (e.g., block A1) may be added to the coordinates of the current block.
[0140] Next, the encoding / decoding apparatus may derive the collocated sub-blocks (col sub-blocks) of the motion shift on the col picture, and may obtain motion information (motion vector, reference index, etc.) for each col sub-block. For example, the encoding / decoding apparatus may derive each col sub-block on the col picture corresponding to the motion shift position at each sub-block position within the current block (i.e., the position indicated by the motion vector of the spatially adjacent block (e.g., A1)). Furthermore, the motion information of each col sub-block may be used as motion information for each sub-block of the current block (i.e., an sbTMVP candidate).
[0141] Furthermore, scaling can be applied to the motion vector of the col sub-block. Scaling can be performed based on the temporal distance difference between the reference picture of the col block and the reference picture of the current block. Therefore, scaling can be expressed as temporal motion scaling, and thus the reference picture of the current block and the reference picture of the temporal motion vector can be arranged. In this case, the encoding / decoding apparatus can obtain the scaled motion vector of the col sub-block as motion information for each sub-block of the current block.
[0142] In addition, in deriving the sbTMVP candidate, motion information may not exist in the col sub-block. In this case, for the col sub-block for which motion information does not exist, basic motion information (or default motion information) may be derived. The basic motion information may be used as the motion information for the sub-block of the current block. The basic motion information may be derived from a block located at the center of the col block (i.e., the col CU including the col sub-block). For example, motion information (e.g., a motion vector) may be derived from a block including a sample located at the bottom right of four samples located at the center of the col block and may be used as the basic motion information.
[0143] As described above, in the case of an affine mode or sbTMVP mode that derives motion information in sub-block units, affine merge candidates and sbTMVP candidates can be derived, and a sub-block-based merge candidate list can be configured based on these candidates. In this case, flag information indicating whether the affine mode or sbTMVP mode is enabled or disabled can be signaled. If the sbTMVP mode is enabled based on the flag information, the sbTMVP candidate derived as described above can be added to the first ranking of the sub-block-based merge candidate list. In addition, the affine merge candidate can be added to the next entry of the sub-block-based merge candidate list. In this case, the maximum number of candidates of the sub-block-based merge candidate list can be 5.
[0144] Furthermore, in the case of the sbTMVP mode, the sub-block size may be fixed, and may be fixed to, for example, 8×8 size. Furthermore, the sbTMVP mode may be applied only to blocks having both a width and a height equal to or greater than 8.
[0145] Meanwhile, in the current VVC standard, as shown in Table 1, sub-block based temporal motion information candidates (sbTMVP candidates) may be derived.
[0146] [Table 1]
[0147]
[0148]
[0149]
[0150]
[0151]
[0152]
[0153] When deriving sbTMVP candidates according to the method illustrated in Table 1, the default MV and the sub-block MV may be considered. In this case, the default MV may be referred to as temporal merging basic motion data or basic motion vector (basic motion information) based on the sub-block. Referring to Table 1, the default MV may correspond to ctrMV (or ctrMVLX) in Table 1. The sub-block MV may correspond to mvSbCol (or mvLXSbcol) in Table 1.
[0154] For example, if a subblock or subblock MV is available according to the sbTMVP derivation process, the subblock MV may be assigned to the corresponding subblock, or if the subblock or subblock MV is not available, a default MV may be used as the corresponding subblock MV for the corresponding subblock. In this case, the default MV may derive motion information from a position corresponding to the center pixel position of the corresponding block (i.e., col CU) on the col picture, and each subblock MV may derive motion information from the upper left position of the corresponding subblock (i.e., col subblock) on the col picture. In this case, the corresponding block (i.e., col CU) may be derived from the motion shift position based on the motion vector (i.e., temporal MV) of the spatially neighboring block A1, as described above in Figure 11 As described in.
[0155] Figure 11 is a diagram for schematically describing a method for calculating corresponding positions for deriving a default MV and a sub-block MV based on a block size in the sbTMVP derivation process.
[0156] Figure 11 Pixels (samples) drawn by dotted lines in indicate corresponding positions of each subblock used to derive each subblock MV, and pixels (samples) drawn by solid lines illustrate corresponding positions of the CU used to derive the default MV.
[0157] For example, reference Figure 11 (a), if the current block (i.e., the current CU) has an 8×8 size, the motion information of the sub-block may be derived based on the upper-left sample position within the sub-block having the 8×8 size, and the default motion information of the sub-block may be derived based on the center sample position within the current block (i.e., the current CU) having the 8×8 size.
[0158] Alternatively, for example, reference Figure 11(b), if the current block (i.e., the current CU) has a size of 16×8, the motion information of each sub-block can be derived based on the upper-left sample position within each sub-block having a size of 8×8, and the default motion information of each sub-block can be derived based on the center sample position within the current block (i.e., the current CU) having a size of 16×8.
[0159] Alternatively, for example, reference Figure 11 (c), if the current block (i.e., the current CU) has an 8×16 size, the motion information of each sub-block can be derived based on the upper-left sample position within each sub-block having the 8×8 size, and the default motion information of each sub-block can be derived based on the center sample position within the current block (i.e., the current CU) having the 8×16 size.
[0160] Alternatively, for example, reference Figure 11 (d), if the current block (i.e., the current CU) has a size of 16×16, the motion information of each sub-block can be derived based on the upper-left sample position within each sub-block having a size of 8×8, and the default motion information of each sub-block can be derived based on the center sample position within the current block (i.e., the current CU) having a size of 16×16.
[0161] from Figure 11 It can be seen that since the motion information of the sub-block tends to the upper left pixel position, there is a problem in that the sub-block MV is derived at a position far away from the position where the default MV indicating the representative motion information of the current CU is derived. As an example, in Figure 11 In the case of an 8×8 block shown in (a), one CU includes one sub-block, but there is a contradiction in that the sub-block MV and the default MV are represented as different motion information. In addition, since the methods of calculating the corresponding positions of the sub-block and the current CU block are different (i.e., the corresponding position used to derive the MV of the sub-block is the upper left sample position, and the corresponding position used to derive the default MV is the center sample position), an additional module may be required when implementing in hardware (H / W).
[0162] Therefore, in order to improve the problem, this document proposes a solution for unifying a method for deriving the corresponding position of the CU for the default MV and a method for deriving the corresponding position of the sub-block for each sub-block MV in the process of deriving sbTMVP candidates. According to the embodiment of this document, there is a unified effect in that, from the perspective of hardware (H / W), only one module for deriving each corresponding position based on the block size can be used. For example, since the method of calculating the corresponding position if the block size is a 16×16 block and the method of calculating the corresponding position if the block size is an 8×8 block can be implemented identically, there is a simplification effect from the hardware implementation aspect. In this case, a 16×16 block can represent a CU, and an 8×8 block can represent each sub-block.
[0163] As an embodiment, in deriving sbTMVP candidates, the center sample position may be used as a corresponding position for deriving motion information of a sub-block and a corresponding position for deriving default motion information, and may be implemented as shown in Table 2 below.
[0164] The following Table 2 is an illustration of an example of a method of deriving motion information of a sub-block and default motion information according to an embodiment of this document.
[0165] [Table 2]
[0166]
[0167]
[0168]
[0169]
[0170]
[0171]
[0172] Referring to Table 2, when deriving the sbTMVP candidate, the position of the current block (i.e., the current CU) including the subblock can be derived. The upper left sample position (xCtb, yCtb) of the coding tree block (or coding tree unit) including the current block and the lower right center sample position (xCtr, yCtr) of the current block can be derived as shown in equations (8-514) to (8-517) in Table 2. In this case, the positions (xCtb, yCtb) and (xCtr, yCtr) can be calculated based on the upper left sample position (xCb, yCb) of the current block relative to the upper left sample of the current picture.
[0173] In addition, a col block (i.e., col CU) on the col picture positioned corresponding to the current block (i.e., current CU) including the subblock may be derived. In this case, the position of the col block may be set to (xColCtrCb, yColCtrCb). The position may indicate the position of the col block, which includes the position (xCtr, yCtr) of the upper left sample of the col picture relative to the col picture within the col picture.
[0174] In addition, basic motion data (i.e., default motion information) for sbTMVP can be derived. The basic motion data may include a default MV (e.g., CtrMvLX). For example, a col block on a col picture may be derived. In this case, the position of the col block may be derived as (xColCb, yColCb). The position may be a position where a motion shift (e.g., tempMv) has been applied to the derived col block position (xColCtrCb, yColCtrCb). As described above, motion shifting may be performed by adding a motion vector (e.g., tempMv) derived from a spatially neighboring block (e.g., an A1 block) of the current block to the current col block position (xColCtrCb, yColCtrCb). Next, a default MV (e.g., ctrMvLX) may be derived based on the position (xColCb, yColCb) of the motion-shifted col block. In this case, the default MV (e.g., ctrMvLX) may represent a motion vector derived from a position corresponding to the lower-right center sample of the col block.
[0175] In addition, the col sub-block on the col picture corresponding to the sub-block in the current block (denoted as the current sub-block) can be derived. First, the position of each current sub-block can be derived. The position of each of the sub-blocks can be represented as (xSb, ySb). The position (xSb, ySb) can represent the position of the current sub-block based on the upper left sample of the current picture. For example, the position (xSb, ySb) of the current sub-block can be calculated as in equations (8-523) to (8-524) of Table 2, which can represent the lower right center sample position of the sub-block. Next, the position of each col sub-block in the col sub-block on the col picture can be derived. The position of each col sub-block can be represented as (xColSb, yColSb). The position (xColSb, yColSb) can be the position where the motion shift (e.g., tempMv) is applied to the position (xSb, ySb) of the current sub-block. As described above, motion shifting can be performed by adding a motion vector (e.g., tempMv) derived from a spatially neighboring block (e.g., A1 block) of the current block to the position (xSb, ySb) of the current subblock. Next, motion information (e.g., motion vector mvLXSbCol, flag availableFlagLXSbCol indicating availability) of the col subblock can be derived based on the position (xColSb, yColSb) of each of the motion-shifted col subblocks.
[0176] In this case, if a col subblock is unavailable among the col subblocks (for example, when availableFlagLXSbCol is 0), basic motion data (ie, default motion information) may be used for the unavailable col subblock. For example, a default MV (eg, ctrMvLX) may be used as a motion vector (eg, mvLXSbCol) for the unavailable col subblock.
[0177] Figure 12 is an example diagram for schematically describing a unified method for deriving corresponding positions of a default MV and a sub-block MV based on block size during sbTMVP derivation.
[0178] Figure 12 The pixels (samples) drawn with dotted lines in the figure indicate the corresponding positions within each sub-block for deriving the MV of each sub-block, and the pixels (samples) drawn with solid lines illustrate the corresponding positions of the CU for deriving the default MV.
[0179] For example, reference Figure 12 (a), if the current block (i.e., current CU) has an 8×8 size, the motion information can be derived from the col sub-block at the corresponding position on the col picture based on the lower right center sample position within the sub-block with an 8×8 size, and can be used as the motion information of the current sub-block. The motion information can be derived from the col block at the corresponding position on the col picture (i.e., col CU) based on the lower right center sample position within the current block (i.e., current CU) with an 8×8 size, and can be used as the default motion information of the current sub-block. In this case, as Figure 16 As shown, the motion information of the current sub-block and the default motion information can be derived from the same sample position (the same corresponding position).
[0180] Alternatively, for example, reference Figure 12 (b), if the current block (i.e., current CU) has a size of 16×8, the motion information can be derived from the col sub-block at the corresponding position on the col picture based on the lower-right center sample position within the sub-block with a size of 8×8, and can be used as the motion information of the current sub-block. The motion information can be derived from the col block at the corresponding position on the col picture (i.e., col CU) based on the lower-right center sample position within the current block (i.e., current CU) with a size of 16×8, and can be used as the default motion information of the current sub-block.
[0181] Alternatively, for example, reference Figure 12(c), if the current block (i.e., current CU) has an 8×16 size, the motion information may be derived from the col sub-block at the corresponding position on the col picture based on the lower-right center sample position within the sub-block having an 8×8 size, and may be used as the motion information of the current sub-block. The motion information may be derived from the col block at the corresponding position on the col picture (i.e., col CU) based on the lower-right center sample position within the current block (i.e., current CU) having an 8×16 size, and may be used as the default motion information of the current sub-block.
[0182] Alternatively, for example, reference Figure 12 (d), if the current block (i.e., current CU) has a size equal to or greater than 16×16, the motion information may be derived from the col sub-block at the corresponding position on the col picture based on the lower-right center sample position within the sub-block having an 8×8 size, and may be used as the motion information of the current sub-block. The motion information may be derived from the col block at the corresponding position on the col picture (i.e., col CU) based on the lower-right center sample position within the current block (i.e., current CU) having a size of 16×16 (or 16×16 or greater), and may be used as the default motion information of the current sub-block.
[0183] However, the aforementioned embodiments of this document are merely examples, and the default motion information and the motion information of the current sub-block may be derived based on another sample position (i.e., the lower right sample position) other than the center position. For example, the default motion information may be derived based on the upper left sample position of the current CU, and the motion information of the current sub-block may be derived based on the upper left sample position of the sub-block.
[0184] If the embodiments of this document are implemented as hardware, pipelines such as Figures 20 and 21 can be configured because the same H / W module can be used to derive motion information (temporal motion).
[0185] Figure 13 and Figure 14 is an example diagram schematically illustrating a configuration of a pipeline, through which corresponding positions for deriving a default MV and a sub-block MV can be unified and calculated in the sbTMVP derivation process.
[0186] refer to Figure 13 and Figure 14 , the corresponding position calculation module can calculate the corresponding positions for deriving the default MV and the sub-block MV. Figure 13 and Figure 14As shown, when the position (posX, posY) and block size (blkszX, blkszY) of the block are input to the corresponding position calculation module, the center position of the input block (i.e., the lower right sample position) can be output. When the position and block size of the current CU are input to the corresponding position calculation module, the center position of the col block on the col picture (i.e., the lower right sample position) can be output, that is, the corresponding position for deriving the default MV. Alternatively, when the position and block size of the current sub-block are input to the corresponding position calculation module, the center position of the col sub-block on the col picture (i.e., the lower right sample position) can be output, that is, the corresponding position for deriving the current sub-block MV.
[0187] As described above, when the corresponding position for deriving the default MV and the sub-block MV is output from the corresponding position calculation module, the motion vector (i.e., time mv) derived from the corresponding position can be patched. In addition, the sub-block-based temporal motion information (i.e., sbTMVP candidate) can be derived based on the patched motion vector (i.e., time mv). For example, as in Figure 13 and 14 In the embodiment, depending on the H / W implementation, the sbTMVP candidates can be derived in parallel based on the clock cycle, or can be derived sequentially.
[0188] The following drawings are written to describe the detailed examples of this document. The names of the detailed devices or detailed terms or names (for example, grammatical names / grammatical names) written in the drawings are exemplary, and therefore the technical features of this document are not limited to the detailed names used in the following drawings.
[0189] Figure 15 and Figure 16 An example of a video / image encoding method and related components according to an embodiment of the present disclosure is schematically shown.
[0190] Figure 15 The method disclosed in Figure 2 Specifically, Figure 15 Steps S1500 to S1540 can be performed by Figure 2 The predictor 220 disclosed in (more specifically, the inter-frame predictor 221) performs, Figure 15 Step S1550 can be performed by Figure 2 The residual processor 230 disclosed in Figure 15 Step S1560 can be performed by Figure 2 The entropy encoder 240 disclosed in is executed. In addition, Figure 15 The method disclosed in may include the above-mentioned embodiments of the present disclosure. Figure 15 In the present invention, any redundant detailed description of the embodiments will be omitted or briefly described.
[0191] refer to Figure 15 , the encoding apparatus may derive the positions of the sub-blocks included in the current block ( S1500 ).
[0192] Here, the current block may be referred to as a current coding unit CU or a current coding block CB, and a subblock included in the current block may be referred to as a current coding subblock.
[0193] In an embodiment, the encoding apparatus may derive the position of the current subblock in the current block.
[0194] For example, the encoding device may derive the position of the current subblock on the current picture based on the center sample position of the current subblock. In this case, the center sample position may represent the position of the lower right center sample located at the lower right among the four samples located at the center.
[0195] Meanwhile, the upper-left sample position used in the present disclosure may be referred to as a top-left sample position or an upper-left sample position, etc., and the lower-right center sample position may be referred to as a below-right center sample position, a center-lower-right sample position, a bottom-right center sample position or a center-bottom-right sample position, etc.
[0196] The encoding apparatus may derive a reference subblock on a collocated reference picture for a subblock within a current block ( S1510 ).
[0197] Here, the collocated reference picture refers to a reference picture used to derive the temporal motion information (ie, sbTMVP) as described above, and may represent the col picture described above. The reference subblock may represent the col subblock described above.
[0198] In an embodiment, the encoding device may derive a reference sub-block on a collocated reference picture based on the position of the current sub-block within the current block. For example, the encoding device may derive a reference sub-block on a collocated reference picture based on the center sample position (e.g., the bottom right center sample position) of the current sub-block.
[0199] For example, the encoding device may first specify the position of the current block and then specify the position of the sub-block within the current block. As explained with reference to Table 2 above, the position of the current block may be represented based on the upper left sample position (xCtb, yCtb) of the coding tree block and the lower right center sample position (xCtr, yCtr) of the current block. The position of the current sub-block within the current block may be represented as (xSb, ySb), and the position (xSb, ySb) may represent the lower right center sample position of the current sub-block. Here, the lower right center sample position (xSb, ySb) of the sub-block may be calculated based on the upper left sample position of the sub-block and the sub-block size, and may be calculated as in Equations 8-523 and 8-524 of Table 2 above.
[0200] In addition, the encoding apparatus may derive a reference subblock on a collocated reference picture based on the lower-right center sample position of the current subblock within the current block. As explained with reference to Table 2 above, the reference subblock may be represented as a position (xColSb, yColSb) on a collocated reference picture, and the position (xColSb, yColSb) on the collocated reference picture may be derived based on the lower-right center sample position (xSb, ySb) of the current subblock within the current block.
[0201] In addition, motion shifting may be applied when deriving the reference subblock. The encoding apparatus may perform motion shifting based on a motion vector derived from a spatially neighboring block of the current block. The spatially neighboring block of the current block may be a left neighboring block located on the left side of the current block (e.g., Figure 9 and Figure 10 ). In this case, if the left neighboring block (e.g., the A1 block) is available, a motion vector can be derived from the left neighboring block, or if the left neighboring block is not available, a zero vector can be derived. Here, the availability of the spatial neighboring block can be determined by reference picture information, prediction mode information, position information, etc. of the spatial neighboring block. For example, if the reference picture of the spatial neighboring block is the same as the reference picture of the current block, it can be determined that the spatial neighboring block is available. Alternatively, if the spatial neighboring block is coded in intra-frame prediction mode or the spatial neighboring block is located outside the current picture / tile, it can be determined that the spatial neighboring block is unavailable.
[0202] That is, the encoding apparatus may apply motion shifting (i.e., the motion vector of the spatially neighboring block (e.g., the A1 block)) to the lower right center sample position (xSb, ySb) of the current subblock within the current block, and may derive the reference subblock on the collocated reference picture based on the motion shifted position. In this case, the position (xColSb, yColSb) of the reference subblock may be expressed as a position obtained by motion shifting from the lower right center sample position (xSb, ySb) of the current subblock within the current block to the position indicated by the motion vector of the spatially neighboring block (e.g., the A1 block), and may be calculated as in Equations 8-525 and 8-526 of Table 2 above.
[0203] The encoding apparatus may derive sbTMVP (sub-block temporal motion vector predictor) candidates based on the reference sub-block ( S1520 ).
[0204] Meanwhile, in the present disclosure, the sbTMVP candidate may be replaced or used interchangeably with a sub-block-based temporal motion information candidate or a sub-block unit temporal motion information candidate, a sub-block-based temporal motion vector predictor candidate, etc. That is, if motion information is derived for each sub-block to perform prediction as described above, the sbTMVP candidate may be derived, and motion prediction may be performed at the sub-block level (or sub-coding unit (sub-CU) level) based on the sbTMVP candidate.
[0205] In an embodiment, the encoding device may derive an sbTMVP candidate based on the motion vector of the reference subblock. For example, the encoding device may derive an sbTMVP candidate based on the motion vector of the reference subblock derived based on whether the reference subblock is available. If the reference subblock is available, the motion vector of the available reference subblock may be derived as the sbTMVP candidate. If the reference subblock is not available, the base motion vector may be derived as the sbTMVP candidate.
[0206] Here, the basic motion vector may correspond to the above-mentioned default motion vector and may be derived on a collocated reference picture based on the position of the current block. In this case, the position of the current block may be derived based on the center sample position within the current block (e.g., the bottom right center sample position).
[0207] In deriving the base motion vector, in an embodiment, the encoding device may specify the position of a reference coding block on a collocated reference picture based on the lower-right center sample position of the current block, and derive the base motion vector based on the position of the reference coding block. The reference coding block may refer to a col block located on a collocated reference picture corresponding to the current block including the subblock. As explained with reference to Table 2 above, the position of the reference coding block may be represented as (xColCtrCb, yColCtrCb), and the position (xColCtrCb, yColCtrCb) may represent the position of the reference coding block that covers the position (xCtr, yCtr) within the collocated reference picture relative to the upper-left sample of the collocated reference picture. The position (xCtr, yCtr) may represent the lower-right center sample position of the current block.
[0208] Furthermore, in deriving the base motion vector, motion shifting may be applied to the position (xColCtrCb, yColCtrCb) of the reference coding block. Motion shifting may be performed by adding the motion vector derived from a spatially neighboring block (e.g., the A1 block) of the current block as described above to the position (xColCtrCb, yColCtrCb) of the reference coding block covering the bottom-right center sample. The encoding apparatus may derive the base motion vector based on the position (xColCb, yColCb) of the motion-shifted reference coding block. That is, the base motion vector may be a motion vector derived from a motion-shifted position on a collocated reference picture based on the bottom-right center sample position of the current block.
[0209] Meanwhile, the availability of a reference subblock can be determined based on whether the reference subblock is located outside a collocated reference picture or based on a motion vector. For example, an unavailable reference subblock may include a reference subblock located outside a collocated reference picture or a reference subblock for which a motion vector is unavailable. For example, if the reference subblock is based on intra mode, IBC (Intra Block Copy) mode, or palette mode, the reference subblock may be a subblock for which a motion vector is unavailable. Alternatively, if the reference coding block covering the modified position derived based on the position of the reference subblock is based on intra mode, IBC mode, or palette mode, the reference subblock may be a subblock for which a motion vector is unavailable.
[0210] In this case, as an embodiment, the motion vector of the available reference subblock may be derived based on the motion vector of the block covering the modified position derived based on the upper left sample position of the reference subblock. For example, as shown in Table 2 above, the modified position may be derived by the equation ((xColSb>>3)<<3, (yColSb>>3)<<3). Here, xColSb and yColSb may represent the x-coordinate and y-coordinate of the upper left sample position of the reference subblock, respectively, and >> may represent an arithmetic right shift, and << may represent an arithmetic left shift.
[0211] Meanwhile, as described above, in deriving the sbTMVP candidate, it can be seen that the motion vector for the reference sub-block is derived based on the position of the sub-block within the current block, and the basic motion vector is derived based on the position of the current block. Figure 12 As shown in FIG, for a current block of size 8×8, the motion vector for the reference subblock and the basic motion vector may be derived based on the lower right center sample position of the current block. For a current block of size greater than 8×8, the motion vector for the reference subblock may be derived based on the lower right center sample position of the subblock within the current block, and the basic motion vector may be derived based on the lower right center sample position of the current block.
[0212] The encoding apparatus may derive motion information for a subblock within the current block based on the sbTMVP candidate ( S1530 ).
[0213] In an embodiment, the encoding device may derive the motion vector of the reference subblock as the motion information (e.g., motion vector) of the current subblock within the current block. As described above, the encoding device may derive an sbTMVP candidate based on the motion vector of the available reference subblock or the basic motion vector, and may use the motion vector derived as the sbTMVP candidate as the motion vector for the current subblock.
[0214] The encoding apparatus may generate a prediction sample of the current block based on motion information of a subblock within the current block ( S1540 ).
[0215] In an embodiment, the encoding device may generate prediction samples based on the motion vector of the current sub-block. Specifically, the encoding device may select the best motion information based on the RD (rate-distortion) cost and generate prediction samples based on the information. For example, if the motion information derived for each sub-block of the current block (i.e., sbTMVP) is selected as the best motion information, the encoding device may generate prediction samples for the current block based on the motion information derived for the sub-blocks of the current block.
[0216] The encoding apparatus may generate information about residual samples derived based on the prediction samples ( S1550 ), and may encode image information including the information about the residual samples ( S1560 ).
[0217] That is, the encoding device may derive residual samples based on the initial samples of the current block and the predicted samples of the current block. In addition, the encoding device may generate information about the residual samples. Here, the information about the residual samples may include information derived by performing transformation and quantization on the residual samples (such as value information, position information, transformation scheme, transformation kernel, and quantization parameter of quantized transformation coefficients).
[0218] The encoding device may encode information about the residual samples and output it as a bitstream, and may transmit it to the decoding device through a network or a storage medium.
[0219] Figure 17 and Figure 18 An example of a video / image decoding method and related components according to an embodiment of the present disclosure is schematically shown.
[0220] Figure 17 The method disclosed in Figure 3 Specifically, Figure 17 Steps S1700 to S1740 can be performed by Figure 3 The predictor 330 disclosed in (more specifically, the inter-frame predictor 332) is executed, and Figure 17 Step S1750 can be performed by Figure 3 In addition, Figure 17 The method disclosed in may include the above-mentioned embodiments of the present disclosure. Figure 17 In the present invention, any redundant detailed description of the embodiments will be omitted or briefly described.
[0221] refer to Figure 17 , the decoding apparatus may derive the position of the sub-block included in the current block ( S1700 ).
[0222] Here, the current block may be referred to as a current coding unit CU or a current coding block CB, and a subblock included in the current block may be referred to as a current coding subblock.
[0223] In an embodiment, the decoding apparatus may derive the position of the current sub-block in the current block.
[0224] For example, the decoding device may derive the position of the current subblock on the current picture based on the center sample position of the current subblock. In this case, the center sample position may represent the position of the lower right center sample located at the lower right among the four samples located at the center.
[0225] Meanwhile, the upper-left sample position used in the present disclosure may be referred to as a top-left sample position or an upper-left sample position, etc., and the lower-right center sample position may be referred to as a below-right center sample position, a center-lower-right sample position, a bottom-right center sample position or a center-bottom-right sample position, etc.
[0226] The decoding apparatus may derive a reference subblock on a collocated reference picture for a subblock within a current block ( S1710 ).
[0227] Here, the collocated reference picture refers to a reference picture used to derive the temporal motion information (ie, sbTMVP) as described above, and may represent the col picture described above. The reference sub-block may represent the col sub-block described above.
[0228] In an embodiment, the decoding device may derive a reference sub-block on a collocated reference picture based on the position of the current sub-block within the current block. For example, the decoding device may derive a reference sub-block on a collocated reference picture based on the center sample position (e.g., the bottom right center sample position) of the current sub-block.
[0229] For example, the decoding device may first specify the position of the current block and then specify the position of the sub-block within the current block. As explained with reference to Table 2 above, the position of the current block may be represented based on the upper left sample position (xCtb, yCtb) of the coding tree block and the lower right center sample position (xCtr, yCtr) of the current block. The position of the current sub-block within the current block may be represented as (xSb, ySb), and the position (xSb, ySb) may represent the lower right center sample position of the current sub-block. Here, the lower right center sample position (xSb, ySb) of the sub-block may be calculated based on the upper left sample position of the sub-block and the sub-block size, and may be calculated as in Equations 8-523 and 8-524 of Table 2 above.
[0230] In addition, the decoding apparatus may derive a reference subblock on a collocated reference picture based on the lower-right center sample position of the current subblock within the current block. As explained with reference to Table 2 above, the reference subblock may be represented as a position (xColSb, yColSb) on a collocated reference picture, and the position (xColSb, yColSb) on the collocated reference picture may be derived based on the lower-right center sample position (xSb, ySb) of the current subblock within the current block.
[0231] In addition, motion shifting may be applied when deriving the reference subblock. The decoding apparatus may perform motion shifting based on a motion vector derived from a spatially neighboring block of the current block. The spatially neighboring block of the current block may be a left neighboring block located on the left side of the current block (e.g., Figure 9 and Figure 10 ). In this case, if the left neighboring block (e.g., the A1 block) is available, a motion vector can be derived from the left neighboring block, or if the left neighboring block is not available, a zero vector can be derived. Here, the availability of the spatial neighboring block can be determined by reference picture information, prediction mode information, position information, etc. of the spatial neighboring block. For example, if the reference picture of the spatial neighboring block is the same as the reference picture of the current block, it can be determined that the spatial neighboring block is available. Alternatively, if the spatial neighboring block is coded in intra-frame prediction mode or the spatial neighboring block is located outside the current picture / tile, it can be determined that the spatial neighboring block is unavailable.
[0232] That is, the decoding apparatus may apply motion shifting (i.e., the motion vector of the spatially neighboring block (e.g., the A1 block)) to the lower right center sample position (xSb, ySb) of the current subblock within the current block, and may derive the reference subblock on the collocated reference picture based on the motion shifted position. In this case, the position (xColSb, yColSb) of the reference subblock may be expressed as a position obtained by motion shifting from the lower right center sample position (xSb, ySb) of the current subblock within the current block to the position indicated by the motion vector of the spatially neighboring block (e.g., the A1 block), and may be calculated as in Equations 8-525 and 8-526 of Table 2 above.
[0233] The decoding apparatus may derive a sbTMVP (sub-block temporal motion vector predictor) candidate based on the reference sub-block ( S1720 ).
[0234] Meanwhile, in the present disclosure, the sbTMVP candidate may be replaced or used interchangeably with a sub-block-based temporal motion information candidate or a sub-block unit temporal motion information candidate, a sub-block-based temporal motion vector predictor candidate, etc. That is, if motion information is derived for each sub-block to perform prediction as described above, the sbTMVP candidate may be derived, and motion prediction may be performed at the sub-block level (or sub-coding unit (sub-CU) level) based on the sbTMVP candidate.
[0235] In an embodiment, the decoding device may derive an sbTMVP candidate based on the motion vector of the reference sub-block. For example, the decoding device may derive an sbTMVP candidate based on the motion vector of the reference sub-block derived based on whether the reference sub-block is available. If the reference sub-block is available, the motion vector of the available reference sub-block may be derived as the sbTMVP candidate. If the reference sub-block is unavailable, the base motion vector may be derived as the sbTMVP candidate.
[0236] Here, the basic motion vector may correspond to the above-mentioned default motion vector and may be derived on a collocated reference picture based on the position of the current block. In this case, the position of the current block may be derived based on the center sample position within the current block (e.g., the bottom right center sample position).
[0237] In deriving the base motion vector, in an embodiment, the decoding device may specify the position of a reference coding block on a collocated reference picture based on the lower-right center sample position of the current block, and derive the base motion vector based on the position of the reference coding block. The reference coding block may refer to a col block located on the collocated reference picture corresponding to the current block including the subblock. As explained with reference to Table 2 above, the position of the reference coding block may be represented as (xColCtrCb, yColCtrCb), and the position (xColCtrCb, yColCtrCb) may represent the position of the reference coding block that covers the position (xCtr, yCtr) within the collocated reference picture relative to the upper-left sample of the collocated reference picture. The position (xCtr, yCtr) may represent the lower-right center sample position of the current block.
[0238] Furthermore, in deriving a base motion vector, motion shifting may be applied to the position (xColCtrCb, yColCtrCb) of the reference coding block. Motion shifting may be performed by adding the motion vector derived from a spatially neighboring block (e.g., the A1 block) of the current block as described above to the position (xColCtrCb, yColCtrCb) of the reference coding block covering the bottom-right center sample. The decoding device may derive the base motion vector based on the position (xColCb, yColCb) of the motion-shifted reference coding block. That is, the base motion vector may be a motion vector derived from a motion-shifted position on a collocated reference picture based on the bottom-right center sample position of the current block.
[0239] Meanwhile, the availability of a reference subblock can be determined based on whether the reference subblock is located outside a collocated reference picture or based on a motion vector. For example, an unavailable reference subblock may include a reference subblock located outside a collocated reference picture or a reference subblock for which a motion vector is unavailable. For example, if the reference subblock is based on intra mode, IBC (Intra Block Copy) mode, or palette mode, the reference subblock may be a subblock for which a motion vector is unavailable. Alternatively, if the reference coding block covering the modified position derived based on the position of the reference subblock is based on intra mode, IBC mode, or palette mode, the reference subblock may be a subblock for which a motion vector is unavailable.
[0240] In this case, as an embodiment, the motion vector of the available reference subblock may be derived based on the motion vector of the block covering the modified position derived based on the upper left sample position of the reference subblock. For example, as shown in Table 2 above, the modified position may be derived by the equation ((xColSb>>3)<<3, (yColSb>>3)<<3). Here, xColSb and yColSb may represent the x-coordinate and y-coordinate of the upper left sample position of the reference subblock, respectively, and >> may represent an arithmetic right shift, and << may represent an arithmetic left shift.
[0241] Meanwhile, as described above, in deriving the sbTMVP candidate, it can be seen that the motion vector of the reference sub-block is derived based on the position of the sub-block within the current block, and the basic motion vector is derived based on the position of the current block. Figure 12 As shown, for a current block of size 8×8, the motion vector for the reference subblock and the base motion vector can be derived based on the lower right center sample position of the current block. For a current block of size greater than 8×8, the motion vector for the reference subblock can be derived based on the lower right center sample position of the subblock within the current block, and the base motion vector can be derived based on the lower right center sample position of the current block.
[0242] The decoding apparatus may derive motion information for a subblock within the current block based on the sbTMVP candidate ( S1730 ).
[0243] In an embodiment, the decoding device may derive the motion vector of the reference subblock as the motion information (e.g., motion vector) of the current subblock within the current block. As described above, the decoding device may derive an sbTMVP candidate based on the motion vector of the available reference subblock or the basic motion vector, and may use the motion vector derived as the sbTMVP candidate as the motion vector for the current subblock.
[0244] The decoding apparatus may generate a prediction sample of the current block based on a motion vector for the current subblock within the current block ( S1740 ).
[0245] In an embodiment, in the case of a prediction mode (ie, sbTMVP mode) in which prediction is performed based on sub-block unit motion information for the current block, the decoding device may generate a prediction sample of the current block based on the motion information for the current sub-block derived above.
[0246] The decoding apparatus may generate reconstructed samples based on the predicted samples ( S1750 ).
[0247] In an embodiment, the decoding apparatus may use the prediction samples directly as the reconstructed samples, or may generate the reconstructed samples by adding the residual samples to the prediction samples, according to the prediction mode.
[0248] If there are residual samples for the current block, the decoding device may receive information about the residual for the current block. The information about the residual may include transform coefficients related to the residual samples. The decoding device may derive residual samples (or residual sample arrays) for the current block based on the residual information. The decoding device may generate reconstructed samples based on the predicted samples and the residual samples, and derive a reconstructed block or a reconstructed picture based on the reconstructed samples. The decoding device may then apply an in-loop filtering process such as a deblocking filter and / or an SAO process to the reconstructed picture as described above to improve subjective / objective image quality when necessary.
[0249] In the above-mentioned embodiments, although these methods have been described based on flowcharts in the form of a series of steps or units, the embodiments of this document are not limited to the order of these steps, and some of these steps may be performed in an order different from that of other steps or may be performed simultaneously with other steps. In addition, it will be understood by those skilled in the art that the steps shown in the flowcharts are not exclusive, and without affecting the scope of rights of this document, these steps may include additional steps or one or more steps in the flowcharts may be deleted.
[0250] The above-mentioned method according to this document may be implemented in software form, and the encoding device and / or decoding device according to this document may be included in an apparatus for performing image processing, such as a TV, computer, smart phone, set-top box, or display device.
[0251] In this document, when the embodiments are implemented in software form, the above-mentioned methods can be implemented as modules (programs, functions, etc.) for performing the above-mentioned functions. The modules can be stored in a memory and executed by a processor. The memory can be arranged inside or outside the processor and connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, a microprocessor, a controller or a chip. For example, the functional units illustrated in the accompanying drawings can be implemented and executed on a computer, a processor, a microprocessor, a controller or a chip. In this case, the information (e.g., information about instructions) or the algorithm used for such implementation can be stored in a digital storage medium.
[0252] In addition, the decoding device and encoding device to which this document is applied may be included in multimedia broadcast transmission and reception devices, mobile communication terminals, home theater video devices, digital theater video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video on demand (VoD) service providing devices, over-the-top (OTT) video devices, Internet streaming service providing devices, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, video phone video devices, transportation terminals (e.g., vehicle (including autonomous vehicle) terminals, aircraft terminals, and ship terminals), and medical video devices, and may be used to process video signals or data signals. For example, over-the-top (OTT) video devices may include game consoles, Blueray players, Internet access TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0253] In addition, the processing method of applying this document can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. The multimedia data with a data structure according to this document can also be stored in a computer-readable recording medium. Computer-readable recording media include all kinds of storage devices storing computer-readable data. Computer-readable recording media may include, for example, Blueray discs (BD), universal serial buses (USB), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. In addition, computer-readable recording media include media implemented in the form of carrier waves (for example, transmitted over the Internet). In addition, the bit stream generated using the encoding method can be stored in a computer-readable recording medium or can be transmitted through wired and wireless communication networks.
[0254] In addition, the embodiments of this document may be implemented as a computer program product using program code. The program code may be executed by a computer according to the embodiments of this document. The program code may be stored on a carrier wave that can be read by a computer.
[0255] Figure 19 An example of a content streaming system to which the embodiments disclosed in this document can be applied is illustrated.
[0256] refer to Figure 19 The content streaming media system to which the embodiment of the present invention is applied may basically include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0257] The encoding server compresses content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream, and transmits the bitstream to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a camcorder directly generates a bitstream, the encoding server may be omitted.
[0258] A bitstream may be generated by applying the encoding method or the bitstream generation method of the embodiments of this document, and the streaming server may temporarily store the bitstream in a process of transmitting or receiving the bitstream.
[0259] The streaming server transmits multimedia data to user devices via a web server based on user requests. The web server also serves as a medium for informing users of services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control commands and responses between devices in the content streaming system.
[0260] The streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined time.
[0261] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, touch-screen PCs, tablet PCs, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc.
[0262] The individual servers in the content streaming system may operate as distributed servers, in which case data received from the individual servers may be distributed.
[0263] The claims described herein can be combined in various ways. For example, the technical features of the method claims of this specification can be combined and implemented as a device, and the technical features of the device claims of this specification can be combined and implemented as a method. Furthermore, the technical features of the method claims of this specification and the technical features of the device claims of this specification can be combined and implemented as a device, and the technical features of the method claims of this specification and the technical features of the device claims of this specification can be combined and implemented as a method.
Claims
1. A method for decoding an image, the method comprising: Export the position of the current sub-block within the current block; deriving a collocated sub-block within the collocated picture based on a position of the current sub-block; deriving a sub-block based temporal motion information candidate based on the motion vector of the collocated sub-block; deriving a motion vector of the current sub-block based on the sub-block-based temporal motion information candidate; generating a prediction sample of the current block based on the motion vector of the current sub-block; as well as Based on the predicted samples, reconstructed samples are generated, Wherein, based on the center sample position of the current sub-block, the position of the current sub-block is derived, wherein, based on the center sample position of the current sub-block, the collocated sub-block is derived within the collocated picture, The center sample position represents the position of the lower right center sample located at the lower right among the four samples located at the center. wherein the sub-block based temporal motion information candidate is derived based on the motion vector of the collocated sub-block derived based on the availability of the collocated sub-block, For the available collocated sub-blocks, the motion vectors of the available collocated sub-blocks are derived as the sub-block-based temporal motion information candidates. wherein the motion vector of the available collocated sub-block is derived based on the motion vector of the block covering the modified position derived based on the upper left sample position of the collocated sub-block, and Wherein, the modified position is derived by the equation ((xColSb>>3)<<3, (yColSb>>3)<<3), Here, xColSb and yColSb represent the x-coordinate and y-coordinate of the sample position of the collocated sub-block derived based on the center sample position of the current sub-block, respectively, and >> represents an arithmetic right shift, and << represents an arithmetic left shift.
2. A method for encoding an image, the method comprising: Export the position of the current sub-block within the current block; deriving a collocated sub-block within the collocated picture based on a position of the current sub-block; deriving a sub-block based temporal motion information candidate based on the motion vector of the collocated sub-block; deriving a motion vector of the current sub-block based on the sub-block-based temporal motion information candidate; generating a prediction sample of the current block based on the motion vector of the current sub-block; generating information about residual samples derived based on the prediction samples, and encoding image information including information about the residual samples, Wherein, based on the center sample position of the current sub-block, the position of the current sub-block is derived, wherein, based on the center sample position of the current sub-block, the collocated sub-block is derived within the collocated picture, The center sample position represents the position of the lower right center sample located at the lower right among the four samples located at the center. wherein the sub-block based temporal motion information candidate is derived based on the motion vector of the collocated sub-block derived based on the availability of the collocated sub-block, For the available collocated sub-blocks, the motion vectors of the available collocated sub-blocks are derived as the sub-block-based temporal motion information candidates. wherein the motion vector of the available collocated sub-block is derived based on the motion vector of the block covering the modified position derived based on the upper left sample position of the collocated sub-block, and Wherein, the modified position is derived by the equation ((xColSb>>3)<<3, (yColSb>>3)<<3), Here, xColSb and yColSb represent the x-coordinate and y-coordinate of the sample position of the collocated sub-block derived based on the center sample position of the current sub-block, respectively, and >> represents an arithmetic right shift, and << represents an arithmetic left shift.
3. A method for transmitting data for image information, the method comprising: obtaining a bitstream of image information including information about residual samples; as well as sending data of a bitstream including the image information, the image information including information about the residual samples, The bitstream is generated based on the following steps: deriving a position of a current subblock within a current block; deriving a collocated subblock within a collocated picture based on the position of the current subblock; deriving a subblock-based temporal motion information candidate based on a motion vector of the collocated subblock; deriving a motion vector of the current subblock based on the subblock-based temporal motion information candidate; generating prediction samples of the current block based on the motion vector of the current subblock; generating information about residual samples derived based on the prediction samples; and encoding image information including the information about the residual samples. Wherein, based on the center sample position of the current sub-block, the position of the current sub-block is derived, wherein, based on the center sample position of the current sub-block, the collocated sub-block is derived within the collocated picture, The center sample position represents the position of the lower right center sample located at the lower right among the four samples located at the center. wherein the sub-block based temporal motion information candidate is derived based on the motion vector of the collocated sub-block derived based on the availability of the collocated sub-block, For the available collocated sub-blocks, the motion vectors of the available collocated sub-blocks are derived as the sub-block-based temporal motion information candidates. wherein the motion vector of the available collocated sub-block is derived based on the motion vector of the block covering the modified position derived based on the upper left sample position of the collocated sub-block, and Wherein, the modified position is derived by the equation ((xColSb>>3)<<3, (yColSb>>3)<<3), Here, xColSb and yColSb represent the x-coordinate and y-coordinate of the sample position of the collocated sub-block derived based on the center sample position of the current sub-block, respectively, and >> represents an arithmetic right shift, and << represents an arithmetic left shift.
Citation Information
Patent Citations
Method for sub-PU motion information inheritance in 3D video coding
KR1020160117489A
KR20190053238A