Inter prediction based image or video coding using sbtmvp
The sbTMVP method is used to derive the temporal motion vectors of sub-blocks in high-resolution images and videos, solving the problem of high transmission and storage costs in existing technologies and achieving more efficient image and video compression and transmission.
Patent Information
- Application Number
- CN202510979132.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-13
- Filing Date
- 2020-06-15
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies struggle to effectively compress and transmit high-resolution, high-quality image and video data, especially in immersive media such as virtual reality, artificial reality, and holograms, leading to increased transmission and storage costs.
The method based on sub-block temporal motion vector prediction (sbTMVP) is adopted. By deriving a reference sub-block at the center position of the current sub-block, the motion vectors of available and unavailable reference sub-blocks are used to improve the efficiency of inter-frame prediction. The corresponding positions at the sub-compilation block level and the compilation block level are unified to calculate the temporal motion vector.
It improves image and video compression efficiency, reduces computational complexity, simplifies hardware implementation, and enhances prediction performance.
Smart Images

Figure CN120935366A_ABST
Abstract
Description
[0001] This application is a divisional application of patent application No. 202080050139.0 (PCT / KR2020 / 007713), filed on January 10, 2022, with a filing date of June 15, 2020, entitled "Image or video compilation based on inter-frame prediction using SBTMVP". Technical Field
[0002] This invention relates to video or image compilation, for example, an inter-frame prediction-based image or video compilation technique using sub-block-based temporal motion vector prediction (sbTMVP). Background Technology
[0003] There has been a growing demand for high-resolution, high-quality images and videos, such as ultra-high-definition (HUD) images and 4K or 8K or higher video, across various fields. As image and video data becomes higher resolution and higher quality, the relative amount of information or bits transmitted increases compared to existing image and video data. Therefore, transmission and storage costs increase if media such as existing wired or wireless broadband lines are used to transmit image data or if existing storage media are used to store image and video data.
[0004] Furthermore, there has been a growing interest in and demand for immersive media such as virtual reality (VR), artificial reality (AR) content, or holograms. The broadcasting of images and videos, such as game graphics, which differ from the image characteristics of real-world images, is also increasing.
[0005] Therefore, in order to effectively compress and transmit or store and play back information of high-resolution and high-quality images and videos with such various characteristics, efficient image and video compression technologies are needed.
[0006] Furthermore, to improve image / video compilation efficiency, a sub-block-based temporal motion vector prediction technique is discussed. Therefore, a scheme is needed to efficiently perform the process of patching motion vectors of sub-block units in sub-block-based temporal motion vector prediction. Summary of the Invention
[0007] Technical issues
[0008] The purpose of this document is to provide a method and apparatus for improving the efficiency of video / image compilation.
[0009] Another objective of this document is to provide a method and apparatus for effective inter-frame prediction.
[0010] Another objective of this document is to provide a method and apparatus for improving prediction performance by deriving sub-block-based temporal motion vectors.
[0011] Another objective of this document is to provide a method and apparatus for efficiently deriving the corresponding positions of sub-blocks to derive time motion vectors based on the sub-blocks.
[0012] Another objective of this document is to provide a method and apparatus for unifying corresponding positions at the sub-compilation block level and corresponding positions at the compilation block level to derive sub-block-based time motion vectors.
[0013] Technical solution
[0014] According to embodiments of this disclosure, a reference subblock for the current subblock can be derived based on the position of the lower right sample among four samples located at the center of the current subblock in the subblock temporal motion vector prediction (sbTMVP).
[0015] According to embodiments of this disclosure, sbTMVP candidates can be derived based on the availability of reference subblocks for the current subblock; for available reference subblocks, the motion vectors of the available reference subblocks are derived as sbTMVP candidates, and for unavailable reference subblocks, the basic motion vectors are derived as sbTMVP candidates.
[0016] According to embodiments of this disclosure, an unavailable reference sub-block may include a reference sub-block located outside a reference image or a reference sub-block whose motion vectors are unavailable, and for a reference sub-block that is an intra-frame mode, IBC (intra-block copy) mode or palette mode, the reference sub-block may be a sub-block whose motion vectors are unavailable.
[0017] According to embodiments of this disclosure, a video / image decoding method performed by a decoding device is provided. The video / image decoding method may include the methods disclosed in the embodiments of this disclosure.
[0018] According to embodiments of this disclosure, a decoding apparatus is provided for performing video / image decoding. The decoding apparatus can perform the methods disclosed in the embodiments of this disclosure.
[0019] According to embodiments of this disclosure, a video / image encoding method performed by an encoding apparatus is provided. The video / image encoding method may include the methods disclosed in the embodiments of this disclosure.
[0020] According to embodiments of this disclosure, an encoding apparatus is provided for performing video / image encoding. The encoding apparatus can perform the methods disclosed in the embodiments of this disclosure.
[0021] According to embodiments of the present disclosure, a computer-readable digital storage medium is provided that stores encoded video / image information generated according to a video / image encoding method disclosed in at least one embodiment of the present disclosure.
[0022] According to embodiments of the present disclosure, a computer-readable digital storage medium stores encoded information or encoded video / image information that enables a decoding device to perform the video / image decoding method disclosed in at least one embodiment of the present disclosure.
[0023] Beneficial effects
[0024] This document can have various effects. For example, it can improve overall image / video compression efficiency. Furthermore, through efficient inter-frame prediction, computational complexity can be reduced and overall compilation efficiency can be improved. Additionally, efficiency in terms of complexity and prediction performance can be improved because the corresponding positions of sub-blocks used to derive sub-block-based temporal motion vectors are efficiently computed in sub-block-based temporal motion vector prediction (sbTMVP). Moreover, because the methods for computing corresponding positions at the sub-compilation block level and at the compilation block level to derive sub-block-based temporal motion vectors are unified, hardware implementation simplification is achieved.
[0025] The effects achievable through the detailed embodiments in this document are not limited to those listed. For example, various technical effects may exist that can be understood or derived by those skilled in the art from this document. Therefore, the detailed effects in this document are not limited to those explicitly described herein, but may include various effects that can be understood or derived from the technical features of this document. Attached Figure Description
[0026] Figure 1 The illustrations represent examples of video / image compilation systems to which embodiments of this document can be applied.
[0027] Figure 2 This is a schematic diagram illustrating the configuration of a video / image encoding apparatus to which embodiments of this document may be applied.
[0028] Figure 3 This is a schematic diagram illustrating the configuration of a video / image decoding apparatus to which embodiments of this document can be applied.
[0029] Figure 4 The illustration shows an example of a video / image coding method based on inter-frame prediction, and Figure 5 The illustration shows an example of an inter-frame prediction unit in a coding apparatus.
[0030] Figure 6 An example of a video / image decoding method based on inter-frame prediction is shown, and Figure 7 The illustration shows an example of an inter-frame prediction unit in a decoding device.
[0031] Figure 8 An exemplary illustration shows the spatial neighbor blocks and temporal neighbor blocks of the current block.
[0032] Figure 9 An exemplary diagram is shown that is used to derive temporal neighbor blocks based on sub-block temporal motion information candidates (sbTMVP candidates).
[0033] Figure 10 This is a schematic diagram illustrating the process of deriving sub-block-based temporal motion information candidates (sbTMVP candidates).
[0034] Figure 11 This is a schematic diagram illustrating the method used in the sbTMVP export process to calculate the corresponding positions of the default MV and sub-block MV based on the block size.
[0035] Figure 12 This is an exemplary view schematically illustrating a method for integrating the corresponding positions of the default MV and sub-block MV based on the block size during the sbTMVP export process.
[0036] Figure 13 and Figure 14 This is an exemplary view schematically illustrating the pipeline configuration used in the sbTMVP export process to calculate the corresponding positions of the default MV and sub-block MV based on the block size.
[0037] Figure 15 and Figure 16 Examples of video / image encoding methods and related components according to one or more embodiments of the present disclosure are illustrated schematically.
[0038] Figure 17 and Figure 18 Examples of video / image decoding methods and related components according to one or more embodiments of the present disclosure are illustrated schematically.
[0039] Figure 19 An example of a content streaming system to which the embodiments disclosed in this document can be applied is illustrated. Detailed Implementation
[0040] This document can be modified in various ways and can have various implementations, and specific implementations are illustrated and described in detail in the accompanying drawings. However, this is not intended to limit this document to the specific implementation. The terminology generally used in this specification is used to describe specific implementations and not to limit the technical spirit of this document. Unless otherwise expressly indicated in the context, singular expressions include plural expressions. Terms such as "comprising" or "having" in this specification should be understood to indicate the presence of the features, numbers, steps, operations, elements, components, or combinations thereof described in this specification, without excluding the possibility of the presence or addition of one or more features, numbers, steps, operations, elements, components, or combinations thereof.
[0041] Furthermore, for ease of description in relation to different features and functions, the elements in the accompanying drawings described in this document are illustrated independently. This does not imply that each element is implemented as a separate piece of hardware or separate piece of software. For example, at least two elements may be combined to form a single element, or a single element may be divided into multiple elements. Embodiments in which elements are combined and / or separated are also included within the scope of this document, unless they depart from its spirit.
[0042] This document relates to video / image coding. For example, the methods / exercises disclosed in this document can be applied to methods disclosed in Universal Video Coding (VVC). Furthermore, the methods / exercises disclosed in this document can be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second-generation Audio Video Coding Standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268, etc.).
[0043] This document presents various embodiments of video / image compilation, and unless otherwise stated, these embodiments may be performed in combination with each other.
[0044] In this document, video can refer to a collection of images over time. An image typically refers to a unit representing an image within a specific time period, and a slice / tile is a unit that constitutes part of an image during compilation. A slice / tile can include one or more Compilation Tree Units (CTUs). An image can consist of one or more slices / tiles. A tile is a rectangular area of CTUs within a specific tile column and row in an image. A tile column is a rectangular area where the height of a CTU is equal to the height of the image and the width is specified by a syntax element in the image parameter set. A tile row is a rectangular area where the height of a CTU is specified by a syntax element in the image parameter set and the width is equal to the width of the image. A tile scan is a specific ordering of the CTUs that divide an image: CTUs can be ordered consecutively by CTU raster scan within a tile, and tiles in an image can be ordered consecutively by the raster scan of the image's tiles. A slice includes an integer number of complete tiles or an integer number of consecutive complete CTU rows that can be exclusively contained within a single NAL unit of an image.
[0045] Furthermore, an image can be divided into two or more sub-images. A sub-image can be a rectangular region of one or more slices within the image.
[0046] A pixel or cell (pel) can refer to the smallest unit that makes up a picture (or image). Alternatively, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Alternatively, a sample can refer to a pixel value in the spatial domain, or, when a pixel value is transformed to the frequency domain, to the transform coefficients in the frequency domain.
[0047] A unit can represent a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the context, units and terms such as blocks and regions may be used interchangeably. Typically, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0048] Furthermore, in this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantization transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or, for the sake of consistency, may still be referred to as transform coefficients.
[0049] In this document, quantization transform coefficients and transform coefficients can be referred to as transform coefficients and scaling transform coefficients, respectively. In this case, residual information can include information about the transform coefficients, and this information can be sent as a signal using residual compilation syntax. Transform coefficients can be derived based on the residual information (or information about the transform coefficients), and scaling transform coefficients can be derived by performing an inverse transform (scaling) on the transform coefficients. Residual samples can be derived based on the inverse transform of the scaled transform coefficients. This can also be applied / expressed in other parts of this document.
[0050] In this document, the term "A or B" may mean "A only", "B only", or "both A and B". In other words, in this document, the term "A or B" may be interpreted as indicating "A and / or B". For example, in this document, the term "A, B, or C" may mean "A only", "B only", "C only", or "any combination of A, B, and C".
[0051] The forward slash " / " or comma used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0052] In this document, "at least one of A and B" can mean "only A", "only B" or "both A and B". Furthermore, in this document, expressions such as "at least one of A or B" or "at least one of A and / or B" can be interpreted as "at least one of A and B".
[0053] Furthermore, in this document, "at least one of A, B, and C" may mean "A only", "B only", "C only" or "any combination of A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".
[0054] Furthermore, the parentheses used in this document may mean "for example". Specifically, when expressing "prediction (intra-frame prediction)", it may be indicated that "intra-frame prediction" is presented as an example of "prediction". In other words, the term "prediction" in this document is not limited to "intra-frame prediction", and it may be indicated that "intra-frame prediction" is presented as an example of "prediction". Moreover, even when expressing "prediction (i.e., intra-frame prediction)", it may be indicated that "intra-frame prediction" is presented as an example of "prediction".
[0055] In this document, a technical feature explained separately in a single figure may be implemented individually or simultaneously.
[0056] In the following, preferred embodiments of this document will be described in more detail with reference to the accompanying drawings. In the following drawings, the same reference numerals are used for the same elements, and repeated descriptions of the same elements may be omitted.
[0057] Figure 1 Examples of video / image compilation systems to which the implementation methods described in this document can be applied are illustrated.
[0058] refer to Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device via digital storage media or a network in the form of files or streams.
[0059] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0060] Video sources can be obtained through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, video / image archives including previously captured video / images, etc. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images can be generated by computers, etc. In this case, the video / image capture process can be replaced by processes that generate related data.
[0061] An encoding device can encode input video / images. The encoding device can perform a series of processes such as prediction, transformation, and quantization for compression and compilation efficiency. The encoded data (encoded video / image information) can be output as a bitstream.
[0062] A transmitter can send encoded video / image information or data, output as a bitstream, to a receiver of a receiving device via a digital storage medium or network, either as a file or a stream. Digital storage media can include various media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiver can receive / extract the bitstream and send the received / extracted bitstream to a decoding device.
[0063] Decoding devices can decode video / images by performing a series of processes such as dequantization, inverse transform, and prediction, which correspond to the operations of encoding devices.
[0064] The renderer can render decoded video / images. The rendered video / images can then be displayed on a monitor.
[0065] Figure 2 This is a schematic diagram illustrating the configuration of a video / image encoding apparatus to which this document can be applied. In the following text, the encoding apparatus may include an image encoding apparatus and / or a video encoding apparatus.
[0066] refer to Figure 2 The encoding apparatus 200 may include an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 described above may be constituted by one or more hardware components (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded image buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.
[0067] Image segmenter 210 can segment an input image (or picture or frame) input to encoding device 200 into one or more processing units. As an example, a processing unit may be referred to as a compilation unit (CU). In this case, starting from a compilation tree unit (CTU) or a maximum compilation unit (LCU), the compilation unit can be recursively segmented according to a quadtree-binary-tritree (QTBTTT) structure. For example, a compilation unit can be divided into multiple compilation units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, a quadtree structure can be applied first, followed by a binary tree structure and / or a ternary tree structure. Alternatively, a binary tree structure can be applied first. The compilation process according to this document can be performed based on the final compilation unit that has not been further segmented. In this case, based on the compilation efficiency according to image characteristics, the maximum compilation unit can be directly used as the final compilation unit. Alternatively, the compilation unit can be recursively segmented into compilation units of even greater depth as needed, such that the optimally sized compilation unit can be used as the final compilation unit. Here, the compilation process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may also include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can be divided or segmented from the final compilation unit described above. The prediction unit may be a unit for predicting samples, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving the residual signal from the transformation coefficients.
[0068] Depending on the context, units and terms such as blocks and regions can be used interchangeably. Typically, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. Samples can usually represent pixels or pixel values, and can represent pixel / pixel values only for the luminance component, or only for the chrominance component. Samples can be used as a term corresponding to pixels or cells (pel) of a picture (or image).
[0069] In the encoding apparatus 200, a residual signal (residual block, residual sample array) is generated by subtracting the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 from the input image signal (original block, original sample array), and the generated residual signal is sent to the converter 232. In this case, as shown, the unit in the encoder 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be called the subtractor 231. The predictor can perform prediction on the processing target block (hereinafter referred to as the "current block") and can generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As discussed later in the description of each prediction mode, the predictor can generate various prediction-related information such as prediction mode information and send the generated information to the entropy encoder 240. The prediction information can be encoded in the entropy encoder 240 and output as a bitstream.
[0070] Intra-predictor 222 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the reference samples can be located near or separate from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Depending on the level of detail in the prediction direction, the directional modes can include, for example, 33 or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes can be used depending on the settings. Intra-predictor 222 can determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.
[0071] Inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted based on the correlation between motion information of neighboring blocks and the current block, on a block, sub-block, or sample basis. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same as or different from each other. The temporally neighboring block may be referred to as a juxtaposed reference block, juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use motion information of neighboring blocks as motion information of the current block. In skip mode, unlike merge mode, residual signals cannot be sent. In motion information prediction (motion vector prediction, MVP) mode, motion vectors of neighboring blocks can be used as motion vector prediction terms, and the motion vector of the current block can be indicated by sending the motion vector difference as a signal.
[0072] Predictor 220 can generate prediction signals based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction to the prediction of a block, and can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Alternatively, the predictor can perform prediction on blocks based on an intra-frame block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video compilation, such as in games like Screen Content Compilation (SCC). Although IBC essentially performs prediction in the current image, its execution is similar to inter-frame prediction in that it derives a reference block in the current image. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. The palette mode can be considered an example of intra-frame compilation or intra-frame prediction. When applying a palette mode, sample values in the image can be signaled based on information about the palette index and palette table.
[0073] The predicted signal generated by the predictor (including inter-frame predictor 221 and / or intra-frame predictor 222) can be used to generate a reconstructed signal or a residual signal. Transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform technique can include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transform obtained based on the predicted signal generated using all previously reconstructed pixels. Furthermore, the transform process can be applied to square pixel blocks of the same size, or to blocks of variable size that are not square.
[0074] Quantizer 233 quantizes the transform coefficients and sends them to entropy encoder 240, which encodes the quantized signal (information about the quantized transform coefficients) and outputs the encoded signal in a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order and generate information about the quantized transform coefficients based on this one-dimensional vector form. Entropy encoder 240 can perform various encoding methods such as exponential Golomb, context-adaptive variable-length compilation (CAVLC), and context-adaptive binary arithmetic compilation (CABAC). Entropy encoder 240 can encode information required for video / image reconstruction, other than the quantized transform coefficients (e.g., values of syntax elements), either together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form on a unit-by-unit basis in the Network Abstraction Layer (NAL). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS). Additionally, the video / image information may include general constraint information. In this document, information and / or syntax elements sent from the encoding device to / signaled to the decoding device may be included in the video / image information. The video / image information can be encoded using the encoding process described above and included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks, communication networks, and / or the like, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that sends the signal output from the entropy encoder 240 or a memory (not shown) that stores it may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoder 240.
[0075] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transform to the quantized transform coefficients via dequantizer 234 and inverse transformer 235, the residual signal (residual block or residual sample) can be reconstructed. Adder 155 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222, thereby generating a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). When there is no residual for the processing target block, as in the case of applying skip mode, the prediction block can be used as the reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current image, and, as described later, can be used for inter-frame prediction of the next image by filtering.
[0076] In addition, Luminance Mapping and Chromaticity Scaling (LMCS) can be applied in image encoding and / or reconstruction processing.
[0077] Filter 260 can improve subjective / objective video quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be stored in memory 270, specifically in the DPB of memory 270. Various filtering methods can include, for example, deblocking filtering, sample adaptive offsetting, adaptive loop filtering, bilateral filtering, etc. As discussed later in the description of each filtering method, filter 260 can generate various filtering-related information and send the generated information to entropy encoder 240. The filtering information can be encoded in entropy encoder 240 and output as a bitstream.
[0078] The modified reconstructed image sent to memory 270 can be used as a reference image in inter-frame predictor 221. Accordingly, the encoding device can avoid prediction mismatch in the encoding device 100 and the decoding device when applying inter-frame prediction, and can also improve encoding efficiency.
[0079] The memory 270DPB can store modified reconstructed images for use as reference images in the inter-frame predictor 221. The memory 270 can store motion information of blocks in the current image from which motion information has been derived (or encoded) and / or motion information of blocks in reconstructed images. The stored motion information can be sent to the inter-frame predictor 221 to be used as motion information for neighboring blocks or temporally neighboring blocks. The memory 270 can store reconstructed samples of reconstructed blocks in the current image and send them to the intra-frame predictor 222.
[0080] Figure 3This is a schematic diagram illustrating the configuration of a video / image decoding apparatus to which this document can be applied. In the following text, the decoding apparatus may include an image decoding apparatus and / or a video decoding apparatus.
[0081] refer to Figure 3 The video decoding apparatus 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to embodiments, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above may be constituted by one or more hardware components (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0082] When the input includes a bitstream containing video / image information, the decoding device 300 can interact with the data already prepared therein. Figure 2 The processing of video / image information in the encoding apparatus correspondingly reconstructs the image. For example, the decoding apparatus 300 can derive units / blocks based on information related to block segmentation obtained from the bitstream. The decoding apparatus 300 can perform decoding by using processing units applied in the encoding apparatus. Therefore, the decoding processing unit can be, for example, a compilation unit, which can be segmented from a compilation tree unit or a maximum compilation unit along a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transformation units can be derived from the compilation unit. The reconstructed image signal decoded and output by the decoding apparatus 300 can be reproduced by a reproduction apparatus.
[0083] Decoding device 300 is capable of receiving data from a source in the form of a bitstream. Figure 2The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc. In addition, the video / image information may also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or general constraint information. In this document, the information and / or syntax elements sent / received by the signal, as described subsequently, can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on a compilation method such as Exponential Columbus compilation, CAVLC, or CABAC, and can output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information about the target syntax element and the decoding information of neighboring and target blocks, or information about symbols / bins decoded in previous steps, predict the bin generation probability based on the determined context model, and perform arithmetic decoding on the bins to generate symbols corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model after determining it using information about symbols / bins decoded for the context model of the next symbol / bin. Prediction information from the information decoded in the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantization transform coefficients) and associated parameter information that have undergone entropy decoding in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). Additionally, filtering information from the information decoded in the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not shown) that receives the signal output from the encoding device can also configure the decoding device 300 as an internal / external component, and the receiver can be a component of the entropy decoder 310. Additionally, the decoding device according to this document can be referred to as a video / image / picture decoding device, and the decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0084] The dequantizer 321 can output transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into two-dimensional blocks. In this case, the rearrangement can be performed based on the order of coefficient scans already performed in the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0085] The inverse converter 322 obtains the residual signal (residual block, residual sample array) by performing an inverse transformation on the transformation coefficients.
[0086] The predictor can perform predictions on the current block and generate a prediction block that includes prediction samples for the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the predictions output from the entropy decoder 310, and can determine a specific intra-frame / inter-frame prediction mode.
[0087] Predictor 320 can generate prediction signals based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction not only to the prediction of a block, but also to both intra-frame and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Additionally, the predictor can perform block prediction based on an intra-frame block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video compilation, such as in games or games with screen content compilation (SCC). Although IBC essentially performs prediction within the current frame, its execution is similar to inter-frame prediction in that it derives a reference block from the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. The palette mode can be considered an example of intra-frame compilation or intra-frame prediction. When a palette mode is applied, information about the palette table and palette index can be included in the video / image information and transmitted as a signal.
[0088] The intra-predictor 331 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the reference samples can be located near or separate from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-predictor 331 can determine the prediction mode applied to the current block by using the prediction modes applied to neighboring blocks.
[0089] Inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted based on the correlation between motion information of neighboring blocks and the current block, on a block, sub-block, or sample basis. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference image index of the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information about the prediction may include information indicating the mode of inter-frame prediction used for the current block.
[0090] Adder 340 adds the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (inter-frame predictor 332 or intra-frame predictor 331) to generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array). When there is no residual for processing the target block, as in the case of applying skip mode, the prediction block can be used as the reconstruction block.
[0091] Adder 340 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and can be output by filtering as described below, or it can be used for inter-frame prediction of the next image.
[0092] In addition, Luminance Mapping and Chromaticity Scaling (LMCS) can be applied to image decoding processing.
[0093] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 360, specifically in the DPB of memory 360. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.
[0094] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame predictor 332. Memory 360 can store motion information of blocks from which motion information in the current image is derived (or decoded) and / or motion information of blocks in reconstructed images. The stored motion information can be sent to inter-frame predictor 332 to be used as motion information for spatially or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and transmit the reconstructed samples to intra-frame predictor 331.
[0095] In this disclosure, the embodiments described in the filter 260, inter-frame predictor 221, and intra-frame predictor 222 of the encoding apparatus 200 can be the same as or applied to the filter 350, inter-frame predictor 332, and intra-frame predictor 331 of the decoding apparatus 300, respectively. This can also be applied to unit 332 and intra-frame predictor 331.
[0096] As described above, during video compilation, prediction is performed to improve compression efficiency. A prediction block, i.e., a target compiled block, can be generated by prediction, including prediction samples for the current block. In this case, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is similarly derived in both the encoding and decoding devices. The encoding device can improve image compilation efficiency by signaling information (residual information) about the residual between the original block (not the original block) and the prediction block and the original sample values. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and generate a reconstructed image including the reconstructed block.
[0097] Residual information can be generated through transformation and quantization processes. For example, the encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transform process on the residual samples (residual sample array) included in the residual block, derive quantized transform coefficients by performing a quantization process on the transform coefficients, and send the relevant residual information to the decoding device as a signal (via bitstream). In this case, the residual information can include information such as value information, position information, transform scheme, transform kernel, and quantization parameters of the quantized transform coefficients. The decoding device can perform dequantization / inverse transform processes based on the residual information and derive residual samples (or residual blocks). The decoding device can generate a reconstructed image based on the prediction block and the residual block. Furthermore, the encoding device can derive residual blocks by performing dequantization / inverse transform on the quantized transform coefficients for inter-frame prediction reference of subsequent images and can generate a reconstructed image.
[0098] Furthermore, as mentioned above, when performing prediction on the current block, either intra-frame prediction or inter-frame prediction can be applied. The following section will describe the case where inter-frame prediction is applied to the current block.
[0099] The predictor (more specifically, the inter-frame predictor) of the encoding / decoding device can derive prediction samples by performing inter-frame prediction on a block-by-block basis. Inter-frame prediction can represent a prediction derived by relying on data elements (e.g., sample values or motion information) of images other than the current image. When applying inter-frame prediction to the current block, a prediction block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) on a reference image indicated by a reference image index, specified by a motion vector. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample-by-sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information can include motion vectors and reference image indices. Motion information can further include inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When applying inter-frame prediction, neighboring blocks can include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block can be the same as or different from each other. Temporally neighboring blocks can be referred to by names such as juxtaposed reference blocks, juxtaposed CUs (colCU), etc., and the reference picture including the temporally neighboring blocks can be referred to as the juxtaposed picture (colPic). For example, a motion information candidate list can be configured based on the neighboring blocks of the current block, and flags or index information indicating which candidate to select (use) can be signaled to derive the motion vector and / or reference picture index of the current block. Inter-frame prediction can be performed based on various prediction modes, and for example, in skip mode and merge mode, the motion information of the current block can be the same as the motion information of the selected neighboring blocks. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, the motion vectors of the selected neighboring blocks can be used as motion vector predictors, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived by using the sum of the motion vector predictor and the motion vector difference.
[0100] Motion information may further include L0 motion information and / or L1 motion information based on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). The L0 direction motion vector may be referred to as the L0 motion vector or MVL0, and the L1 direction motion vector may be referred to as the L1 motion vector or MVL1. Prediction based on the L0 motion vector may be referred to as L0 prediction, prediction based on the L1 motion vector may be referred to as L1 prediction, and prediction based on both L0 and L1 motion vectors may be referred to as bi-prediction. Here, the L0 motion vector may indicate the motion vector associated with the reference image list L0, and the L1 motion vector may indicate the motion vector associated with the reference image list L1. The reference image list L0 may include images preceding the current image in output order, and the reference image list L1 may include images following the current image in output order as reference images. The previous image may be referred to as the forward (reference) image, and the subsequent image may be referred to as the backward (reference) image. The reference image list L0 may further include images following the current image in output order as reference images. In this scenario, previous images can be indexed first in the reference image list L0, and then subsequent images can be indexed. The reference image list L1 can further include images preceding the current image in the output order as reference images. In this case, subsequent images can be indexed first in the reference image list L1, and then previous images can be indexed. Here, the output order can correspond to the image order count (POC) order.
[0101] Furthermore, various inter-frame prediction modes can be used to predict the current block in an image. For example, modes such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, merge with MVD (MMVD) mode, and history motion vector prediction (HMVP) mode can be used. Decoder-side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, dual prediction with CU-level weights (BCW), bidirectional optical flow (BDOF), etc., can be further used as additional modes. Affine mode can also be referred to as affine motion prediction mode. MVP mode can also be referred to as advanced motion vector prediction (AMVP) mode. In this document, some modes and / or motion information candidates derived from some modes can also be included in one of the motion information-related candidates in other modes. For example, HMVP candidates can be added to the merge candidates of merge / skip mode, or they can be added to the MVP candidates of MVP mode. If an HMVP candidate is used as a motion information candidate for merge mode or skip mode, the HMVP candidate can be referred to as an HMVP merge candidate.
[0102] Prediction mode information, indicating the inter-frame prediction mode of the current block, can be signaled from the encoding device to the decoding device. In this case, the prediction mode information can be included in the bitstream and received by the decoding device. The prediction mode information may include index information indicating one of several candidate modes. Alternatively, the inter-frame prediction mode can be indicated by hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, whether a skip mode is applied can be indicated by signaling a skip flag, whether a merge mode is applied can be indicated by signaling a merge flag when a skip mode is not applied, and flags for additional differentiation can be signaled to indicate the application of the MVP mode or, when the merge mode is not applied. Affine modes can be signaled as independent modes or as subordinate modes to the merge or MVP modes. For example, affine modes may include an affine merge mode and an affine MVP mode.
[0103] Furthermore, when applying inter-frame prediction to the current block, motion information of the current block can be used. The encoding device can derive the optimal motion information for the current block through a motion estimation process. For example, the encoding device can search for highly correlated similar reference blocks in a predetermined search range in a reference image using original blocks from the original image used for the current block, in fractional pixel units, and derive motion information from the searched reference blocks. Block similarity can be derived based on the difference of phase-based sample values. For example, block similarity can be calculated based on the sum of absolute differences (SAD) between the current block (or its template) and a reference block (or its template). In this case, motion information can be derived based on the reference block with the smallest SAD in the search region. The derived motion information can be sent as a signal to the decoding device based on the inter-frame prediction mode, according to various methods.
[0104] A prediction block for the current block can be derived based on motion information derived from the inter-frame prediction mode. The prediction block can include prediction samples (an array of prediction samples) for the current block. An interpolation process can be performed when the motion vector (MV) of the current block indicates fractional sample units, and the prediction samples for the current block can be derived from reference samples of fractional sample units in the reference image through the interpolation process. When affine inter-frame prediction is applied to the current block, prediction samples can be generated based on sample / sub-block units MV. When applying dual prediction, the prediction samples derived from a weighted sum or weighted average of prediction samples derived from L0 prediction (i.e., predictions using reference images in the reference image list L0 and MVL0) and prediction samples derived (based on phase) from L1 prediction (i.e., predictions using reference images in the reference image list L1 and MVL1) can be used as the prediction samples for the current block. When applying dual prediction, if the reference images used for L0 prediction and L1 prediction are located in different time directions based on the current image (i.e., if the predictions correspond to dual prediction and bidirectional prediction), this can be called true dual prediction.
[0105] Reconstructed samples and reconstructed images can be generated based on the exported predicted samples, and then processes such as in-loop filtering can be performed as described above.
[0106] Figure 4 The illustration shows an example of a video / image coding method based on inter-frame prediction, and Figure 5 The illustration shows an example of an inter-frame prediction unit in a coding apparatus. Figure 5 The inter-frame prediction unit in the coding device can also be used as a... Figure 2 The inter-frame prediction unit 221 of the coding device 200 is the same as or corresponds to that of the encoding device 200.
[0107] refer to Figure 4 and Figure 5 The encoding device performs inter-frame prediction on the current block (S400). The encoding device can derive the inter-frame prediction mode and motion information of the current block, and generate prediction samples for the current block. Here, the inter-frame prediction mode determination process, the motion information derivation process, and the prediction sample generation process can be executed simultaneously, and any one of these processes can be executed earlier than the other processes.
[0108] For example, the inter-frame prediction unit 221 of the coding apparatus may include a prediction mode determination unit 221-1, a motion information derivation unit 221-2, and a prediction sample derivation unit 221-3. The prediction mode determination unit 221-1 can determine the prediction mode for the current block, the motion information derivation unit 221-2 can derive the motion information of the current block, and the prediction sample derivation unit 221-3 can derive the prediction samples of the current block. For example, the inter-frame prediction unit 221 of the coding apparatus can search for blocks similar to the current block in a predetermined region (search region) of a reference image through motion estimation, and derive a reference block whose difference from the current block is the smallest or equal to or less than a predetermined standard. Based on this, a reference image index indicating the location of the reference block can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The coding apparatus can determine the mode applicable to the current block from various prediction modes. The coding apparatus can compare the RD costs of various prediction modes and determine the optimal prediction mode for the current block.
[0109] For example, when a skip mode or merge mode is applied to the current block, the encoding device can configure a merge candidate list (described below) and export a reference block from among the reference blocks indicated by the merge candidates included in the merge candidate list whose difference from the current block is the smallest, equal to, or less than a predetermined standard. In this case, a merge candidate associated with the exported reference block can be selected, and merge index information indicating the selected merge candidate can be generated and sent to the decoding device as a signal. Motion information of the current block can be exported using the motion information of the selected merge candidate.
[0110] As another example, when the (A)MVP mode is applied to the current block, the encoding device can configure an (A)MVP candidate list and use the motion vector of the selected MVP candidate from the motion vector predictor (MVP) candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, the motion vector of a reference block derived through motion estimation can be used as the motion vector of the current block, and the MVP candidate with the motion vector having the smallest difference from the motion vector of the current block can become the selected MVP candidate. The motion vector difference (MVD) can be derived, which is the difference obtained by subtracting the MVP from the motion vector of the current block. In this case, information about the MVD can be signaled to the decoding device. Furthermore, when the (A)MVP mode is applied, the value of the reference picture index can be configured as reference picture index information and sent separately to the decoding device.
[0111] The encoding device can derive residual samples based on the predicted samples (S410). The encoding device can derive residual samples by comparing the initial samples with the predicted samples of the current block.
[0112] The encoding device encodes image information including prediction information and residual information (S420). The encoding device can output the encoded image information in the form of a bitstream. The prediction information may include information about prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and information about motion information, as information related to the prediction process. The information about motion information may include candidate selection information (e.g., merge index, MVP flag, or MVP index), which is information used to derive motion vectors. In addition, the information about motion information may include information about MVD and / or reference image index information. Furthermore, the information about motion information may include information indicating whether L0 prediction, L1 prediction, or dual prediction is applied. The residual information is information about residual samples. The residual information may include information about the quantization transform coefficients used for the residual samples.
[0113] The output bitstream can be stored in (digital) storage media and transmitted to the decoding device or transmitted to the decoding device via a network.
[0114] Simultaneously, as mentioned above, the encoding device can generate a reconstructed image (including reconstructed samples and reconstructed blocks) based on the reference samples and residual samples. This is to derive the same prediction result as that performed by the decoding device, and therefore, compilation efficiency can be improved. Thus, the encoding device can store the reconstructed image (or reconstructed samples or reconstructed blocks) in memory and use the reconstructed image as a reference image. As mentioned above, the in-loop filtering process can be further applied to the reconstructed image.
[0115] Figure 6 The illustration shows an example of a video / image decoding method based on inter-frame prediction, and Figure 7 The illustration shows an example of an inter-frame prediction unit in a decoding apparatus. Figure 7 The inter-frame prediction unit in the decoding device can also be used as a... Figure 3 The inter-frame prediction unit 332 of the decoding device 300 is the same as or corresponds to that of the decoding device 300.
[0116] refer to Figure 6 and Figure 7 The decoding device can perform operations corresponding to those performed by the encoding device. Based on the received prediction information, the decoding device can perform predictions on the current block and derive prediction samples.
[0117] Specifically, the decoding device can determine the prediction mode of the current block based on the received prediction information (S600). The decoding device can determine which inter-frame prediction mode to apply to the current block based on the prediction mode information in the prediction information.
[0118] For example, a merge flag can be used to determine whether a merge mode or (A)MVP mode is applied to the current block. Alternatively, a mode index can be used to select one of a variety of inter-frame prediction mode candidates. Inter-frame prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include the various inter-frame prediction modes described above.
[0119] The decoding device derives motion information for the current block based on the determined inter-frame prediction mode (S610). For example, when a skip mode or merge mode is applied to the current block, the decoding device can configure a merge candidate list and select a merge candidate from among the merge candidates included in the merge candidate list. Here, the selection can be performed based on selection information (merge index). The motion information for the current block can be derived using the motion information of the selected merge candidate. The motion information of the selected merge candidate can be used as the motion information for the current block.
[0120] As another example, when the (A)MVP mode is applied to the current block, the decoding device can configure an (A)MVP candidate list and use the motion vector of the selected MVP candidate from the motion vector predictor (MVP) candidates included in the (A)MVP candidate list as the MVP of the current block. Here, selection can be performed based on selection information (MVP flag or MVP index). In this case, the MVD of the current block can be derived based on information about the MVD, and the motion vector of the current block can be derived based on the MVP and MVD of the current block. Furthermore, the reference image index of the current block can be derived based on reference image index information. The image indicated by the reference image index in the reference image list of the current block can be exported as the reference image referenced by the inter-frame prediction of the current block.
[0121] Simultaneously, motion information for the current block can be exported without a candidate list configuration, and in this case, the motion information for the current block can be exported according to the process disclosed in the prediction mode. In this scenario, the candidate list configuration can be omitted.
[0122] The decoding device can generate prediction samples for the current block based on the motion information of the current block (S620). In this case, a reference image can be derived based on the reference image index of the current block, and prediction samples for the current block can be derived by using samples of the reference block indicated by the motion vector of the current block on the reference image. In some cases, a prediction sample filtering process for all or some of the prediction samples for the current block can be further performed.
[0123] For example, the inter-frame prediction unit 332 of the decoding device may include a prediction mode determination unit 332-1, a motion information derivation unit 332-2, and a prediction sample derivation unit 332-3. The prediction mode determination unit 332-1 can determine the prediction mode for the current block based on the received prediction mode information, the motion information derivation unit 332-2 can derive the motion information (motion vector and / or reference image index) of the current block based on information about the received motion information, and the prediction sample derivation unit 332-3 can derive the prediction sample of the current block.
[0124] The decoding device generates residual samples for the current block based on the received residual information (S630). The decoding device can generate reconstructed samples for the current block based on the predicted samples and residual samples, and generate a reconstructed image based on the generated reconstructed samples (S640). Thereafter, as described above, the in-loop filtering process can be further applied to the reconstructed image.
[0125] As described above, the inter-frame prediction process may include an inter-frame prediction mode determination step, a motion information derivation step depending on the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information. The inter-frame prediction process can be performed by the encoding and decoding devices described above.
[0126] Figure 8 An exemplary illustration shows the spatial neighboring block and the temporal neighboring block of the current block.
[0127] refer to Figure 8 A spatial neighbor block refers to a neighboring block located around the current block 800 (which is the target of the current inter-frame prediction), and may include neighboring blocks located to the left of the current block 800 or to the top of the current block 800. For example, a spatial neighbor block may include the lower-left neighbor block, the left neighbor block, the upper-right neighbor block, the upper neighbor block, and the upper-left neighbor block of the current block 800. Figure 8 The spatial neighboring block is illustrated as "S".
[0128] According to an exemplary embodiment, the encoding device / decoding device can detect available neighboring blocks by searching spatial neighboring blocks of the current block in a predetermined order (e.g., the lower left neighboring block, the left neighboring block, the upper right neighboring block, the upper neighboring block, and the upper left neighboring block), and export the motion information of the detected neighboring blocks as spatial motion information candidates.
[0129] A temporally neighboring block is a block located on a different image than the current image (i.e., a reference image) that includes the current block 800, and specifically refers to a juxtaposed block of the current block 800 within the reference image. Here, the reference image can be located before or after the current image in the Picture Order Count (POC). Furthermore, the reference image used to derive the temporally neighboring block can be called a juxtaposed reference image or a col image (juxtaposed image). Additionally, a juxtaposed block can refer to a block located in the col image at a position corresponding to the position of the current block 800, and is referred to as a col block. For example, as... Figure 8 As shown, a temporally neighboring block may include a col block located in the reference image (i.e., the col image) corresponding to the position of the lower right corner sample of the current block 800 (i.e., the col block including the lower right corner sample) and / or a col block located in the reference image (i.e., the col image) corresponding to the position of the center lower right sample of the current block 800 (i.e., the col block including the center lower right sample). Figure 8 The time-adjacent block is illustrated as "T".
[0130] According to an exemplary embodiment, the encoding / decoding device can detect available blocks by searching temporally neighboring blocks (e.g., the col block including the lower right sample and the col block including the center lower right sample) in a predetermined order, and derive the motion information of the detected blocks as temporal motion information candidates. As described above, the technique using temporally neighboring blocks can be referred to as Temporal Motion Vector Prediction (TMVP). Furthermore, the temporal motion information candidates can be referred to as TMVP candidates.
[0131] Simultaneously, prediction can also be performed by deriving motion information on a sub-block basis according to the inter-frame prediction mode. For example, in affine mode or TMVP mode, motion information can be derived on a sub-block basis. Specifically, the method for deriving temporal motion information candidates on a sub-block basis can be called sub-block-based temporal motion vector prediction (sbTMVP).
[0132] sbTMVP is a method that uses the motion field within a col image to improve the motion vector prediction (MVP) and merging patterns of compilation units within the current image. The col image for sbTMVP can be the same as that used by TMVP. However, in TMVP, motion prediction is performed at the compilation unit (CU) level. Conversely, in sbTMVP, motion prediction can be performed at the sub-block level or the sub-compilation unit (sub-CU) level. Furthermore, in TMVP, temporal motion information is derived from col blocks within the col image (in this case, the col block corresponds to the lower right corner sample position of the current block or the center lower right sample position of the current block). In sbTMVP, temporal motion information is derived after applying motion shift from the col image. In this case, motion shift can include the process of obtaining a motion vector from one of the spatially neighboring blocks of the current block and shifting by that motion vector.
[0133] Figure 9 An exemplary illustration shows a spatial neighbor block that can be used to derive sub-block-based temporal motion information candidates (sbTMVP candidates).
[0134] refer to Figure 9 Spatial neighbor blocks can include at least one of the following: the bottom-left neighbor block A0, the left neighbor block A1, the top-right neighbor block B0, and the top neighbor block B1 of the current block. In some cases, spatial neighbor blocks can further include, in addition to, the following: Figure 9 Another neighboring block besides the one shown, or may not include. Figure 9 The specific neighboring blocks shown are among the neighboring blocks. Furthermore, a spatial neighboring block may include only a specific neighboring block, and for example, it may include only the left neighboring block A1 of the current block.
[0135] For example, the encoding / decoding device can, while searching for spatially neighboring blocks in a predetermined search order, first detect the motion vectors of available spatially neighboring blocks, and determine the block in the reference image at the position indicated by the motion vector of the spatially neighboring block as the col block (i.e., the juxtaposed reference block). In this case, the motion vector of the spatially neighboring block can be represented as a temporal motion vector (time MV).
[0136] In this scenario, the availability of a spatial neighbor block can be determined using its reference image information, prediction mode information, and location information. For example, if the reference image of a spatial neighbor block is the same as the reference image of the current block, then the corresponding spatial neighbor block is considered available. Alternatively, if the spatial neighbor block is compiled using an intra-frame prediction mode or is located outside the current image / tile, then the corresponding spatial neighbor block is considered unavailable.
[0137] Furthermore, the search order for spatially adjacent blocks can be defined differently, and can be in, for example, the order of A1, B1, B0, and A0. Alternatively, the availability of A1 can be determined by searching only A1.
[0138] Figure 10 This is a diagram used to schematically describe the process of deriving sub-block-based temporal motion information candidates (sbTMVP candidates).
[0139] refer to Figure 10 First, the encoding / decoding device can determine whether the spatial neighboring block (e.g., block A1) of the current block is available. For example, if the reference image of the spatial neighboring block (e.g., block A1) uses a col image, it can be determined that the spatial neighboring block (e.g., block A1) is available, and the motion vector of the spatial neighboring block (e.g., block A1) can be derived. In this case, the motion vector of the spatial neighboring block (e.g., block A1) can be represented as a time MV (tempMV), and the motion vector can be used in motion shifting. Alternatively, if it is determined that the spatial neighboring block (e.g., block A1) is not available, the time MV (i.e., the motion vector of the spatial neighboring block) can be set to a zero vector. In other words, in this case, a motion vector set to (0,0) can be applied to motion shifting.
[0140] Next, the encoding / decoding device can apply motion shift based on the motion vector of a spatially neighboring block (e.g., block A1). For example, the motion shift can be shifted (e.g., A1') to the position indicated by the motion vector of the spatially neighboring block (e.g., block A1). That is, by applying motion shift, the motion vector of the spatially neighboring block (e.g., block A1) can be added to the coordinates of the current block.
[0141] Next, the encoder / decoder can export the motion-shifted juxtaposed sub-blocks (col sub-blocks) on the col image and obtain motion information (motion vectors, reference indices, etc.) for each col sub-block. For example, the encoder / decoder can export each col sub-block on the col image corresponding to the motion-shifted position at each sub-block position within the current block (i.e., the position indicated by the motion vector of the spatially neighboring block (e.g., A1)). Furthermore, the motion information of each col sub-block can be used as motion information for each sub-block of the current block (i.e., sbTMVP candidates).
[0142] Furthermore, scaling can be applied to the motion vectors of the `col` sub-blocks. Scaling can be performed based on the temporal distance difference between the reference image of the `col` sub-block and the reference image of the current block. Therefore, scaling can be represented as temporal motion scaling, thus allowing the placement of the reference image of the current block and the reference image of the temporal motion vector. In this case, the encoding / decoding device can obtain the scaled motion vectors of the `col` sub-blocks as motion information for each sub-block of the current block.
[0143] Furthermore, in deriving sbTMVP candidates, motion information may not exist in the col sub-block. In this case, for the col sub-block where motion information is absent, basic motion information (or default motion information) can be derived. Basic motion information can be used as motion information for sub-blocks within the current block. Basic motion information can be derived from the block located at the center of the col block (i.e., the colCU that includes the col sub-block). For example, motion information (e.g., motion vectors) can be derived from the block that includes the sample located in the lower right of the four samples located at the center of the col block, and this can be used as basic motion information.
[0144] As described above, in the case of affine mode or sbTMVP mode where motion information is derived on a sub-block basis, affine merge candidates and sbTMVP candidates can be derived, and a sub-block-based merge candidate list can be configured based on these candidates. In this case, flag information indicating whether the affine mode or sbTMVP mode is enabled or disabled can be sent using signals. If the sbTMVP mode is enabled based on the flag information, the sbTMVP candidates derived as described above can be added to the first order of the sub-block-based merge candidate list. Furthermore, affine merge candidates can be added to the next entry in the sub-block-based merge candidate list. In this case, the maximum number of candidates in the sub-block-based merge candidate list can be 5.
[0145] Furthermore, in sbTMVP mode, the size of sub-blocks can be fixed, and can be fixed to, for example, 8x8. Additionally, sbTMVP mode can be applied only to blocks with a width and height equal to or greater than 8.
[0146] Meanwhile, in the current VVC standard, as shown in Table 1, candidates for temporal motion information based on sub-blocks (sbTMVP candidates) can be derived.
[0147] [Table 1]
[0148]
[0149]
[0150]
[0151] When deriving sbTMVP candidates according to the method illustrated in Table 1, default MV and sub-block MV can be considered. In this case, the default MV can be referred to as sub-block-based time-merged basic motion data or basic motion vectors (basic motion information). Referring to Table 1, the default MV can correspond to ctrMV (or ctrMVLX) in Table 1. The sub-block MV can correspond to mvSbCol (or mvLXSbcol) in Table 1.
[0152] For example, if a sub-block or sub-block MV is available according to the sbTMVP export process, the sub-block MV can be assigned to the corresponding sub-block; or if a sub-block or sub-block MV is not available, the default MV can be used as the corresponding sub-block MV for the corresponding sub-block. In this case, the default MV can derive motion information from the position corresponding to the center pixel position of the corresponding block (i.e., colCU) on the col image, and each sub-block MV can derive motion information from the upper-left position of the corresponding sub-block (i.e., col sub-block) on the col image. In this case, the corresponding block (i.e., colCU) can be derived from the motion shift position based on the motion vector (i.e., time MV) of the spatially neighboring block A1, as described above. Figure 11 As described in [the text].
[0153] Figure 11 This is a diagram used to schematically illustrate the method for calculating the corresponding positions of the default MV and sub-block MV based on the block size during the sbTMVP export process.
[0154] Figure 11 The pixels (samples) drawn by dashed lines indicate the corresponding positions of each sub-block used to derive each sub-block MV, and the pixels (samples) drawn by solid lines illustrate the corresponding positions of the CUs used to derive the default MV.
[0155] For example, refer to Figure 11 (a) If the current block (i.e., the current CU) has a size of 8×8, the motion information of the sub-block can be derived based on the top-left sample position within the sub-block with a size of 8×8, and the default motion information of the sub-block can be derived based on the center sample position within the current block (i.e., the current CU) with a size of 8×8.
[0156] Alternative locations, for example, refer to Figure 11 (b) If the current block (i.e., the current CU) has a size of 16×8, the motion information of each sub-block can be derived based on the top-left sample position within each sub-block with a size of 8×8, and the default motion information of each sub-block can be derived based on the center sample position within the current block (i.e., the current CU) with a size of 16×8.
[0157] Alternative locations, for example, refer to Figure 11(c) If the current block (i.e., the current CU) has a size of 8×16, the motion information of each sub-block can be derived based on the top-left sample position within each sub-block with a size of 8×8, and the default motion information of each sub-block can be derived based on the center sample position within the current block (i.e., the current CU) with a size of 8×16.
[0158] Alternative locations, for example, refer to Figure 11 (d) If the current block (i.e., the current CU) has a size of 16×16, the motion information of each sub-block can be derived based on the top-left sample position within each sub-block with a size of 8×8, and the default motion information of each sub-block can be derived based on the center sample position within the current block (i.e., the current CU) with a size of 16×16.
[0159] from Figure 11 As can be seen, because the motion information of the sub-blocks tends to be skewed towards the top-left pixel position, the following problem exists: the sub-block MV is exported at a position far from the default MV that indicates the representative motion information of the current CU. As an example, in Figure 11 In the case of the 8×8 block shown in (a), a CU includes a sub-block, but there is a contradiction that the sub-block MV and the default MV are represented with different motion information. In addition, since the corresponding positions of the sub-block and the current CU block are calculated in different ways (i.e., the corresponding position used to derive the MV of the sub-block is the top-left sample position, and the corresponding position used to derive the default MV is the center sample position), additional modules may be required in the hardware (H / W) implementation.
[0160] Therefore, to improve the problem, this document proposes a scheme for uniformly deriving the corresponding positions of the CUs for the default MV and the corresponding positions of the sub-blocks for each sub-block MV during the process of deriving sbTMVP candidates. According to the embodiments in this document, the uniform effect is that, from a hardware (H / W) point of view, only one module can be used to derive each corresponding position based on the block size. For example, since the method for calculating the corresponding positions can be implemented identically if the block size is 16×16 and if the block size is 8×8, there is a simplification effect in terms of hardware implementation. In this case, a 16×16 block can represent the CU, and an 8×8 block can represent each sub-block.
[0161] As an example, in exporting sbTMVP candidates, the center sample position can be used as the corresponding position for exporting the motion information of the sub-block and the corresponding position for exporting the default motion information, and can be implemented as shown in Table 2 below.
[0162] Table 2 below illustrates an example of a method for deriving motion information and default motion information of a sub-block according to an embodiment of this document.
[0163] [Table 2]
[0164]
[0165]
[0166]
[0167] Referring to Table 2, when deriving sbTMVP candidates, the position of the current block (i.e., the current CU) including its sub-blocks can be derived. The top-left sample position (xCtb, yCtb) of the compilation tree block (or compilation tree unit) including the current block and the bottom-right center sample position (xCtr, yCtr) of the current block can be derived from equations (8-514) to (8-517) in Table 2. In this case, the positions (xCtb, yCtb) and (xCtr, yCtr) can be calculated based on the top-left sample position (xCb, yCb) of the current block relative to the top-left sample of the current image.
[0168] Furthermore, it is possible to export the col block (i.e., col CU) on the col image that corresponds to the current block (i.e., the current CU) including the sub-blocks. In this case, the position of the col block can be set to (xColCtrCb, yColCtrCb). The position can represent the location of the col block, which includes the position (xCtr, yCtr) within the col image relative to the top-left sample of the col image.
[0169] Furthermore, basic motion data (i.e., default motion information) for sbTMVP can be exported. Basic motion data can include a default motion vector (MV) (e.g., CtrMvLX). For example, a col block on a col image can be exported. In this case, the position of the col block can be exported as (xColCb, yColCb). This position can be the position where a motion shift (e.g., tempMv) has been applied to the exported col block position (xColCtrCb, yColCtrCb). As described above, motion shift can be performed by adding a motion vector (e.g., tempMv) derived from a spatially neighboring block (e.g., block A1) to the current col block position (xColCtrCb, yColCtrCb). Next, a default MV (e.g., ctrMvLX) can be derived based on the position of the motion-shifted col block (xColCb, yColCb). In this case, the default MV (e.g., ctrMvLX) can represent a motion vector derived from the position corresponding to the lower right center sample of the col block.
[0170] Furthermore, the col sub-blocks on the col image corresponding to the sub-blocks in the current block (denoted as the current sub-block) can be derived. First, the position of each current sub-block can be derived. The position of each sub-block can be represented as (xSb, ySb). The position (xSb, ySb) can represent the position of the current sub-block based on the upper left sample of the current image. For example, the position (xSb, ySb) of the current sub-block can be calculated as shown in equations (8-523) to (8-524) in Table 2, which can represent the lower right center sample position of the sub-block. Next, the position of each col sub-block in the col sub-blocks on the col image can be derived. The position of each col sub-block can be represented as (xColSb, yColSb). The position (xColSb, yColSb) can be the position of the current sub-block's position (xSb, ySb) after applying a motion shift (e.g., tempMv). As described above, motion shifting can be performed by adding a motion vector (e.g., tempMv) derived from the spatial neighboring block of the current block (e.g., block A1) to the position (xSb, ySb) of the current sub-block. Next, motion information (e.g., motion vector mvLXSbCol, availableFlagLXSbCol) of the col sub-block can be derived based on the position (xColSb, yColSb) of each of the motion-shifted col sub-blocks.
[0171] In this case, if a col sub-block is unavailable (e.g., when availableFlagLXSbCol is 0), the basic motion data (i.e., the default motion information) can be used for the unavailable col sub-block. For example, the default MV (e.g., ctrMvLX) can be used as the motion vector (e.g., mvLXSbCol) for the unavailable col sub-block.
[0172] Figure 12 This is an example diagram used to schematically illustrate a unified method for exporting the corresponding positions of the default MV and sub-block MV based on the block size during the sbTMVP export process.
[0173] Figure 12 The pixels (samples) marked with dashed lines indicate the corresponding positions within each sub-block used to derive the MV of each sub-block, while the pixels (samples) marked with solid lines illustrate the corresponding positions of the CU used to derive the default MV.
[0174] For example, refer to Figure 12(a) If the current block (i.e., the current CU) has a size of 8×8, motion information can be derived from the corresponding col sub-block in the col image based on the bottom-right center sample position within the 8×8-sized sub-block, and can be used as the motion information for the current sub-block. Motion information can be derived from the corresponding col block (i.e., colCU) in the col image based on the bottom-right center sample position within the current block (i.e., the current CU) with a size of 8×8, and can be used as the default motion information for the current sub-block. In this case, as... Figure 16 As shown, the motion information of the current sub-block and the default motion information can be derived from the same sample position (the same corresponding position).
[0175] Alternative locations, for example, refer to Figure 12 (b) If the current block (i.e., the current CU) has a size of 16×8, then motion information can be derived from the corresponding position of the col sub-block in the col image based on the bottom-right center sample position within the sub-block of size 8×8, and can be used as the motion information for the current sub-block. Motion information can be derived from the corresponding position of the col block (i.e., colCU) in the col image based on the bottom-right center sample position within the current block (i.e., the current CU) of size 16×8, and can be used as the default motion information for the current sub-block.
[0176] Alternative locations, for example, refer to Figure 12 (c) If the current block (i.e., the current CU) has a size of 8×16, then motion information can be derived from the corresponding position of the col sub-block in the col image based on the bottom-right center sample position within the 8×8 sub-block, and can be used as the motion information for the current sub-block. Motion information can also be derived from the corresponding position of the col block (i.e., colCU) in the col image based on the bottom-right center sample position within the current block (i.e., the current CU) with a size of 8×16, and can be used as the default motion information for the current sub-block.
[0177] Alternative locations, for example, refer to Figure 12 (d) If the current block (i.e., the current CU) has a size equal to or greater than 16×16, then motion information can be derived from the corresponding position of the col sub-block in the col image based on the bottom-right center sample position within the sub-block of size 8×8, and can be used as the motion information for the current sub-block. Motion information can also be derived from the corresponding position of the col block (i.e., col CU) in the col image based on the bottom-right center sample position within the current block (i.e., the current CU) of size 16×16 (or 16×16 or larger), and can be used as the default motion information for the current sub-block.
[0178] However, the embodiments described above in this document are merely examples, and the default motion information and the motion information of the current sub-block can be derived based on a sample position other than the center position (i.e., the lower right sample position). For example, the default motion information can be derived based on the upper left sample position of the current CU, and the motion information of the current sub-block can be derived based on the upper left sample position of the sub-block.
[0179] If the embodiments in this document are implemented as hardware, then configurations such as Figure 13 and 14 The pipeline can export motion information (time motion) using the same H / W module.
[0180] Figure 13 and Figure 14 This is an example diagram illustrating the configuration of a pipeline that allows for the unified calculation of the corresponding positions used to export the default MV and sub-block MV during the sbTMVP export process.
[0181] refer to Figure 13 and Figure 14 The corresponding position calculation module can calculate the corresponding positions used to derive the default MV and sub-block MV. For example, such as... Figure 13 and Figure 14 As shown, when the block position (posX, posY) and block size (blkszX, blkszY) are input to the corresponding position calculation module, the center position of the input block (i.e., the lower right sample position) can be output. When the current CU position and block size are input to the corresponding position calculation module, the center position of the col block on the col image (i.e., the lower right sample position) can be output, which is used to derive the corresponding position of the default MV. Alternatively, when the current sub-block position and block size are input to the corresponding position calculation module, the center position of the col sub-block on the col image (i.e., the lower right sample position) can be output, which is used to derive the corresponding position of the current sub-block MV.
[0182] As described above, when the corresponding position is output from the corresponding position calculation module for deriving the default MV and sub-block MV, the motion vector (i.e., time mv) derived from the corresponding position can be repaired. Furthermore, sub-block-based temporal motion information (i.e., sbTMVP candidate) can be derived based on the repaired motion vector (i.e., time mv). For example, as in... Figure 13 and 14 Depending on the H / W implementation, sbTMVP candidates can be derived in parallel based on clock cycles, or they can be derived sequentially.
[0183] The following figures are provided to illustrate detailed examples of this document. The names of detailed devices or detailed terms or names (e.g., names of grammars / grammar names) written in the figures are exemplary, and therefore the technical features of this document are not limited to the detailed names used in the following figures.
[0184] Figure 15 and Figure 16 Examples of video / image encoding methods and related components according to embodiments of the present disclosure are illustrated schematically.
[0185] Figure 15 The method disclosed in the article can be derived from Figure 2 The encoding device 200 disclosed herein is executed. Specifically, Figure 15 Steps S1500 to S1540 can be performed by Figure 2 The predictor 220 (more specifically, inter-frame predictor 221) disclosed in the document executes, Figure 15 Step S1550 can be performed by Figure 2 The residual processor 230 disclosed in the paper executes, and Figure 15 Step S1560 can be performed by Figure 2 The entropy encoder 240 disclosed in the document is executed. Furthermore... Figure 15 The methods disclosed herein may include the embodiments described above. Therefore, in Figure 15 In this document, any redundant detailed descriptions related to the embodiments will be omitted or briefly omitted.
[0186] refer to Figure 15 The encoding device can derive the positions of the sub-blocks included in the current block (S1500).
[0187] Here, the current block can be referred to as the current compilation unit CU or the current compilation block CB, and the sub-blocks included in the current block can be referred to as the current compilation sub-blocks.
[0188] In an embodiment, the encoding device can derive the position of the current sub-block within the current block.
[0189] For example, the encoding device can derive the position of the current sub-block in the current image based on the center sample position of the current sub-block. In this case, the center sample position can represent the position of the bottom-right center sample among the four samples located at the center.
[0190] Meanwhile, the upper left sample position used in this disclosure can be referred to as the top-left sample position or the upper-left sample position, etc., and the lower right center sample position can be referred to as the lower right center sample position, the center lower right sample position, the bottom right center sample position, or the center bottom right sample position, etc.
[0191] The encoding device can derive a reference sub-block on the juxtaposed reference picture for the sub-blocks within the current block (S1510).
[0192] Here, the juxtaposed reference image refers to the reference image used to derive the time motion information (i.e., sbTMVP) as described above, and can represent the aforementioned col image. The reference sub-block can represent the aforementioned col sub-block.
[0193] In an embodiment, the encoding device can derive a reference sub-block on the juxtaposed reference image based on the position of the current sub-block within the current block. For example, the encoding device can derive a reference sub-block on the juxtaposed reference image based on the center sample position of the current sub-block (e.g., the lower right center sample position).
[0194] For example, the encoding device can first specify the position of the current block, and then specify the positions of the sub-blocks within the current block. As explained in Table 2 above, the position of the current block can be represented based on the top-left sample position (xCtb, yCtb) of the compilation tree block and the bottom-right center sample position (xCtr, yCtr) of the current block. The position of the current sub-block within the current block can be represented as (xSb, ySb), and this position (xSb, ySb) can represent the bottom-right center sample position of the current sub-block. Here, the bottom-right center sample position (xSb, ySb) of the sub-block can be calculated based on the top-left sample position and the size of the sub-block, and can be calculated as in Equations 8-523 and 8-524 in Table 2 above.
[0195] Furthermore, the encoding device can derive a reference sub-block on the juxtaposed reference image based on the lower right center sample position of the current sub-block within the current block. As explained in Table 2 above, the reference sub-block can be represented as a position (xColSb, yColSb) on the juxtaposed reference image, and the position (xColSb, yColSb) on the juxtaposed reference image can be derived based on the lower right center sample position (xSb, ySb) of the current sub-block within the current block.
[0196] Furthermore, motion shifting can be applied when deriving a reference sub-block. The encoding device can perform motion shifting based on motion vectors derived from spatially neighboring blocks of the current block. The spatially neighboring block of the current block can be the left-hand neighboring block located to the left of the current block (e.g., Figure 9 and Figure 10(e.g., block A1 as depicted in the image). In this case, if the left neighboring block (e.g., block A1) is available, the motion vector can be derived from the left neighboring block, or if the left neighboring block is unavailable, the zero vector can be derived. Here, the availability of the spatial neighboring block can be determined by the reference image information, prediction mode information, position information, etc. of the spatial neighboring block. For example, if the reference image of the spatial neighboring block is the same as the reference image of the current block, it can be determined that the spatial neighboring block is available. Alternatively, if the spatial neighboring block is compiled with an intra-frame prediction mode or the spatial neighboring block is located outside the current image / tile, it can be determined that the spatial neighboring block is unavailable.
[0197] In other words, the encoding device can apply a motion shift (i.e., the motion vector of a spatially neighboring block (e.g., block A1)) to the lower right center sample position (xSb, ySb) of the current sub-block within the current block, and can derive and place a reference sub-block on the reference image based on the motion shift position. In this case, the position (xColSb, yColSb) of the reference sub-block can be represented as the position obtained by motion shifting from the lower right center sample position (xSb, ySb) of the current sub-block within the current block to the position indicated by the motion vector of the spatially neighboring block (e.g., block A1), and can be calculated as in equations 8-525 and 8-526 in Table 2 above.
[0198] The encoding device can derive sbTMVP (sub-block time motion vector predictor) candidates (S1520) based on reference sub-blocks.
[0199] Furthermore, in this disclosure, the sbTMVP candidate can be used interchangeably or alternatively with sub-block-based temporal motion information candidates, sub-block unit temporal motion information candidates, sub-block-based temporal motion vector predictor candidates, etc. That is, if motion information is derived for each sub-block to perform the prediction as described above, sbTMVP candidates can be derived, and motion prediction can be performed at the sub-block level (or sub-compilation unit (sub-CU) level) based on sbTMVP candidates.
[0200] In an embodiment, the encoding device can derive sbTMVP candidates based on the motion vectors of a reference sub-block. For example, the encoding device can derive sbTMVP candidates based on the motion vectors of a reference sub-block derived from whether the reference sub-block is available. If the reference sub-block is available, the motion vectors of the available reference sub-blocks can be derived as sbTMVP candidates. If the reference sub-block is unavailable, the basic motion vectors can be derived as sbTMVP candidates.
[0201] Here, the basic motion vector can correspond to the default motion vector mentioned above, and can be derived on the juxtaposed reference image based on the position of the current block. In this case, the position of the current block can be derived based on the center sample position within the current block (e.g., the lower right center sample position).
[0202] In deriving the basic motion vector, in an embodiment, the encoding device can specify the position of a reference compilation block on a juxtaposed reference image based on the lower right center sample position of the current block, and derive the basic motion vector based on the position of the reference compilation block. The reference compilation block can refer to a col block located on a juxtaposed reference image corresponding to the current block, which includes sub-blocks. As explained with reference to Table 2 above, the position of the reference compilation block can be represented as (xColCtrCb, yColCtrCb), and the position (xColCtrCb, yColCtrCb) can represent the position of the reference compilation block covering the position (xCtr, yCtr) within the juxtaposed reference image relative to the upper left sample. The position (xCtr, yCtr) can represent the lower right center sample position of the current block.
[0203] Furthermore, in deriving the basic motion vector, motion shifting can be applied to the position (xColCtrCb, yColCtrCb) of the reference compilation block. Motion shifting can be performed by adding the motion vector derived from the spatially neighboring block (e.g., block A1) of the current block, as described above, to the position (xColCtrCb, yColCtrCb) of the reference compilation block covering the lower right center sample. The encoding apparatus can derive the basic motion vector based on the position (xColCb, yColCb) of the motion-shifted reference compilation block. That is, the basic motion vector can be a motion vector derived from the motion-shifted position on the juxtaposed reference image, based on the lower right center sample position of the current block.
[0204] Additionally, the availability of a reference sub-block can be determined based on whether it is located outside the juxtaposed reference image or based on motion vectors. For example, unavailable reference sub-blocks may include reference sub-blocks located outside the juxtaposed reference image or reference sub-blocks whose motion vectors are unavailable. For example, if a reference sub-block is based on intra-frame mode, IBC (Intra-Block Copy) mode, or palette mode, then the reference sub-block may be a sub-block whose motion vectors are unavailable. Alternatively, if a reference compilation block that overrides a modified position derived from the location of the reference sub-block is based on intra-frame mode, IBC mode, or palette mode, then the reference sub-block may be a sub-block whose motion vectors are unavailable.
[0205] In this case, as an example, the motion vector of the available reference sub-block can be derived based on the motion vector of the block whose modified position is derived from the top-left sample position of the reference sub-block. For example, as shown in Table 2 above, the modified position can be derived by the equation ((xColSb>>3)<<3,(yColSb>>3)<<3). Here, xColSb and yColSb can represent the x-coordinate and y-coordinate of the top-left sample position of the reference sub-block, respectively, and >> can represent an arithmetic right shift, and << can represent an arithmetic left shift.
[0206] Furthermore, as mentioned above, in deriving the sbTMVP candidate, it can be seen that motion vectors for reference sub-blocks are derived based on the positions of sub-blocks within the current block, and basic motion vectors are derived based on the position of the current block. For example, as... Figure 12 As shown, for a current block of size 8×8, motion vectors for reference sub-blocks and basic motion vectors can be derived based on the bottom right center sample position of the current block. For a current block larger than 8×8, motion vectors for reference sub-blocks can be derived based on the bottom right center sample position of sub-blocks within the current block, and basic motion vectors can also be derived based on the bottom right center sample position of the current block.
[0207] The encoding device can derive motion information for sub-blocks within the current block based on sbTMVP candidates (S1530).
[0208] In an embodiment, the encoding device can derive the motion vector of a reference sub-block into motion information (e.g., motion vector) of the current sub-block within the current block. As described above, the encoding device can derive sbTMVP candidates based on the motion vectors of available reference sub-blocks or basic motion vectors, and can use the motion vectors derived as sbTMVP candidates as motion vectors for the current sub-block.
[0209] The encoding device can generate a prediction sample for the current block based on the motion information of the sub-blocks within the current block (S1540).
[0210] In this embodiment, the encoding device can generate prediction samples based on the motion vector of the current sub-block. Specifically, the encoding device can select the optimal motion information based on the RD (rate distortion) cost and generate prediction samples based on that information. For example, if the motion information derived for each sub-block of the current block (i.e., sbTMVP) is selected as the optimal motion information, the encoding device can generate prediction samples for the current block based on the motion information derived for the sub-blocks of the current block.
[0211] The encoding device can generate information about residual samples derived from the predicted samples (S1550), and can encode image information including information about the residual samples (S1560).
[0212] In other words, the encoding device can derive residual samples based on the initial samples and predicted samples of the current block. Furthermore, the encoding device can generate information about the residual samples. This information can include information derived by performing transformations and quantizations on the residual samples (such as value information, location information, transformation scheme, transformation kernel, and quantization parameters of the quantization transformation coefficients).
[0213] The encoding device can encode information about the residual samples and output it as a bit stream, which can be sent to the decoding device via a network or storage medium.
[0214] Figure 17 and Figure 18 Examples of video / image decoding methods and related components according to embodiments of the present disclosure are illustrated schematically.
[0215] Figure 17 The method disclosed in the article can be derived from Figure 3 The decoding device 300 disclosed in the document executes this. Specifically, Figure 17 Steps S1700 to S1740 can be performed by Figure 3 The predictor 330 (more specifically, the inter-frame predictor 332) disclosed in the document is executed, and Figure 17 Step S1750 can be performed by Figure 3 The adder 340 disclosed in the document is executed. Furthermore... Figure 17 The methods disclosed herein may include the embodiments described above. Therefore, in Figure 17 In this document, any redundant detailed descriptions related to the embodiments will be omitted or briefly omitted.
[0216] refer to Figure 17 The decoding device can derive the positions of the sub-blocks included in the current block (S1700).
[0217] Here, the current block can be referred to as the current compilation unit CU or the current compilation block CB, and the sub-blocks included in the current block can be referred to as the current compilation sub-blocks.
[0218] In one embodiment, the decoding device can derive the position of the current sub-block within the current block.
[0219] For example, the decoding device can derive the position of the current sub-block in the current image based on the center sample position of the current sub-block. In this case, the center sample position can represent the position of the bottom-right center sample among the four samples located at the center.
[0220] Meanwhile, the upper left sample position used in this disclosure can be referred to as the top-left sample position or the upper-left sample position, etc., and the lower right center sample position can be referred to as the lower right center sample position, the center lower right sample position, the bottom right center sample position, or the center bottom right sample position, etc.
[0221] The decoding device can export a reference sub-block on the juxtaposed reference picture for the sub-blocks within the current block (S1710).
[0222] Here, the juxtaposed reference image refers to the reference image used to derive the time motion information (i.e., sbTMVP) as described above, and can represent the col image mentioned above. The reference sub-block can represent the col sub-block mentioned above.
[0223] In an embodiment, the decoding device can derive a reference sub-block on the juxtaposed reference image based on the position of the current sub-block within the current block. For example, the decoding device can derive a reference sub-block on the juxtaposed reference image based on the center sample position of the current sub-block (e.g., the lower right center sample position).
[0224] For example, the decoding device can first specify the position of the current block, and then specify the positions of the sub-blocks within the current block. As explained in Table 2 above, the position of the current block can be represented based on the top-left sample position (xCtb, yCtb) of the compilation tree block and the bottom-right center sample position (xCtr, yCtr) of the current block. The position of the current sub-block within the current block can be represented as (xSb, ySb), and this position (xSb, ySb) can represent the bottom-right center sample position of the current sub-block. Here, the bottom-right center sample position (xSb, ySb) of the sub-block can be calculated based on the top-left sample position of the sub-block and the size of the sub-block, and can be calculated as in equations 8-523 and 8-524 of Table 2 above.
[0225] Furthermore, the decoding device can derive a reference sub-block on the juxtaposed reference image based on the lower right center sample position of the current sub-block within the current block. As explained in Table 2 above, the reference sub-block can be represented as a position (xColSb, yColSb) on the juxtaposed reference image, and its position (xColSb, yColSb) on the juxtaposed reference image can be derived based on the lower right center sample position (xSb, ySb) of the current sub-block within the current block.
[0226] Furthermore, motion shifting can be applied when deriving a reference sub-block. The decoding device can perform motion shifting based on motion vectors derived from spatially neighboring blocks of the current block. A spatially neighboring block of the current block can be a left-neighboring block located to the left of the current block (e.g., Figure 9and Figure 10 (e.g., block A1 as depicted in the image). In this case, if the left neighboring block (e.g., block A1) is available, the motion vector can be derived from the left neighboring block, or if the left neighboring block is unavailable, the zero vector can be derived. Here, the availability of the spatial neighboring block can be determined by the reference image information, prediction mode information, position information, etc. of the spatial neighboring block. For example, if the reference image of the spatial neighboring block is the same as the reference image of the current block, it can be determined that the spatial neighboring block is available. Alternatively, if the spatial neighboring block is compiled with an intra-frame prediction mode or the spatial neighboring block is located outside the current image / tile, it can be determined that the spatial neighboring block is unavailable.
[0227] In other words, the decoding device can apply motion shift (i.e., the motion vector of a spatially neighboring block (e.g., block A1)) to the lower right center sample position (xSb, ySb) of the current sub-block within the current block, and can derive and place a reference sub-block on the reference image based on the motion shift position. In this case, the position (xColSb, yColSb) of the reference sub-block can be represented as the position obtained by motion shifting from the lower right center sample position (xSb, ySb) of the current sub-block within the current block to the position indicated by the motion vector of the spatially neighboring block (e.g., block A1), and can be calculated as in equations 8-525 and 8-526 in Table 2 above.
[0228] The decoding device can derive sbTMVP (sub-block time motion vector predictor) candidates (S1720) based on the reference sub-block.
[0229] Furthermore, in this disclosure, the sbTMVP candidate can be used interchangeably or alternatively with sub-block-based temporal motion information candidates, sub-block unit temporal motion information candidates, sub-block-based temporal motion vector predictor candidates, etc. That is, if motion information is derived for each sub-block to perform the prediction as described above, sbTMVP candidates can be derived, and motion prediction can be performed at the sub-block level (or sub-compilation unit (sub-CU) level) based on sbTMVP candidates.
[0230] In an embodiment, the decoding device can derive sbTMVP candidates based on the motion vectors of a reference sub-block. For example, the decoding device can derive sbTMVP candidates based on the motion vectors of a reference sub-block derived from whether the reference sub-block is available. If the reference sub-block is available, the motion vectors of the available reference sub-blocks can be derived as sbTMVP candidates. If the reference sub-block is unavailable, the basic motion vectors can be derived as sbTMVP candidates.
[0231] Here, the basic motion vector can correspond to the default motion vector mentioned above, and can be derived on the juxtaposed reference image based on the position of the current block. In this case, the position of the current block can be derived based on the center sample position within the current block (e.g., the lower right center sample position).
[0232] In deriving the basic motion vector, in this embodiment, the decoding device can specify the position of a reference compilation block on the juxtaposed reference image based on the lower right center sample position of the current block, and derive the basic motion vector based on the position of the reference compilation block. The reference compilation block can refer to a col block located on the juxtaposed reference image corresponding to the current block, which includes sub-blocks. As explained with reference to Table 2 above, the position of the reference compilation block can be represented as (xColCtrCb, yColCtrCb), and the position (xColCtrCb, yColCtrCb) can represent the position of the reference compilation block covering the position (xCtr, yCtr) within the juxtaposed reference image relative to the upper left sample. The position (xCtr, yCtr) can represent the lower right center sample position of the current block.
[0233] Furthermore, in deriving the basic motion vector, motion shifting can be applied to the position (xColCtrCb, yColCtrCb) of the reference compilation block. Motion shifting can be performed by adding the motion vector derived from the spatially neighboring block (e.g., block A1) of the current block, as described above, to the position (xColCtrCb, yColCtrCb) of the reference compilation block covering the lower right center sample. The decoding apparatus can derive the basic motion vector based on the position (xColCb, yColCb) of the motion-shifted reference compilation block. That is, the basic motion vector can be a motion vector derived from the motion-shifted position on the juxtaposed reference image, based on the lower right center sample position of the current block.
[0234] Additionally, the availability of a reference sub-block can be determined based on whether it is located outside the juxtaposed reference image or based on motion vectors. For example, unavailable reference sub-blocks may include reference sub-blocks located outside the juxtaposed reference image or reference sub-blocks whose motion vectors are unavailable. For example, if a reference sub-block is based on intra-frame mode, IBC (Intra-Block Copy) mode, or palette mode, then the reference sub-block may be a sub-block whose motion vectors are unavailable. Alternatively, if a reference compilation block that overrides a modified position derived from the location of the reference sub-block is based on intra-frame mode, IBC mode, or palette mode, then the reference sub-block may be a sub-block whose motion vectors are unavailable.
[0235] In this case, as an example, the motion vector of the available reference sub-block can be derived based on the motion vector of the block whose modified position is derived from the top-left sample position of the reference sub-block. For example, as shown in Table 2 above, the modified position can be derived by the equation ((xColSb>>3)<<3,(yColSb>>3)<<3). Here, xColSb and yColSb can represent the x-coordinate and y-coordinate of the top-left sample position of the reference sub-block, respectively, and >> can represent an arithmetic right shift, and << can represent an arithmetic left shift.
[0236] Furthermore, as mentioned above, in deriving the sbTMVP candidate, it can be seen that the motion vector of the reference sub-block is derived based on the position of the sub-block within the current block, and the basic motion vector is derived based on the position of the current block. For example, as... Figure 12 As shown, for a current block of size 8×8, the motion vector for the reference sub-block and the basic motion vector can be derived based on the bottom right center sample position of the current block. For a current block larger than 8×8, the motion vector for the reference sub-block can be derived based on the bottom right center sample position of the sub-blocks within the current block, and the basic motion vector can also be derived based on the bottom right center sample position of the current block.
[0237] The decoding device can derive motion information for sub-blocks within the current block based on the sbTMVP candidate (S1730).
[0238] In an embodiment, the decoding device can derive the motion vector of a reference sub-block into motion information (e.g., motion vector) of the current sub-block within the current block. As described above, the decoding device can derive sbTMVP candidates based on the motion vector of the available reference sub-block or the basic motion vector, and can use the motion vector derived as the sbTMVP candidate as the motion vector for the current sub-block.
[0239] The decoding device can generate a prediction sample for the current block based on the motion vector of the current sub-block within the current block (S1740).
[0240] In an embodiment, in the case of a prediction mode that performs prediction based on the sub-block unit motion information for the current block (i.e., sbTMVP mode), the decoding device can generate a prediction sample for the current block based on the motion information for the current sub-block derived above.
[0241] The decoding device can generate reconstructed samples based on the predicted samples (S1750).
[0242] In this embodiment, the decoding device can directly use the predicted sample as the reconstructed sample according to the prediction pattern, or it can generate the reconstructed sample by adding the residual sample to the predicted sample.
[0243] If residual samples exist for the current block, the decoding device can receive information about the residuals used for the current block. This information may include transform coefficients associated with the residual samples. Based on the residual information, the decoding device can derive residual samples (or an array of residual samples) for the current block. The decoding device can generate reconstructed samples based on the predicted samples and residual samples, and derive a reconstructed block or reconstructed image based on the reconstructed samples. Then, the decoding device can apply in-loop filtering processes such as deblocking filtering and / or SAO processes to the reconstructed image as described above to improve subjective / objective image quality where necessary.
[0244] In the embodiments described above, although these methods have been described based on flowcharts in the form of a series of steps or units, the embodiments described herein are not limited to the order of these steps, and some of these steps may be performed in a different order than the other steps or may be performed simultaneously with the other steps. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and without affecting the scope of the claims of this document, these steps may include additional steps or one or more steps in the flowchart may be deleted.
[0245] The methods described above in this document can be implemented in software, and the encoding and / or decoding devices according to this document can be included in devices for performing image processing, such as TVs, computers, smartphones, set-top boxes, or display devices.
[0246] In this document, when the implementation is carried out in software form, the methods mentioned above can be implemented as modules (programs, functions, etc.) for performing the functions mentioned above. Modules can be stored in memory and executed by a processor. Memory can be located inside or outside the processor and connected to the processor by various known means. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. Memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the implementations described in this document can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the figures can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information about instructions) or algorithms used for such implementation can be stored in a digital storage medium.
[0247] Furthermore, the decoding and encoding devices using this document can be included in multimedia broadcasting transmitting and receiving equipment, mobile communication terminals, home theater video equipment, digital cinema video equipment, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, video telephony devices, transportation terminals (e.g., vehicle (including autonomous vehicle) terminals, aircraft terminals, and ship terminals), and medical video equipment, and can be used to process video signals or data signals. For example, over-the-top (OTT) video devices can include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0248] Furthermore, the processing methods described in this document can be generated in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data with data structures according to this document can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices storing computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. In addition, computer-readable recording media include media implemented in the form of a carrier wave (e.g., transmitted via the Internet). Furthermore, bitstreams generated using encoding methods can be stored in computer-readable recording media or transmitted via wired and wireless communication networks.
[0249] Furthermore, the embodiments described in this document can be implemented as a computer program product using program code. The program code can be executed by a computer according to the embodiments described in this document. The program code can be stored on a carrier wave that can be read by a computer.
[0250] Figure 19 The illustrations are examples of content streaming systems that can be applied to the embodiments disclosed in this document.
[0251] refer to Figure 19 The content streaming media system using embodiments of the present invention can basically include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.
[0252] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then transmits this bitstream to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted.
[0253] Bitstreams can be generated using the encoding methods or bitstream generation methods described in this document, and the streaming server can temporarily store the bitstreams during the sending or receiving of bitstreams.
[0254] A streaming server sends multimedia data to a user's device via a web server based on a user's request, and the web server acts as a medium for informing the user of services. When a user requests a desired service from the web server, the web server forwards it to the streaming server, which then sends the multimedia data to the user. In this scenario, the content streaming system may include a separate control server. In this case, the control server manages the commands / responses between devices within the content streaming system.
[0255] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.
[0256] Examples of user equipment may include mobile phones, smartphones, laptops, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, touchscreen PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc.
[0257] In a content streaming system, each server can operate as a distributed server, in which case the data received from each server can be distributed.
[0258] The claims described herein can be combined in various ways. For example, the technical features of the method claims of this specification can be combined to implement an apparatus, and the technical features of the apparatus claims of this specification can be combined and implemented as a method. Furthermore, the technical features of the method claims and the apparatus claims of this specification can be combined to implement an apparatus, and the technical features of the method claims and the apparatus claims of this specification can be combined and implemented as a method.
Claims
1. An image decoding method performed by a decoding device, the method comprising: Based on the center sample position of the current sub-block, derive the position of the current sub-block within the current block; Based on the position of the current sub-block, export and place the reference sub-block on the reference image; Based on the motion vector of the reference sub-block, a candidate sub-block time motion vector predictor (sbTMVP) is derived; Based on the sbTMVP candidate, the motion vector of the current sub-block is derived; Based on the motion vector of the current sub-block, generate a prediction sample for the current block; Based on the residual information obtained from the bit stream, a residual sample of the current block is generated; as well as Based on the predicted samples and the residual samples, reconstructed samples are generated. Wherein, the center sample position refers to the position of the bottom right center sample among the four samples located at the center. Specifically, for the available reference sub-block, the motion vector of the available reference sub-block is derived as the sbTMVP candidate. Specifically, based on the motion vector of the block whose modified position is derived from the upper left sample position of the reference sub-block, the motion vector of the usable reference sub-block is derived, and... The modified position is derived by the equation ((xColSb >> 3) << 3, (yColSb >> 3) << 3), where xColSb and yColSb represent the x and y coordinates of the sample position of the reference sub-block derived from the center sample position of the current sub-block, respectively, and >> represents an arithmetic right shift and << represents an arithmetic left shift.
2. An image encoding method performed by an encoding device, the method comprising: Based on the center sample position of the current sub-block, derive the position of the current sub-block within the current block; Based on the position of the current sub-block, export and place the reference sub-block on the reference image; Based on the motion vector of the reference sub-block, a candidate sub-block time motion vector predictor (sbTMVP) is derived; Based on the sbTMVP candidate, the motion vector of the current sub-block is derived; Based on the motion vector of the current sub-block, generate a prediction sample for the current block; Generate residual samples derived from the predicted samples; as well as The image information, including information about the residual samples, is encoded. The central sample position refers to the position of the lower right center sample among the four samples located at the center. Specifically, for the available reference sub-block, the motion vector of the available reference sub-block is derived as the sbTMVP candidate. Specifically, based on the motion vector of the block whose modified position is derived from the upper left sample position of the reference sub-block, the motion vector of the usable reference sub-block is derived, and... The modified position is derived by the equation ((xColSb >> 3) << 3, (yColSb >> 3) << 3), where xColSb and yColSb represent the x and y coordinates of the sample position of the reference sub-block derived from the center sample position of the current sub-block, respectively, and >> represents an arithmetic right shift and << represents an arithmetic left shift.
3. A non-transitory computer-readable digital storage medium storing a bitstream generated by an image encoding method, the method comprising: Based on the center sample position of the current sub-block, derive the position of the current sub-block within the current block; Based on the position of the current sub-block, export and place the reference sub-block on the reference image; Based on the motion vector of the reference sub-block, a candidate sub-block time motion vector predictor (sbTMVP) is derived; Based on the sbTMVP candidate, the motion vector of the current sub-block is derived; Based on the motion vector of the current sub-block, generate a prediction sample for the current block; Generate residual samples derived from the predicted samples; as well as The image information, including information about the residual samples, is encoded to output the bitstream. The central sample position refers to the position of the lower right center sample among the four samples located at the center. Specifically, for the available reference sub-block, the motion vector of the available reference sub-block is derived as the sbTMVP candidate. Specifically, based on the motion vector of the block whose modified position is derived from the upper left sample position of the reference sub-block, the motion vector of the usable reference sub-block is derived, and... The modified position is derived by the equation ((xColSb >> 3) << 3, (yColSb >> 3) << 3), where xColSb and yColSb represent the x and y coordinates of the sample position of the reference sub-block derived from the center sample position of the current sub-block, respectively, and >> represents an arithmetic right shift and << represents an arithmetic left shift.