Image encoding / decoding method and apparatus, and recording medium storing bitstream

WO2026197745A1PCT designated stage Publication Date: 2026-09-24LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/004262
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-16
Filing Date
2026-03-16
Publication Date
2026-09-24

Smart Images

  • Figure KR2026004262_24092026_PF_FP_ABST
    Figure KR2026004262_24092026_PF_FP_ABST
Patent Text Reader

Abstract

An image decoding method and apparatus according to the present disclosure may determine a co-located block of a current block, obtain motion information of the co-located block, derive a temporal motion vector predictor for the current block on the basis of the motion information of the co-located block, and derive a prediction sample of the current block on the basis of the temporal motion vector predictor for the current block. Here, the motion information of the co-located block may be obtained on the basis of a position that is derived on the basis of a storage unit.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding / decoding method and device, and a recording medium storing a bitstream

[0001] The present invention relates to a video encoding / decoding method and apparatus, and a recording medium storing a bitstream.

[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition) video, has been increasing across various application fields, and accordingly, high-efficiency video compression technologies are being discussed.

[0003] Various image compression technologies exist, such as inter-prediction methods that predict pixel values ​​in the current picture from previous or subsequent pictures, intra-prediction methods that predict pixel values ​​in the current picture using pixel information within the current picture, and entropy coding techniques that assign short codes to values ​​with high frequency and long codes to values ​​with low frequency; by utilizing these image compression technologies, image data can be effectively compressed for transmission or storage.

[0004] The present disclosure provides a method and apparatus for deriving a temporal motion vector predictor (TMVP) of a current block.

[0005] An image decoding method and apparatus according to the present disclosure can determine a corresponding position block of a current block, acquire motion information of the corresponding position block, derive a time motion vector predictor of the current block based on the motion information of the corresponding position block, and derive a prediction sample of the current block based on the time motion vector predictor of the current block. Here, the motion information of the corresponding position block can be acquired based on a position derived based on a storage unit.

[0006] In the image decoding method and apparatus according to the present disclosure, the storage unit may be determined based on whether a specific condition is satisfied, and the specific condition may be a result of comparing the number of samples of the current picture with a predetermined number.

[0007] In the image decoding method and apparatus according to the present disclosure, the storage unit may be determined as a 4x4 unit when the number of samples in the current picture is less than a predetermined number, and as an 8x8 unit when the number of samples in the current picture is greater than or equal to the predetermined number.

[0008] In the image decoding method and apparatus according to the present disclosure, the position derived based on the storage unit may be a coordinate derived by performing a shift operation based on a storage unit variable on the reference position coordinate of the corresponding position block, and the reference position coordinate of the corresponding position block may be any one of the coordinate of a sample at a position corresponding to the coordinate of the lower-right position sample of the current block within the corresponding position picture of the current picture, or the coordinate of a sample at a position corresponding to the coordinate of the center position sample of the current block within the corresponding position picture.

[0009] In the image decoding method and apparatus according to the present disclosure, the position derived based on the storage unit may be a coordinate obtained by shifting the reference position coordinate of the corresponding position block by the value of the storage unit variable.

[0010] In the image decoding method and apparatus according to the present disclosure, the storage unit variable may be determined based on the value of a condition fulfillment flag indicating whether the specific condition is satisfied, and when the condition fulfillment flag is 1, the storage unit variable may be determined as 2, and when the condition fulfillment flag is 0, the storage unit variable may be determined as 3.

[0011] In the image decoding method and apparatus according to the present disclosure, when the number of samples of the current picture is less than 2,073,600, the condition fulfillment flag may be set to 1, and when the number of samples of the current picture is greater than or equal to 2,073,600, the condition fulfillment flag may be set to 0.

[0012] In the image decoding method and apparatus according to the present disclosure, the width of the storage unit may be determined based on whether a specific width condition is satisfied, and the height of the storage unit may be determined based on whether a specific height condition is satisfied, the specific width condition may be a result of comparing the width of the current picture with a predetermined width, and the specific height condition may be a result of comparing the height of the current picture with a predetermined height.

[0013] In the image decoding method and apparatus according to the present disclosure, the position derived based on the storage unit may be a coordinate derived by performing a shift operation on the reference position coordinate of the corresponding position block based on at least one of a storage unit width variable or a storage unit height variable, and the reference position coordinate of the corresponding position block may be any one of the coordinate of a sample at a position corresponding to the coordinate of the lower-right position sample of the current block within the corresponding position picture of the current picture, or the coordinate of a sample at a position corresponding to the coordinate of the center position sample of the current block within the corresponding position picture.

[0014] In the image decoding method and apparatus according to the present disclosure, the x-coordinate of a position derived based on the storage unit may be a coordinate obtained by shifting the x-coordinate of the reference position of the corresponding position block by the value of the storage unit width variable, and the y-coordinate of a position derived based on the storage unit may be a coordinate obtained by shifting the y-coordinate of the reference position of the corresponding position block by the value of the storage unit height variable.

[0015] In the image decoding method and apparatus according to the present disclosure, the storage unit width variable may be determined based on the value of a width condition satisfying flag indicating whether the specific width condition is satisfied, and the storage unit height variable may be determined based on the value of a height condition satisfying flag indicating whether the specific height condition is satisfied.

[0016] In the image decoding method and apparatus according to the present disclosure, when the width of the current picture is less than 1920, the width condition fulfillment flag may be set to 1; when the width of the current picture is greater than or equal to 1920, the width condition fulfillment flag may be set to 0; when the height of the current picture is less than 1080, the height condition fulfillment flag may be set to 1; and when the height of the current picture is greater than or equal to 1080, the height condition fulfillment flag may be set to 0.

[0017] The image encoding method and apparatus according to the present disclosure can determine a corresponding position block of a current block, acquire motion information of the corresponding position block, derive a time motion vector predictor of the current block based on the motion information of the corresponding position block, and derive a prediction sample of the current block based on the time motion vector predictor of the current block. Here, the motion information of the corresponding position block can be acquired based on a position derived based on a storage unit.

[0018] A computer-readable digital storage medium is provided that stores encoded video / image information that causes an image decoding method to be performed by a decoding device according to the present disclosure.

[0019] A computer-readable digital storage medium is provided that stores video / image information generated according to the image encoding method according to the present disclosure.

[0020] A method and apparatus for transmitting video / image information generated according to the image encoding method according to the present disclosure are provided.

[0021] According to the present disclosure, by deriving movement information of the current block based on a storage unit derived according to specific conditions, a more accurate prediction can be made, and thereby the coding efficiency of inter-prediction can be improved.

[0022] FIG. 1 illustrates a video / image coding system according to the present disclosure.

[0023] FIG. 2 shows a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and to which encoding of a video / image signal is performed.

[0024] FIG. 3 shows a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and to which decoding of a video / image signal is performed.

[0025] FIG. 4 illustrates a decoding method performed by a decoding device (300) as an embodiment according to the present disclosure.

[0026] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs a decoding method according to the present disclosure.

[0027] FIG. 6 illustrates an encoding method performed by an encoding device (200) as an embodiment according to the present disclosure.

[0028] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs an encoding method according to the present disclosure.

[0029] FIG. 8 shows an example of a content streaming system to which embodiments of the present disclosure can be applied.

[0030] The present disclosure is susceptible to various modifications and may have various embodiments; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the present disclosure. Similar reference numerals have been used for similar components in the description of each drawing.

[0031] Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present disclosure, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.

[0032] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.

[0033] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as “comprising” or “having” are intended to specify the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0034] The present disclosure relates to video / video coding. For example, the methods / embodiments disclosed herein may be applied to methods disclosed in the VVC (versatile video coding) standard. Additionally, the methods / embodiments disclosed herein may be applied to methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVN2 (2nd generation of audio video coding standard), or next-generation video / video coding standards (e.g., H.267 or H.268).

[0035] This specification presents various embodiments regarding video / image coding, and unless otherwise noted, said embodiments may be performed in combination with one another.

[0036] In this specification, "video" may refer to a set of images over time. "Picture" generally refers to a unit representing a single image of a specific time period, and "slice" or "tile" is a unit that constitutes a part of a picture in coding. A slice or tile may contain one or more coding tree units (CTUs). A picture may consist of one or more slices or tiles. A tile is a rectangular area composed of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of ​​CTUs having a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of ​​CTUs having a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged continuously according to the CTU raster scan, whereas tiles within a picture may be arranged continuously according to the tile's raster scan. A single slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that can be exclusively contained in a single NAL unit. Meanwhile, a single picture may be divided into two or more subpictures. A subpicture may be a rectangular area of ​​one or more slices within a picture.

[0037] A pixel, or pel, can refer to the smallest unit that constitutes a picture (or image). Additionally, the term 'sample' may be used as a counterpart to pixel. A sample generally represents a pixel or its value, and it may represent only the pixel / pixel value of the luminance (luma) component or only the pixel / pixel value of the chroma component.

[0038] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of ​​a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.

[0039] In this specification, "A or B" may mean "only A," "only B," or "both A and B." Alternatively, in this specification, "A or B" may be interpreted as "A and / or B." For example, in this specification, "A, B or C" may mean "only A," "only B," "only C," or "any combination of A, B and C."

[0040] A slash ( / ) or a comma used in this specification may mean "and / or." For example, "A / B" may mean "A and / or B." Accordingly, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B or C."

[0041] In this specification, "at least one of A and B" may mean "only A," "only B," or "both A and B." Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted as synonymous with "at least one of A and B."

[0042] Additionally, in this specification, "at least one of A, B and C" may mean "only A," "only B," "only C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."

[0043] Additionally, parentheses used in this specification may mean "for example." Specifically, where indicated as "prediction (intra-prediction)," "intra-prediction" may be proposed as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be proposed as an example of "prediction." Furthermore, even where indicated as "prediction (i.e., intra-prediction)," "intra-prediction" may be proposed as an example of "prediction."

[0044] Technical features described individually within a single drawing in this specification may be implemented individually or simultaneously.

[0045] FIG. 1 illustrates a video / image coding system according to the present disclosure.

[0046] Referring to FIG. 1, the video / image coding system may include a first device (source device) (10) and a second device (receiving device) (20).

[0047] A source device (10) can transmit encoded video / image information or data in the form of a file or streaming to a receiving device (20) via a digital storage medium or network. The source device (10) may include a video source (11), an encoding device (12), and a transmission unit (13). The receiving device (20) may include a receiving unit (21), a decoding device (22), and a renderer (23). The encoding device (12) may be called a video / image encoding device, and the decoding device (22) may be called a video / image decoding device. A transmitter may be included in the encoding device. A receiver may be included in the decoding device. The renderer (23) may include a display unit, and the display unit may be composed of a separate device or an external component.

[0048] A video source (11) can acquire video / image through a process of capturing, synthesizing, or generating video / image. The video source (11) may include a video / image capture device and / or a video / image generation device. The video / image capture device may include one or more cameras, a video / image archive containing previously captured video / image, etc. The video / image generation device may include a computer, a tablet, a smartphone, etc., and can generate video / image (electronically). For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.

[0049] The encoding device (12) can encode an input video / image. The encoding device (12) can perform a series of procedures such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0050] The transmission unit (13) can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit (21) of the receiving device (20) via a digital storage medium or network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (13) may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit (21) can receive / extract the bitstream and transmit it to a decoding device (22).

[0051] The decoding device (22) can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device (12).

[0052] The renderer (23) can render the decoded video / image. The rendered video / image can be displayed through the display unit.

[0053] FIG. 2 shows a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and to which encoding of a video / image signal is performed.

[0054] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-predictor (221) and an intra-predictor (222). The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit (250) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to the embodiment. Additionally, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.

[0055] The image segmentation unit (210) can divide an input image (or picture, frame) input to an encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure.

[0056] For example, a single coding unit may be divided into multiple coding units with a deeper depth based on a quad tree structure, a binary tree structure, and / or a terrestrial structure. In this case, for example, the quad tree structure may be applied first and the binary tree structure and / or terrestrial structure may be applied later. Alternatively, the binary tree structure may be applied before the quad tree structure. A coding procedure according to the present specification may be performed based on a final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units of a lower depth so that a coding unit of the optimal size may be used as the final coding unit. Here, the term "coding procedure" may include procedures such as prediction, transformation, and restoration described below.

[0057] As another example, the processing unit may further include a Prediction Unit (PU) or a Transform Unit (TU). In this case, the Prediction Unit and the Transform Unit may each be divided or partitioned from the aforementioned final coding unit. The Prediction Unit may be a unit for sample prediction, and the Transform Unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from transformation coefficients.

[0058] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. A sample may be used to refer to a single picture (or image) as a term corresponding to a pixel or pel.

[0059] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231).

[0060] The prediction unit (220) performs a prediction for a block to be processed (hereinafter referred to as the current block) and can generate a predicted block containing prediction samples for the current block. The prediction unit (220) can determine whether intra prediction is applied or inter prediction is applied at the current block or CU level. The prediction unit (220) can generate various information regarding the prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit it to the entropy encoding unit (240). The information regarding the prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.

[0061] The intra prediction unit (222) can predict the current block by referencing samples within the current picture. The referenced samples may be located near the current block or at a certain distance from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one DC mode or a planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of fineness of the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional modes may be used. The intra prediction unit (222) may determine the prediction mode applied to the current block by using the prediction mode applied to the template area.

[0062] The inter prediction unit (221) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between the template area and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the template area may include a spatial template area (spatial neighboring block) existing within the current picture and a temporal template area (temporal neighboring block) existing in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal template area may be the same or different. The above temporal template area may be referred to by names such as collocated reference block, collocated CU (colCU), etc., and the reference picture containing the above temporal template area may be referred to as a collocated picture (colPic). For example, the inter prediction unit (221) may construct a list of motion information candidates based on the template areas and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of the template area as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the template area is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0063] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for prediction of a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called a combined inter and intra prediction (CIIP) mode. Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in screen content coding (SCC) for games. IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this specification. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values ​​within the picture can be signaled based on information regarding the palette table and palette index. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or to generate a residual signal.

[0064] The transformation unit (232) can generate transform coefficients by applying a transformation technique to a residual signal. For example, the transformation technique may include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on a prediction signal generated using all previously restored pixels. Additionally, the transformation process may be applied to a pixel block of the same size in a square, or to a block of variable size that is not square.

[0065] The quantization unit (233) quantizes the transformation coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (233) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients.

[0066] The entropy encoding unit (240) can perform various encoding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (240) may encode information required for video / image restoration (e.g., values ​​of syntax elements, etc.) together or separately, in addition to quantized transform coefficients.

[0067] Encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream at the level of a Network Abstraction Layer (NAL) unit. The video / image information may further include information regarding various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Additionally, the video / image information may further include general constraint information. In this specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream. The bitstream may be transmitted over a network or stored on a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) that transmits the signal output from the entropy encoding unit (240) and / or a storage unit (not shown) that stores it may be configured as internal / external elements of the encoding device (200), or the transmission unit may be included in the entropy encoding unit (240).

[0068] Quantized transformation coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transformation coefficients through the inverse quantization unit (234) and the inverse transformation unit (235). An adder (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter-prediction unit (221) or the intra-prediction unit (222). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (250) may be called a reconstruction unit or a reconstruction block generation unit. The generated restoration signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.

[0069] The filtering unit (260) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (260) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (270), specifically in the DPB of memory (270). The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0070] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter-prediction unit (221). Through this, when inter-prediction is applied, the encoding device can avoid prediction mismatches between the encoding device (200) and the decoding device, and can also improve encoding efficiency.

[0071] The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter-prediction unit (221). The memory (270) can store motion information of blocks from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (221) to be used as motion information in a spatial template area or motion information in a temporal template area. The memory (270) can store restoration samples of blocks restored within the current picture and transmit them to the intra-prediction unit (222).

[0072] FIG. 3 shows a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and to which decoding of a video / image signal is performed.

[0073] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-predictor (332) and an intra-predictor (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321).

[0074] The aforementioned entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) may be configured by a single hardware component (e.g., a decoding device chipset or processor) according to an embodiment. Additionally, the memory (360) may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.

[0075] When a bitstream containing video / image information is input, the decoding device (300) can restore the image in correspondence with the process in which the video / image information is processed by the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied by the encoding device. Accordingly, the processing unit for decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a binary tree structure. One or more conversion units may be derived from the coding unit. And, the restored image signal decoded and output through the decoding device (300) can be played back through a playback device.

[0076] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through an entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The decoding device can decode the picture based on information regarding the parameter sets and / or the general constraint information. The signaling / receiving information and / or syntax elements described below in this specification may be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information within the bitstream based on coding methods such as exponential chord coding, CAVLC, or CABAC, and output the values ​​of syntax elements required for image restoration and the quantized values ​​of transformation coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, determine a context model using information on the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information on symbols / bins decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Among the information decoded in the entropy decoding unit (310), information regarding prediction is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and the residual value for which entropy decoding was performed in the entropy decoding unit (310), i.e., quantized transformation coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive residual signals (residual blocks, residual samples, residual sample array). Additionally, among the information decoded in the entropy decoding unit (310), information regarding filtering can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310).

[0077] Meanwhile, the decoding device according to the present specification may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoding device (video / image / picture information decoding device) and a sample decoding device (video / image / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), inverse transform unit (322), adder (340), filtering unit (350), memory (360), inter prediction unit (332), and intra prediction unit (331).

[0078] In the inverse quantization unit (321), the quantized transformation coefficients can be inversely quantized to output transformation coefficients. The inverse quantization unit (321) can rearrange the quantized transformation coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain transformation coefficients.

[0079] In the inverse conversion unit (322), the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0080] The prediction unit (320) can perform a prediction for the current block and generate a predicted block containing prediction samples for the current block. The prediction unit (320) can determine whether an intra prediction or an inter prediction is applied to the current block based on information regarding the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter prediction mode.

[0081] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) may apply intra prediction or inter prediction for prediction of a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called a combined inter and intra prediction (CIIP) mode. Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, such as SCC (screen content coding). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this specification. The palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information regarding the palette table and palette index can be included in the above video / image information and signaled.

[0082] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or at a certain distance from the current block depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and a plurality of directional modes. The intra prediction unit (331) may determine the prediction mode applied to the current block by using the prediction mode applied to the template area.

[0083] The inter prediction unit (332) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between a template area and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the template area may include a spatial template area (spatial neighboring block) existing within the current picture and a temporal template area (temporal neighboring block) existing in the reference picture. For example, the inter prediction unit (332) may construct a motion information candidate list based on the template areas and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-prediction can be performed based on various prediction modes, and information regarding the prediction may include information indicating the inter-prediction mode for the current block.

[0084] The adder (340) can generate a restoration signal (restoration picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter prediction unit (332) and / or the intra prediction unit (331)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the prediction block can be used as the restoration block.

[0085] The addition unit (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra-predicting the next block to be processed within the current picture, may be output after filtering as described below, or may be used for inter-predicting the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.

[0086] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can generate a modified restored picture by applying various filtering methods to the restored picture, and can transmit the modified restored picture to memory (360), specifically to the DPB of memory (360). The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0087] The (modified) restored picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter prediction unit (332) to be used as motion information in a spatial template area or motion information in a temporal template area. The memory (360) can store restoration samples of blocks restored within the current picture and transmit them to the intra prediction unit (331).

[0088] In this specification, the embodiments described in the filtering unit (260), inter prediction unit (221), and intra prediction unit (222) of the encoding device (200) may be applied to the filtering unit (350), inter prediction unit (332), and intra prediction unit (331) of the decoding device (300) in the same or corresponding manner.

[0089] The temporal motion vector predictor (TMVP) of the current block can be derived based on motion information of the corresponding collocated block (hereinafter, collocated block) of the current block. The collocated block may be included in the corresponding collocated picture (hereinafter, collocated picture) of the current picture containing the current block. The collocated picture may be a picture corresponding to the basic motion vector of the current block, or a picture determined based on signaled information. The collocated block may be determined based on the basic motion vector of the current block. The basic motion vector of the current block may be derived based on at least one of the motion vector of the surrounding reference block of the current block or pre-defined offset information.

[0090] The decoding method according to the present disclosure can generate a restored picture based on a restored sample. The restored picture can be stored in a Decoded Picture Buffer (DPB) of a decoding device (300) and used to restore a picture to be decoded thereafter. At this time, at least one of the sample value of the restored picture or the motion information of the prediction block within the restored picture can be stored in the DPB. The motion information of the prediction block can be stored in sample units or prediction block units. Alternatively, the motion information of the prediction block can be stored in predefined units. By storing the motion information of the prediction block in predefined units, the usage of the DPB can be reduced. For example, even if motion vectors are derived in 4x4 prediction block units, the motion information of the restored picture can be integrated and stored in predefined 8x8 units.

[0091] The time motion vector predictor of the current block can be derived based on motion information of the reconstructed picture stored in the DPB. The accuracy of the time motion vector predictor can be determined by the storage unit of the motion information stored in the DPB. Additionally, if the size of the input image changes, the encoding efficiency or decoding efficiency may change depending on the motion information storage unit of the DPB. For example, when acquiring motion information of a corresponding location block (hereinafter, call block) to derive the time motion vector predictor of the current block, the acquired motion information may differ depending on the motion information storage unit. A specific method for deriving the time motion vector predictor of the current block according to the motion information storage unit will be examined in detail with reference to FIG. 4.

[0092] FIG. 4 illustrates an image decoding method performed by a decoding device (300) as an embodiment according to the present disclosure.

[0093] Referring to FIG. 4, the corresponding position block of the current block can be determined (S400).

[0094] A block corresponding to the current block's position (hereinafter referred to as the call block) can be determined based on at least one of the basic motion vector of the current block or the corresponding position picture of the current picture (hereinafter referred to as the call picture). More specifically, a block containing a sample corresponding to coordinates derived based on the basic motion vector within the call picture can be determined as the call block. For example, the call block may be a block containing a sample of coordinates corresponding to the coordinates of the current block within the call picture.

[0095] Referring to FIG. 4, movement information of the corresponding position block can be obtained (S410).

[0096] The time motion vector predictor of the current block can be derived based on the corresponding location block of the current block (hereinafter, call block). More specifically, the time motion vector predictor of the current block can be derived based on the motion information of the call block. In this case, the motion information of the call block can be obtained based on a location derived based on a storage unit.

[0097] A location derived based on a storage unit (hereinafter referred to as a storage location) may be derived based on whether specific conditions are satisfied. For example, if a first condition is satisfied, the storage location may be derived according to a first storage unit, and if the first condition is not satisfied, it may be derived according to whether a second condition is satisfied. Additionally, if the second condition is satisfied, the storage location may be derived according to a second storage unit, and if the second condition is not satisfied, it may be derived according to whether a third condition is satisfied. In this manner, the storage location may be derived according to a hierarchical structure in which multiple conditions are applied sequentially, and if each condition is not satisfied, a corresponding storage unit may be selected based on whether the next condition is satisfied. This process may be performed repeatedly until a predetermined final condition is reached.

[0098] As another example, the storage location may be derived according to a first storage unit if a predetermined condition is satisfied, and may be derived according to a second storage unit if the predetermined condition is not satisfied. As an example, the storage location may be derived based on a comparison of the current picture's resolution (size) and a predetermined resolution. For example, if the current picture's resolution is smaller than the predetermined resolution, the storage location may be derived according to a first storage unit, and if the current picture's resolution is greater than or equal to the predetermined resolution, the storage location may be derived according to a second storage unit. The storage location may be derived based on a comparison of at least one of the width or height, as well as a comparison of the total resolution of the current picture's resolution and the predetermined resolution. Here, the first storage unit may be a 4x4 unit, the second storage unit may be an 8x8 unit, and the predetermined resolution may be a 1920x1080 (FHD) resolution. Alternatively, the first storage unit may be an 8x8 unit, the second storage unit may be a 16x16 unit, and the predetermined resolution may be a 4K resolution. However, the present disclosure is not limited thereto, and the predetermined resolution may be another specific resolution, and the first storage unit and the second storage unit may each be any one of 1x1 to 16x16 units.

[0099] As another example, the width and height of a storage unit for deriving a storage location can be derived based on a comparison of the current picture width with a predetermined width value and a comparison of the current picture height with a predetermined height value, respectively. In this case, the storage unit may be in the shape of a rectangle. For example, if the width of the current picture is smaller than a predetermined width, the width of the storage unit may be determined as a first value, and if the width of the current picture is greater than or equal to a predetermined width, the width of the storage unit may be determined as a second value. Here, the predetermined width may be 1920, the first value may be 4, and the second value may be 8. Additionally, if the height of the current picture is smaller than a predetermined height, the height of the storage unit may be determined as a third value, and if the height of the current picture is greater than or equal to a predetermined height, the height of the storage unit may be determined as a fourth value. Here, the predetermined height may be 1080, the third value may be 4, and the fourth value may be 8. The first value may be set to be the same as the third value or different, and the second value may be set to be the same as the fourth value or different.

[0100] As another example, the shape of the storage unit to determine the storage location can be determined based on a comparison between the ratio of the width and height of the current picture and a predetermined value. For example, if the value obtained by dividing the width of the current picture by the height is greater than 16 / 9, the storage unit may be in the shape of a rectangle. Here, the storage unit may be an 8x4 unit.

[0101] As another example, the storage location may be derived based on a comparison between the number of samples in the current picture and a predetermined number of samples. For example, if the number of samples in the current picture is less than the predetermined number of samples, the storage location may be derived according to a first storage unit, and if the number of samples in the current picture is greater than or equal to the predetermined number of samples, the storage location may be derived according to a second storage unit. Here, the first storage unit may be a 4x4 unit, the second storage unit may be an 8x8 unit, and the predetermined number of samples may be 2,073,600 (1920x1080). However, the present disclosure is not limited thereto, and the predetermined number of samples may be a different number, and the first storage unit and the second storage unit may each be any one of 1x1 to 16x16 units.

[0102] Movement information of a call block can be obtained based on a location derived based on a storage unit. A storage unit block containing the coordinates of a location derived according to the storage unit may be the location where the movement information of the call block is obtained.

[0103] For example, the location (xColCb, yColCb) derived based on the storage unit can be derived as shown in Equation 1.

[0104]

[0105] Here, (x, y) may be reference position coordinates. Additionally, the reference position coordinates may be coordinates corresponding to the coordinates of the current block within the call picture. mvUnit may be a variable having a value corresponding to a storage unit. For example, if the storage units are 1x1, 2x2, 4x4, 8x8, or 16x16 units, respectively, mvUnit may be 0, 1, 2, 3, or 4, respectively. The mvUnit may be referred to as a storage unit variable.

[0106] The above storage unit may be determined based on whether specific conditions are met. For example, the above specific condition may be a result of comparing the number of samples in the current picture with a predetermined number. For example, if the number of samples in the current picture is less than the predetermined number, the storage unit may be determined to be a 4x4 unit, and if the number of samples in the current picture is greater than or equal to the predetermined number, the storage unit may be determined to be an 8x8 unit. The predetermined number may be 2,073,600 (1920x1080).

[0107] Whether the above specific condition is satisfied can be defined by a condition satisfaction flag (mvCompFlag) indicating whether a specific condition for determining a storage unit is satisfied. For example, if mvCompFlag is 1 (when the specific condition is satisfied), the storage unit can be determined as a 4x4 unit, and accordingly, mvUnit can be 2. Also, if mvCompFlag is 0 (when the specific condition is not satisfied), the storage unit can be determined as an 8x8 unit, and accordingly, mvUnit can be 3. The mvUnit derived according to the value of mvCompFlag can derive the location where movement information of the call block is obtained.

[0108] Hereinafter, with reference to Tables 1 to 4, we will specifically examine a method for deriving the location where movement information of a call block is obtained based on storage unit variables. In Tables 1 to 4, the location of at least one of the derivation processes of mvCompFlag or mvUnit may be changed.

[0109] Table 1 is an example of a method for deriving the time motion vector predictor of the current block.

[0110] 8.5.2.11 Derivation process for temporal luma motion vector predictionInputs to this process are:- a luma location ( xCb, yCb ) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture,- a variable cbWidth specifying the width of the current coding block in luma samples,- a variable cbHeight specifying the height of the current coding block in luma samples,- a reference index refIdxLX, with X being 0 or 1.Outputs of this process are:- the motion vector prediction mvLXCol in 1 / 16 fractional-sample accuracy,- the availability flag availableFlagLXCol.The variable currCb specifies the current luma coding block at luma location ( xCb, yCb ).The variable mvCompFlag is equal to 0.When the following conditions is true, mvCompFlag is set equal to 1.- The number of luma samples of the current picture is less than 1920*1080.The motion compression unit is derived as follows.mvUnit= mvCompFlag ? 2 : 3The variables mvLXCol and availableFlagLXCol are derived as follows:- If ph_temporal_mvp_enabled_flag is equal to 0 or ( cbWidth * cbHeight ) is less than or equal to 32 when mvUnit is greater than or equal to 3, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.- Otherwise (ph_temporal_mvp_enabled_flag is equal to 1), the following ordered steps apply:1.The bottom-right collocated motion vector and the bottom and right boundary sample locations are derived as follows:xColBr = xCb + cbWidth (583)yColBr = yCb + cbHeight (584)rightBoundaryPos = sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] ?SubpicRightBoundaryPos : pps_pic_width_in_luma_samples - 1 (585)botBoundaryPos = sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] ?SubpicBotBoundaryPos : pps_pic_height_in_luma_samples - 1 (586)- If yCb >> CtbLog2SizeY is equal to yColBr >> CtbLog2SizeY, yColBr is less than or equal to botBoundaryPos and xColBr is less than or equal to rightBoundaryPos, the following applies:- The luma location ( xColCb, yColCb ) is set equal to ( ( xColBr >> mvUnit ) << mvUnit, ( yColBr >> mvUnit ) << mvUnit ).- The variable colCb specifies the luma coding block covering the location ( xColCb, yColCb ) inside the collocated picture specified by ColPic.- The derivation process for collocated motion vectors as specified in clause 8.5.2.12 is invoked with currCb, colCb, ( xColCb, yColCb ), refIdxLX and sbFlag set equal to 0 as inputs, and the output is assigned to mvLXCol and availableFlagLXCol.- Otherwise, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.2. When availableFlagLXCol is equal to 0, the central collocated motion vector is derived as follows:xColCtr = xCb + ( cbWidth >> 1 ) (587)yColCtr = yCb + ( cbHeight >> 1 ) (588)- The luma location ( xColCb, yColCb ) is set equal ( ( xColCtr >> mvUnit ) << mvUnit, ( yColCtr >> mvUnit ) << mvUnit ).- The variable colCb specifies the luma coding block covering the location ( xColCb, yColCb ) inside the collocated picture specified by ColPic.The derivation process for collocated motion vectors as specified in clause 8.5.2.12 is invoked with currCb, colCb, ( xColCb, yColCb ), refIdxLX and sbFlag set equal to 0 as inputs,and the output is assigned to (...).

[0111] Referring to Table 1, the location (xColCb, yColCb) (hereinafter referred to as the storage location) derived based on the storage unit can be derived according to the storage unit variable (mvUnit). Movement information of the call block can be obtained based on the storage unit block containing the storage location (xColCb, yColCb). The movement information may include a movement vector. More specifically, (xColCb, yColCb) may be a location obtained by shifting the reference position coordinates of the call block according to the storage unit corresponding to mvUnit. For example, (xColCb, yColCb) may be a location obtained by shifting the reference position coordinates of the call block by mvUnit. xColCb may be the x-coordinate of the reference position of the call block shifted to the right by the value of mvUnit and then shifted to the left by the value of mvUnit, and yColCb may be the y-coordinate of the reference position of the call block shifted to the right by the value of mvUnit and then shifted to the left by the value of mvUnit. The reference position coordinates of the call block may be either the coordinates of a sample at a position corresponding to the coordinates of the sample at the bottom-right position of the current block within the call picture (xColBr, yColBr) or the coordinates of a sample at a position corresponding to the coordinates of the sample at the center position of the current block within the call picture (xColCtr, yColCtr).

[0112] mvUnit can be determined based on the value of a condition fulfillment flag (mvCompFlag) indicating whether a specific condition for determining the storage unit is met. Referring to Table 1, mvUnit can be 2 when the value of mvCompFlag is 1, and 3 when the value of mvCompFlag is 0. mvUnit can have a value corresponding to the storage unit. For example, when mvUnit is 0, 1, 2, 3, or 4, the storage units can be 1x1, 2x2, 4x4, 8x8, or 16x16 units, respectively.

[0113] mvCompFlag may indicate whether specific conditions for determining the storage unit are met. Referring to Table 1, mvCompFlag may be set to 1 if the number of luminance samples in the current picture is less than 2,073,600 (1920x1080), and may be set to 0 if the number of luminance samples in the current picture is greater than or equal to 2,073,600.

[0114] A time motion vector predictor for the current block may not be induced when the time motion vector predictor enable flag (ph_temporal_mvp_enabled_flag) is 0, or when mvUnit has a value greater than or equal to 3 when the product of the width and height of the current block is less than or equal to 32. More specifically, when the time motion vector predictor enable flag (ph_temporal_mvp_enabled_flag) is 0, or when mvUnit has a value greater than or equal to 3 when the product of the width and height of the current block is less than or equal to 32, the time motion vector predictor (mvLXcol) is set to 0, and the time motion vector predictor use flag (availableFlagLXCol), which indicates whether a time motion vector predictor is available for the current block, may be set to 0.

[0115] Table 2 is an example of a method to derive subblock-based time merge candidates.

[0116] 8.5.5.3 Derivation process for subblock-based temporal merging candidatesInputs to this process are:- a luma location ( xCb, yCb ) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture,- a variable cbWidth specifying the width of the current coding block in luma samples,- a variable cbHeight specifying the height of the current coding block in luma samples.- the availability flag availableFlagA1of the neighbouring coding unit,- the reference indices refIdxLXA1of the neighbouring coding unit with X = 0..1,- the prediction list utilization flags predFlagLXA1of the neighbouring coding unit with X = 0..1,- the motion vector in 1 / 16 fractional-sample accuracy mvLXA1of the neighbouring coding unit with X = 0..1.Outputs of this process are:- the availability flag availableFlagSbCol,- the number of luma subblocks in horizontal direction numSbX and in vertical direction numSbY,- the reference indices refIdxL0SbCol and refIdxL1SbCol,- the luma motion vectors in 1 / 16 fractional-sample accuracy mvL0SbCol[ xSbIdx ][ ySbIdx ] and mvL1SbCol[ xSbIdx ][ ySbIdx ] with xSbIdx = 0..numSbX - 1, ySbIdx = 0..numSbY - 1,- the prediction list utilization flags predFlagL0SbCol[ xSbIdx ][ ySbIdx ] and predFlagL1SbCol[ xSbIdx ][ ySbIdx ] with xSbIdx = 0..numSbX - 1, ySbIdx = 0..numSbY - 1.The variable mvCompFlag is equal to 0.When the following conditions is true, mvCompFlag is set equal to 1.- The number of luma samples of the current picture is less than 1920*1080.The motion compression unit is derived as follows.mvUnit= mvCompFlag ? 2 : 3The availability flag availableFlagSbCol is derived as follows.- If one or more of the following conditions are true, availableFlagSbCol is set equal to 0.- ph_temporal_mvp_enabled_flag is equal to 0.- sps_sbtmvp_enabled_flag is equal to 0.- cbWidth is less than 8 when mvUnit is equal to 3.- cbHeight is less than 8 when mvUnit is equal to 3.- Otherwise, the following ordered steps apply:1. The location ( xCtb, yCtb ) of the top-left sample of the luma coding tree block that contains the current coding block and the location ( xCtrCb, yCtrCb ) of the below-right center sample of the current luma coding block are derived as follows:xCtb = ( xCb >> CtbLog2SizeY ) << CtbLog2SizeY (702)yCtb = ( yCb >> CtbLog2SizeY ) << CtbLog2SizeY (703)xCtrCb = xCb + ( cbWidth / 2 ) (704)yCtrCb = yCb + ( cbHeight / 2 ) (705)2. The derivation process for subblock-based temporal merging base motion data as specified in clause 8.5.5.4 is invoked with the location ( xCtb, yCtb ), the location ( xCtrCb, yCtrCb ), the availability flag availableFlagA1, and the prediction list utilization flag predFlagLXA1, and the reference index refIdxLXA1, and the motion vector mvLXA1, with X = 0..1 as inputs and the motion vectors ctrMvLX, and the prediction list utilization flags ctrPredFlagLX of the collocated block, with X = 0..1, and the temporal motion vector tempMv as outputs.3. The variable availableFlagSbCol is derived as follows:- If both ctrPredFlagL0 and ctrPredFlagL1 are equal to 0, availableFlagSbCol is set equal to 0.- Otherwise, availableFlagSbCol is set equal to 1.When availableFlagSbCol is equal to 1, the following applies:- The variables numSbX, numSbY, sbWidth, sbHeight and refIdxLXSbCol are derived as follows:numSbX = cbWidth >> 2 (706)numSbY = cbHeight >> 2 (707)sbWidth = cbWidth / numSbX (708)sbHeight = cbHeight / numSbY (709)refIdxLXSbCol = 0 (710)- For xSbIdx = 0..numSbX - 1 and ySbIdx = 0..numSbY - 1, the motion vectors mvLXSbCol[ xSbIdx ][ ySbIdx ] and prediction list utilization flags predFlagLXSbCol[ xSbIdx ][ ySbIdx ] are derived as follows:- The luma location ( xSb, ySb ) specifying the below-right center sample of the current subblock relative to the top-left luma sample of the current picture is derived as follows:xSb = xCb + xSbIdx * sbWidth + sbWidth / 2 (711)ySb = yCb + ySbIdx * sbHeight + sbHeight / 2 (712)- The location ( xColSb, yColSb ) of the collocated subblock inside ColPic is derived as follows.- The following applies:yColSb = Clip3( yCtb, Min( pps_pic_height_in_luma_samples - 1, yCtb + ( 1 << CtbLog2SizeY ) - 1 ), (713)ySb + tempMv[ 1 ] )- If sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] is equal to 1, the following applies:xColSb = Clip3( xCtb, Min( SubpicRightBoundaryPos, xCtb + ( 1 << CtbLog2SizeY ) + 3 ), (714)xSb + tempMv[ 0 ] )- Otherwise ( sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] is equal to 0), the following applies:xColSb = Clip3( xCtb,Min( pps_pic_width_in_luma_samples - 1, xCtb + ( 1 << CtbLog2SizeY ) + 3 ), (715)xSb + tempMv[ 0 ] )- The variable currCb specifies the luma coding block covering the current subblock inside the current picture.- The luma location ( xColCb, yColCb ) is set equal to ( ( xColSb >> mvUnit ) << mvUnit, ( yColSb >> mvUnit ) << mvUnit ).- The variable colCb specifies the luma coding block covering the location ( xColCb, yColCb ) inside the collocated picture specified by ColPic.- The prediction list utilization flags predFlagL0SbCol[ xSbIdx ][ ySbIdx ] and predFlagL1SbCol[ xSbIdx ][ ySbIdx ] are initialized to be equal to FALSE.- The derivation process for collocated motion vectors as specified in clause 8.5.2.12 is invoked with currCb, colCb, ( xColCb, yColCb ), refIdxL0 set equal to 0 and sbFlag set equal to 1 as inputs and the output being assigned to the motion vector of the subblock mvL0SbCol[ xSbIdx ][ ySbIdx ] and predFlagL0SbCol[ xSbIdx ][ ySbIdx ].- When sh_slice_type is equal to B, the derivation process for collocated motion vectors as specified in clause 8.5.2.12 is invoked with currCb, colCb, ( xColCb, yColCb ), refIdxL1 set equal to 0 and sbFlag set equal to 1 as inputs and the output being assigned to the motion vector of the subblock mvL1SbCol[ xSbIdx ][ ySbIdx ] and predFlagL1SbCol[ xSbIdx ][ ySbIdx ].- When predFlagL0SbCol[ xSbIdx ][ ySbIdx ] and predFlagL1SbCol[ xSbIdx ][ ySbIdx ] are both equal to 0, the following applies for X = 0..1:mvLXSbCol[ xSbIdx ][ ySbIdx ] = ctrMvLX (716)predFlagL

[0117] Referring to Table 2, the location (xColCb, yColCb) (hereinafter referred to as the storage location) derived based on the storage unit can be derived according to the storage unit variable (mvUnit). Motion information corresponding to the storage unit block containing the storage location (xColCb, yColCb) can be used as motion information for the sub-block unit. More specifically, (xColCb, yColCb) may be a location obtained by shifting the reference location coordinates (xColSb, yColSb) of the corresponding location sub-block (hereinafter referred to as the call sub-block) within the call picture (colPic) according to the storage unit corresponding to mvUnit. For example, (xColCb, yColCb) may be a location obtained by shifting (xColSb, yColSb) by mvUnit. xColCb may be the value obtained by right-shifting the reference position x-coordinate (xColSb) of the call subblock by the value of mvUnit and then left-shifting it by the value of mvUnit, and yColCb may be the value obtained by right-shifting the reference position y-coordinate (yColSb) of the call subblock by the value of mvUnit and then left-shifting it by the value of mvUnit.

[0118] Referring to Table 2, subblock-based time merge candidates may not be derived if certain conditions are satisfied. More specifically, if certain conditions are satisfied, a subblock-based time merge candidate usage flag (availableFlagSbCol), which indicates whether a subblock-based time merge candidate is used in the current block, may be set to 0. The said certain conditions may include at least one of the following: when the time motion vector predictor enable flag (ph_temporal_mvp_enabled_flag) is 0, when the subblock-based time merge candidate enable flag (sps_sbtmvp_enabled_flag) is 0, when mvUnit is 3 and the width (cbWidth) of the current block is less than 8, or when mvUnit is 3 and the height (cbHeight) of the current block is less than 8.

[0119] Referring to Table 2, the reference position coordinates (xColSb, yColSb) of a call subblock can be derived based on the reference position coordinates (xSb, ySb) of a subblock. The reference position coordinates of a subblock can be derived based on the number of subblock units (numSbX or numSbY). Here, the size of the subblock unit may be 4x4. In this case, the number of subblocks of the current block may be 1.

[0120] The description of mvUnit and mvCompFlag is as referenced in Table 1.

[0121] Table 3 is an example of a method for deriving basic movement information of subblock-based time merge candidates.

[0122] 8.5.5.4 Derivation process for subblock-based temporal merging base motion dataInputs to this process are:- the location ( xCtb, yCtb ) of the top-left sample of the luma coding tree block that contains the current coding block,- the location ( xCtrCb, yCtrCb ) of the top-left sample of the collocated luma coding block that covers the below-right center sample.- the availability flag availableFlagA1of the neighbouring coding unit,- the reference indices refIdxLXA1of the neighbouring coding unit with X = 0..1,- the prediction list utilization flags predFlagLXA1of the neighbouring coding unit with X = 0..1,- the motion vectors in 1 / 16 fractional-sample accuracy mvLXA1of the neighbouring coding unit with X = 0..1.Outputs of this process are:- the motion vectors ctrMvL0 and ctrMvL1,- the prediction list utilization flags ctrPredFlagL0 and ctrPredFlagL1,- the temporal motion vector tempMv.The variable tempMv is set as follows:tempMv[ 0 ] = 0 (718)tempMv[ 1 ] = 0 (719)The variable currPic specifies the current picture.When availableFlagA1is equal to TRUE, the following applies:- If all of the following conditions are true, tempMv is set equal to mvL0A1:- predFlagL0A1is equal to 1,- DiffPicOrderCnt( ColPic, RefPicList[ 0 ][ refIdxL0A1 ] ) is equal to 0,- Otherwise, if all of the following conditions are true, tempMv is set equal to mvL1A1:- sh_slice_type is equal to B,- predFlagL1A1is equal to 1,- DiffPicOrderCnt( ColPic, RefPicList[ 1 ][ refIdxL1A1 ] ) is equal to 0.The rounding process for motion vectors as specified in clause 8.5.2.14 is invoked with mvX set equal to tempMv, rightShift set equal to 4, and leftShift set equal to 0 as inputs and the rounded tempMv as output.The location ( xColCb, yColCb ) of the collocated block inside ColPic is derived as follows.- The following applies:yColCb = Clip3( yCtb,Min( pps_pic_height_in_luma_samples - 1, yCtb + ( 1 << CtbLog2SizeY ) - 1 ), (720)yCtrCb + tempMv[ 1 ] )- If sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] is equal to 1, the following applies:xColCb = Clip3( xCtb,Min( SubpicRightBoundaryPos, xCtb + ( 1 << CtbLog2SizeY ) + 3 ), (721)xCtrCb + tempMv[ 0 ] )- Otherwise ( sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] is equal to 0), the following applies:xColCb = Clip3( xCtb,Min( pps_pic_width_in_luma_samples - 1, xCtb + ( 1 << CtbLog2SizeY ) + 3 ), (722)xCtrCb + tempMv[ 0 ] )The array colPredMode is set equal to the prediction mode array CuPredMode[ 0 ] of the collocated picture specified by ColPic.The motion vectors ctrMvL0 and ctrMvL1, and the prediction list utilization flags ctrPredFlagL0 and ctrPredFlagL1 are derived as follows:The variable mvCompFlag is equal to 0.When the following conditions is true, mvCompFlag is set equal to 1.- The number of luma samples of the current picture is less than 1920*1080.- The motion compression unit is derived according to mvCompFlag- mvUnit= mvCompFlag ? 2 : 3- If colPredMode[ ( xColCb >> mvUnit ) << mvUnit ][ ( yColCb >> mvUnit ) << mvUnit ] is equal to MODE_INTER, the following applies:- The variable currCb specifies the luma coding block covering ( xCtrCb, yCtrCb ) inside the current picture.- The luma location ( xColCb, yColCb ) is set equal to ( ( xColCb >> mvUnit ) << mvUnit, ( yColCb >> mvUnit ) << mvUnit ).- The variable colCb specifies the luma coding block covering the location ( xColCb, yColCb ) inside the collocated picture specified by ColPic.- The prediction list utilization flags ctrPredFlagL0 and ctrPredFlagL1 are initialized to be equal to FALSE.- The derivation process for collocated motion vectors specified in clause 8.5.2.12 is invoked with currCb, colCb, (xColCb, yColCb), refIdxL0 set equal to 0, and sbFlag set equal to 1 as inputs and the output being assigned to ctrMvL0 and ctrPredFlagL0.- When sh_slice_type is equal to B, the derivation process for collocated motion vectors specified in clause 8.5.2.12 is invoked with currCb, colCb, (xColCb, yColCb), refIdxL1 set equal to 0, and sbFlag set equal to 1 as inputs and the output being assigned to ctrMvL1 and ctrPredFlagL1. - Otherwise, the following applies:ctrPredFlagL0 = 0 (723)ctrPredFlagL1 = 0 (724).

[0123] Referring to Table 3, movement information of the current block can be derived from a call block at a specific location. Here, the specific location can be derived based on a storage unit variable (mvUnit) of the movement information. For example, the specific location (xColCb, yColCb) may be a location shifted by mvUnit from a corresponding location within the call block that corresponds to a location derived based on a sample location (or center sample location) around the bottom right of the current block. xColCb may be a location obtained by shifting the x-coordinate of the corresponding location within the call block to the right by the value of mvUnit and then shifting it to the left by the value of mvUnit, and yColCb may be a location obtained by shifting the y-coordinate of the corresponding location within the call block to the right by the value of mvUnit and then shifting it to the left by the value of mvUnit.

[0124] The description of mvUnit and mvCompFlag is as referenced in Table 1.

[0125] Table 4 is an example of a method for deriving a time motion vector predictor for the lower right control point during affine prediction.

[0126] The fourth (collocated bottom-right) control point motion vector cpMvLXCorner[ 3 ], reference index refIdxLXCorner[ 3 ], prediction list utilization flag predFlagLXCorner[ 3 ] and the availability flag availableFlagCorner[ 3 ] with = 0..1 are derived as follows:- The reference indices for the temporal merging candidate, refIdxLXCorner[ 3 ], with X = 0..1, are set equal to 0.- For X = 0..1, the variables mvLXCol and availableFlagLXCol are derived as follows:- If ph_temporal_mvp_enabled_flag is equal to 0, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.- Otherwise (ph_temporal_mvp_enabled_flag is equal to 1), the following applies:xColBr = xCb + cbWidth (761)yColBr = yCb + cbHeight (762)rightBoundaryPos = sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] ? SubpicRightBoundaryPos :pps_pic_width_in_luma_samples - 1 (763)botBoundaryPos = sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] SubpicBotBoundaryPos :pps_pic_height_in_luma_samples - 1 (764)mvUnit= mvCompFlag 2 : 3- If yCb >> CtbLog2SizeY is equal to yColBr >> CtbLog2SizeY, yColBr is less than or equal to botBoundaryPos and xColBr is less than or equal to rightBoundaryPos, the following applies:- The luma location ( xColCb, yColCb ) is set equal to ( ( xColBr >> mvUnit ) << mvUnit, ( yColBr >> mvUnit ) << mvUnit ).- The variable colCb specifies the luma coding block covering the location ( xColCb, yColCb ) inside the collocated picture specified by ColPic.- The derivation process for collocated motion vectors as specified in clause 8.5.2.12 is invoked with currCb, colCb, ( xColCb, yColCb ), refIdxLXCorner[ 3 ] and sbFlag set equal to 0 as inputs, and the output is assigned to mvLXCol and availableFlagLXCol.- Otherwise, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.- The variables availableFlagCorner[ 3 ], predFlagL0Corner[ 3 ], cpMvL0Corner[ 3 ] and predFlagL1Corner[ 3 ] are derived as follows:availableFlagCorner[ 3 ] = availableFlagL0Col (765)predFlagL0Corner[ 3 ] = availableFlagL0Col (766)cpMvL0Corner[ 3 ] = mvL0Col (767)predFlagL1Corner[ 3 ] = 0 (768)- When sh_slice_type is equal to B, the variables availableFlagCorner[ 3 ], predFlagL1Corner[ 3 ] and cpMvL1Corner[ 3 ] are derived as follows:availableFlagCorner[ 3 ] = availableFlagL0Col | | availableFlagL1Col (769)predFlagL1Corner[ 3 ] = availableFlagL1Col (770)cpMvL1Corner[ 3 ] = mvL1Col (771).

[0127] Referring to Table 4, the location (xColCb, yColCb) (hereinafter referred to as the storage location) derived based on the storage unit can be derived according to the storage unit variable (mvUnit). Movement information of the call block can be obtained based on the storage unit block containing the storage location (xColCb, yColCb). The movement information may include a movement vector. More specifically, (xColCb, yColCb) may be a location obtained by shifting the reference position coordinates of the call block according to the storage unit corresponding to mvUnit. xColCb may be a value obtained by shifting the reference position x-coordinate of the call block to the right by the value of mvUnit and then shifting it to the left by the value of mvUnit, and yColCb may be a value obtained by shifting the reference position y-coordinate of the call block to the right by the value of mvUnit and then shifting it to the left by the value of mvUnit. The reference position coordinates of the above call block may be the coordinates of the sample at the position corresponding to the coordinates of the sample at the bottom right of the current block within the call picture (xColBr, yColBr).

[0128] The description of mvUnit and mvCompFlag is as referenced in Table 1.

[0129] Movement information of a call block can be obtained based on a location derived from a storage unit. A storage unit block containing the coordinates of a location derived according to the storage unit may be the location where the movement information of the call block is obtained. In this case, the width and height of the storage unit may differ from each other.

[0130] For example, a position (xColCb, yColCb) derived according to at least one of the width or height of a storage unit can be derived as shown in Equation 2.

[0131]

[0132] Here, (x, y) may be reference position coordinates. Additionally, the reference position coordinates may be coordinates corresponding to the coordinates of the current block within the call picture. mvUnit_x may be a variable having a value corresponding to the width of the storage unit, and mvUnit_y may be a variable having a value corresponding to the height of the storage unit. For example, if the width or height of the storage unit is 1, 2, 4, 8, or 16, mvUnit_x or mvUnit_y may be 0, 1, 2, 3, or 4, respectively. mvUnit_x may be referred to as the storage unit width variable, mvUnit_y may be referred to as the storage unit height variable, and mvUnit_x and mvUnit_y together may be referred to as the storage unit variable.

[0133] The width of the storage unit may be determined based on whether a specific width condition is satisfied, and the height of the storage unit may be determined based on whether a specific height condition is satisfied. For example, the width of the storage unit may be determined based on whether the width of the current picture is smaller than a predetermined width, and the height of the storage unit may be determined based on whether the height of the current picture is smaller than a predetermined height. For example, the width of the storage unit may be determined as 4 if the width of the current picture is less than 1920, and as 8 if the width of the current picture is greater than or equal to 1920. Additionally, the height of the storage unit may be determined as 4 if the height of the current picture is less than 1080, and as 8 if the height of the current picture is greater than or equal to 1080.

[0134] Whether a specific width condition or a specific height condition is satisfied can be defined by a width condition satisfaction flag (mvCompFlagX) or a height condition satisfaction flag (mvCompFlagY) indicating whether a specific condition for determining the width or height of a storage unit is satisfied. For example, if mvCompFlagX is 1 (when a specific width condition is satisfied), the width of the storage unit may be determined to be 4, and accordingly, mvUnit_x may be 2. Also, if mvCompFlagX is 0 (when a specific width condition is not satisfied), the width of the storage unit may be determined to be 8, and accordingly, mvUnit_x may be 3. Also, if mvCompFlagY is 1 (when a specific height condition is satisfied), the height of the storage unit may be determined to be 4, and accordingly, mvUnit_y may be 2. Additionally, when mvCompFlagY is 0 (when a specific height condition is not met), the height of the storage unit may be determined to be 8, and accordingly, mvUnit_y may be 3. At least one of the storage unit variables, mvUnit_x derived according to the value of mvCompFlagX or mvUnit_y derived according to the value of mvCompFlagY, may derive the location for obtaining movement information of the call block. mvUnit_x and mvUnit_y may have different values, and in this case, the shape of the storage unit may be rectangular.

[0135] Hereinafter, with reference to Tables 5 to 8, we will specifically examine a method for deriving the location where movement information of a call block is obtained based on storage unit variables. In Tables 5 to 8, the location of at least one of the derivation processes of mvCompFlagX, mvCompFlagY, mvUnit_x, or mvUnit_y may be changed.

[0136] Table 5 is an example of a method to derive the time motion vector predictor for the current block.

[0137] 8.5.2.11 Derivation process for temporal luma motion vector predictionInputs to this process are:- a luma location ( xCb, yCb ) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture,- a variable cbWidth specifying the width of the current coding block in luma samples,- a variable cbHeight specifying the height of the current coding block in luma samples,- a reference index refIdxLX, with X being 0 or 1.Outputs of this process are:- the motion vector prediction mvLXCol in 1 / 16 fractional-sample accuracy,- the availability flag availableFlagLXCol.The variable currCb specifies the current luma coding block at luma location ( xCb, yCb ).The variable mvCompFlagX, mvCompFlagY is equal to 0.When the following conditions is true, mvCompFlagX is set equal to 1.- The width of the current picture is less than 1920.When the following conditions is true, mvCompFlagY is set equal to 1.- The height of the current picture is less than 1080.The motion compression unit is derived as follows.mvUnit_x= mvCompFlagX ? 2 : 3mvUnit_y= mvCompFlagY ? 2 : 3The variables mvLXCol and availableFlagLXCol are derived as follows:- If ph_temporal_mvp_enabled_flag is equal to 0 or ( cbWidth * cbHeight ) is less than or equal to 32. 2)when mvUnit_x is greater than or equal to 3 and when mvUnit_y is greater than or equal to 3, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.- Otherwise (ph_temporal_mvp_enabled_flag is equal to 1), the following ordered steps apply:3.The bottom-right collocated motion vector and the bottom and right boundary sample locations are derived as follows:xColBr = xCb + cbWidth (583)yColBr = yCb + cbHeight (584)rightBoundaryPos = sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] ?SubpicRightBoundaryPos : pps_pic_width_in_luma_samples - 1 (585)botBoundaryPos = sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] SubpicBotBoundaryPos : pps_pic_height_in_luma_samples - 1 (586)- If yCb >> CtbLog2SizeY is equal to yColBr >> CtbLog2SizeY, yColBr is less than or equal to botBoundaryPos and xColBr is less than or equal to rightBoundaryPos, the following applies:- The luma location ( xColCb, yColCb ) is set equal to ( ( xColBr >> mvUnit_x ) << mvUnit_x, ( yColBr >> mvUnit_y ) << mvUnit_y ).- The variable colCb specifies the luma coding block covering the location ( xColCb, yColCb ) inside the collocated picture specified by ColPic.- The derivation process for collocated motion vectors as specified in clause 8.5.2.12 is invoked with currCb, colCb, ( xColCb, yColCb ), refIdxLX and sbFlag set equal to 0 as inputs, and the output is assigned to mvLXCol and availableFlagLXCol.- Otherwise, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.4. When availableFlagLXCol is equal to 0, the central collocated motion vector is derived as follows:xColCtr = xCb + ( cbWidth >> 1 ) (587)yColCtr = yCb + ( cbHeight >> 1 ) (588)- The luma location ( xColCb, yColCb ) is set equal ( ( xColCtr >> mvUnit_x ) << mvUnit_x, ( yColCtr >> mvUnit_y ) << mvUnit_y ).- The variable colCb specifies the luma coding block covering the location ( xColCb, yColCb ) inside the collocated picture specified by ColPic.The derivation process for collocated motion vectors as specified in clause 8.5.2.12 is invoked with currCb, colCb, ( xColCb, yColCb ), refIdxLX and sbFlag set equal to 0 as inputs, and the output is assigned to(...).

[0138] Referring to Table 5, the location (xColCb, yColCb) (hereinafter referred to as the storage location) derived based on the storage unit can be derived according to at least one of the storage unit width variable (mvUnit_x) or the storage unit height variable (mvUnit_y). Movement information of the call block can be obtained based on the storage unit block containing the storage location (xColCb, yColCb). The movement information may include a movement vector. More specifically, xColCb may be a location obtained by shifting the reference position x-coordinate of the call block according to the storage unit corresponding to mvUnit_x, and yColCb may be a location obtained by shifting the reference position y-coordinate of the call block according to the storage unit corresponding to mvUnit_y. xColCb may be the x-coordinate of the reference position of the call block shifted to the right by the value of mvUnit_x and then shifted to the left by the value of mvUnit_x, and yColCb may be the y-coordinate of the reference position of the call block shifted to the right by the value of mvUnit_y and then shifted to the left by the value of mvUnit_y. The reference position coordinates of the call block may be either the coordinates of a sample at a position corresponding to the coordinates of the sample at the bottom-right position of the current block within the call picture (xColBr, yColBr) or the coordinates of a sample at a position corresponding to the coordinates of the sample at the center position of the current block within the call picture (xColCtr, yColCtr).

[0139] mvUnit_x can be determined based on the value of a width condition fulfillment flag (mvCompFlagX) indicating whether a specific width condition for determining the width of a storage unit is satisfied. Additionally, mvUnit_y can be determined based on the value of a height condition fulfillment flag (mvCompFlagY) indicating whether a specific height condition for determining the height of a storage unit is satisfied. Referring to Table 5, mvUnit_x can be 2 when the value of mvCompFlagX is 1, and can be 3 when the value of mvCompFlagX is 0. Additionally, mvUnit_y can be 2 when the value of mvCompFlagY is 1, and can be 3 when the value of mvCompFlagY is 0. mvUnit_x can be a variable having a value corresponding to the width of a storage unit, and mvUnit_y can be a variable having a value corresponding to the height of a storage unit. For example, if mvUnit_x or mvUnit_y is 0, 1, 2, 3, or 4, respectively, the width or height of the storage unit may be 1, 2, 4, 8, or 16, respectively.

[0140] mvCompFlagX may indicate whether a specific width condition is satisfied for determining the width of a storage unit. The specific width condition may be a result of comparing the width of the current picture with a predetermined width. Referring to Table 5, mvCompFlagX may be set to 1 if the width of the current picture is smaller than the predetermined width, and may be set to 0 if the width of the current picture is greater than or equal to the predetermined width. Here, the predetermined width may be 1920.

[0141] mvCompFlagY may indicate whether a specific height condition is satisfied for determining the height of the storage unit. The specific height condition may be a result of comparing the height of the current picture with a predetermined height. Referring to Table 5, mvCompFlagY may be set to 1 if the height of the current picture is less than the predetermined height, and may be set to 0 if the height of the current picture is greater than or equal to the predetermined height. Here, the predetermined height may be 1080.

[0142] The time motion vector predictor for the current block may not be induced when the time motion vector predictor enable flag (ph_temporal_mvp_enabled_flag) is 0, or when the product of the width and height of the current block is less than or equal to 32, and mvUnit_x has a value greater than or equal to 3, and mvUnit_y has a value greater than or equal to 3. More specifically, when the time motion vector predictor enable flag (ph_temporal_mvp_enabled_flag) is 0, or when the product of the width and height of the current block is less than or equal to 32, and mvUnit_x has a value greater than or equal to 3, and mvUnit_y has a value greater than or equal to 3, the time motion vector predictor (mvLXcol) is set to 0, and the time motion vector predictor enable flag (availableFlagLXCol), which indicates whether the time motion vector predictor is available for the current block, may be set to 0.

[0143] Table 6 is an example of a method to derive subblock-based time merge candidates.

[0144] 8.5.5.3 Derivation process for subblock-based temporal merging candidatesInputs to this process are:- a luma location ( xCb, yCb ) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture,- a variable cbWidth specifying the width of the current coding block in luma samples,- a variable cbHeight specifying the height of the current coding block in luma samples.- the availability flag availableFlagA1of the neighbouring coding unit,- the reference indices refIdxLXA1of the neighbouring coding unit with X = 0..1,- the prediction list utilization flags predFlagLXA1of the neighbouring coding unit with X = 0..1,- the motion vector in 1 / 16 fractional-sample accuracy mvLXA1of the neighbouring coding unit with X = 0..1.Outputs of this process are:- the availability flag availableFlagSbCol,- the number of luma subblocks in horizontal direction numSbX and in vertical direction numSbY,- the reference indices refIdxL0SbCol and refIdxL1SbCol,- the luma motion vectors in 1 / 16 fractional-sample accuracy mvL0SbCol[ xSbIdx ][ ySbIdx ] and mvL1SbCol[ xSbIdx ][ ySbIdx ] with xSbIdx = 0..numSbX - 1, ySbIdx = 0..numSbY - 1,- the prediction list utilization flags predFlagL0SbCol[ xSbIdx ][ ySbIdx ] and predFlagL1SbCol[ xSbIdx ][ ySbIdx ] with xSbIdx = 0..numSbX - 1, ySbIdx = 0..numSbY - 1.The variable mvCompFlagX, mvCompFlagY is equal to 0.When the following conditions is true, mvCompFlagX is set equal to 1.- The width of the current picture is less than 1920.When the following conditions is true, mvCompFlagY is set equal to 1.- The height of the current picture is less than 1080.The motion compression unit is derived as follows.mvUnit_x= mvCompFlagX ? 2 : 3mvUnit_y= mvCompFlagY ? 2 : 3The availability flag availableFlagSbCol is derived as follows.- If one or more of the following conditions are true, availableFlagSbCol is set equal to 0.- ph_temporal_mvp_enabled_flag is equal to 0.- sps_sbtmvp_enabled_flag is equal to 0.- cbWidth is less than 8 when mvUnit_x is greater than or equal to 3.- cbHeight is less than 8 when mvUnit_y is greater than or equal to 3.- Otherwise, the following ordered steps apply:4. The location ( xCtb, yCtb ) of the top-left sample of the luma coding tree block that contains the current coding block and the location ( xCtrCb, yCtrCb ) of the below-right center sample of the current luma coding block are derived as follows:xCtb = ( xCb >> CtbLog2SizeY ) << CtbLog2SizeY (702)yCtb = ( yCb >> CtbLog2SizeY ) << CtbLog2SizeY (703)xCtrCb = xCb + ( cbWidth / 2 ) (704)yCtrCb = yCb + ( cbHeight / 2 ) (705)5.The derivation process for subblock-based temporal merging base motion data as specified in clause 8.5.5.4 is invoked with the location ( xCtb, yCtb ), the location ( xCtrCb, yCtrCb ), the availability flag availableFlagA1, and the prediction list utilization flag predFlagLXA1, and the reference index refIdxLXA1, and the motion vector mvLXA1, with X = 0..1 as inputs and the motion vectors ctrMvLX, and the prediction list utilization flags ctrPredFlagLX of the collocated block, with X = 0..1, and the temporal motion vector tempMv as outputs.6. The variable availableFlagSbCol is derived as follows:- If both ctrPredFlagL0 and ctrPredFlagL1 are equal to 0, availableFlagSbCol is set equal to 0.- Otherwise, availableFlagSbCol is set equal to 1.When availableFlagSbCol is equal to 1, the following applies:- The variables numSbX, numSbY, sbWidth, sbHeight and refIdxLXSbCol are derived as follows:numSbX = cbWidth >> 2 (706)numSbY = cbHeight >> 2 (707)sbWidth = cbWidth / numSbX (708)sbHeight = cbHeight / numSbY (709)refIdxLXSbCol = 0 (710)- For xSbIdx = 0..numSbX - 1 and ySbIdx = 0..numSbY - 1, the motion vectors mvLXSbCol[ xSbIdx ][ ySbIdx ] and prediction list utilization flags predFlagLXSbCol[ xSbIdx ][ ySbIdx ] are derived as follows:- The luma location ( xSb, ySb ) specifying the below-right center sample of the current subblock relative to the top-left luma sample of the current picture is derived as follows:xSb = xCb + xSbIdx * sbWidth + sbWidth / 2 (711)ySb = yCb + ySbIdx * sbHeight + sbHeight / 2 (712)- The location ( xColSb, yColSb ) of the collocated subblock inside ColPic is derived as follows.- The following applies:yColSb = Clip3( yCtb,Min( pps_pic_height_in_luma_samples - 1, yCtb + ( 1 << CtbLog2SizeY ) - 1 ), (713)ySb + tempMv[ 1 ] )- If sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] is equal to 1, the following applies:xColSb = Clip3( xCtb,Min( SubpicRightBoundaryPos, xCtb + ( 1 << CtbLog2SizeY ) + 3 ), (714)xSb + tempMv[ 0 ] )- Otherwise ( sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] is equal to 0), the following applies:xColSb = Clip3( xCtb,Min( pps_pic_width_in_luma_samples - 1, xCtb + ( 1 << CtbLog2SizeY ) + 3 ), (715)xSb + tempMv[ 0 ] )- The variable currCb specifies the luma coding block covering the current subblock inside the current picture.- The luma location ( xColCb, yColCb ) is set equal to ( ( xColSb >> mvUnit_x ) << mvUnit_x, ( yColSb >> mvUnit_y ) << mvUnit_y ).- The variable colCb specifies the luma coding block covering the location ( xColCb, yColCb ) inside the collocated picture specified by ColPic.- The prediction list utilization flags predFlagL0SbCol[ xSbIdx ][ ySbIdx ] and predFlagL1SbCol[ xSbIdx ][ ySbIdx ] are initialized to be equal to FALSE.- The derivation process for collocated motion vectors as specified in clause 8.5.2.12 is invoked with currCb, colCb, ( xColCb, yColCb ), refIdxL0 set equal to 0 and sbFlag set equal to 1 as inputs and the output being assigned to the motion vector of the subblock mvL0SbCol[ xSbIdx ][ ySbIdx ] and predFlagL0SbCol[ xSbIdx ][ ySbIdx ].- When sh_slice_type is equal to B, the derivation process for collocated motion vectors as specified in clause 8.5.2.12 is invoked with currCb, colCb, ( xColCb, yColCb ), refIdxL1 set equal to 0 and sbFlag set equal to 1 as inputs and the output being assigned to the motion vector of the subblock mvL1SbCol[ xSbIdx ][ ySbIdx ] and predFlagL1SbCol[ xSbIdx ][ ySbIdx ].- When predFlagL0SbCol[ xSbIdx ][ ySbIdx ] and predFlagL1SbCol[ xSbIdx ][ ySbIdx ] are both equal to 0, the following applies for X = 0..1:mvLXSbCol[ xSbIdx ][ ySbIdx ] = ctrMvLX (716)predFlagL

[0145] Referring to Table 6, the location (xColCb, yColCb) (hereinafter referred to as the storage location) derived based on the storage unit can be derived according to at least one of the storage unit width variable (mvUnit_x) or the storage unit height variable (mvUnit_y). Motion information corresponding to the storage unit block containing the storage location (xColCb, yColCb) can be used as motion information for the sub-block unit. More specifically, xColCb may be a location obtained by shifting the reference position x-coordinate (xColSb) of the corresponding location sub-block (hereinafter referred to as the call sub-block) within the call picture (colPic) according to the storage unit corresponding to mvUnit_x, and yColCb may be a location obtained by shifting the reference position y-coordinate (yColSb) of the call sub-block according to the storage unit corresponding to mvUnit_y. xColCb may be the value obtained by right-shifting the reference position x-coordinate (xColSb) of the call subblock by the value of mvUnit_x and then left-shifting it by the value of mvUnit_x, and yColCb may be the value obtained by right-shifting the reference position y-coordinate (yColSb) of the call subblock by the value of mvUnit_y and then left-shifting it by the value of mvUnit_y.

[0146] Referring to Table 6, subblock-based time merge candidates may not be derived if certain conditions are satisfied. More specifically, if certain conditions are satisfied, a subblock-based time merge candidate usage flag (availableFlagSbCol), which indicates whether a subblock-based time merge candidate is used in the current block, may be set to 0. The said certain conditions may include at least one of the following: when the time motion vector predictor enable flag (ph_temporal_mvp_enabled_flag) is 0; when the subblock-based time merge candidate enable flag (sps_sbtmvp_enabled_flag) is 0; when mvUnit_x is greater than or equal to 3 and the width (cbWidth) of the current block is less than 8; or when mvUnit_y is greater than or equal to 3 and the height (cbHeight) of the current block is less than 8.

[0147] Referring to Table 6, the reference position coordinates (xColSb, yColSb) of a call subblock can be derived based on the reference position coordinates (xSb, ySb) of a subblock. The reference position coordinates of a subblock can be derived based on the number of subblock units (numSbX or numSbY). Here, the size of the subblock unit may be 4x4. In this case, the number of subblocks of the current block may be 1.

[0148] The descriptions of mvUnit_x, mvUnit_y, mvCompFlagX, and mvCompFlagY are as referenced in Table 5.

[0149] Table 7 is an example of a method for deriving basic movement information of subblock-based time merge candidates.

[0150] 8.5.5.4 Derivation process for subblock-based temporal merging base motion dataInputs to this process are:- the location ( xCtb, yCtb ) of the top-left sample of the luma coding tree block that contains the current coding block,- the location ( xCtrCb, yCtrCb ) of the top-left sample of the collocated luma coding block that covers the below-right center sample.- the availability flag availableFlagA1of the neighbouring coding unit,- the reference indices refIdxLXA1of the neighbouring coding unit with X = 0..1,- the prediction list utilization flags predFlagLXA1of the neighbouring coding unit with X = 0..1,- the motion vectors in 1 / 16 fractional-sample accuracy mvLXA1of the neighbouring coding unit with X = 0..1.Outputs of this process are:- the motion vectors ctrMvL0 and ctrMvL1,- the prediction list utilization flags ctrPredFlagL0 and ctrPredFlagL1,- the temporal motion vector tempMv.The variable tempMv is set as follows:tempMv[ 0 ] = 0 (718)tempMv[ 1 ] = 0 (719)The variable currPic specifies the current picture.When availableFlagA1is equal to TRUE, the following applies:- If all of the following conditions are true, tempMv is set equal to mvL0A1:- predFlagL0A1is equal to 1,- DiffPicOrderCnt( ColPic, RefPicList[ 0 ][ refIdxL0A1 ] ) is equal to 0,- Otherwise, if all of the following conditions are true, tempMv is set equal to mvL1A1:- sh_slice_type is equal to B,- predFlagL1A1is equal to 1,- DiffPicOrderCnt( ColPic, RefPicList[ 1 ][ refIdxL1A1 ] ) is equal to 0.The rounding process for motion vectors as specified in clause 8.5.2.14 is invoked with mvX set equal to tempMv, rightShift set equal to 4, and leftShift set equal to 0 as inputs and the rounded tempMv as output.The location ( xColCb, yColCb ) of the collocated block inside ColPic is derived as follows.- The following applies:yColCb = Clip3( yCtb,Min( pps_pic_height_in_luma_samples - 1, yCtb + ( 1 << CtbLog2SizeY ) - 1 ), (720)yCtrCb + tempMv[ 1 ] )- If sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] is equal to 1, the following applies:xColCb = Clip3( xCtb, Min( SubpicRightBoundaryPos, xCtb + ( 1 << CtbLog2SizeY ) + 3 ), (721)xCtrCb + tempMv[ 0 ] )- Otherwise ( sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] is equal to 0), the following applies:xColCb = Clip3( xCtb,Min( pps_pic_width_in_luma_samples - 1, xCtb + ( 1 << CtbLog2SizeY ) + 3 ), (722)xCtrCb + tempMv[ 0 ] )The array colPredMode is set equal to the prediction mode array CuPredMode[ 0 ] of the collocated picture specified by ColPic.The motion vectors ctrMvL0 and ctrMvL1, and the prediction list utilization flags ctrPredFlagL0 and ctrPredFlagL1 are derived as follows:The variable mvCompFlagX, mvCompFlagY is equal to 0.When the following conditions is true, mvCompFlagX is set equal to 1.- The width of the current picture is less than 1920.When the following conditions is true, mvCompFlagY is set equal to 1.- The height of the current picture is less than 1080.- The motion compression unit is derived according to mvCompFlagX, mvCompFlagY- mvUnit_x = mvCompFlagX ? 2 : 3- mvUnit_y = mvCompFlagY ? 2 : 3- If colPredMode[ ( xColCb >> mvUnit_x ) << mvUnit_x ][ ( yColCb >> mvUnit_y ) << mvUnit_y ] is equal to MODE_INTER, the following applies:- The variable currCb specifies the luma coding block covering ( xCtrCb, yCtrCb ) inside the current picture.- The luma location ( xColCb, yColCb ) is set equal to ( ( xColCb >> mvUnit_x ) << mvUnit_x, ( yColCb >> mvUnit_y ) << mvUnit_y ).- The variable colCb specifies the luma coding block covering the location ( xColCb, yColCb ) inside the collocated picture specified by ColPic.- The prediction list utilization flags ctrPredFlagL0 and ctrPredFlagL1 are initialized to be equal to FALSE.- The derivation process for collocated motion vectors specified in clause 8.5.2.12 is invoked with currCb, colCb, (xColCb, yColCb), refIdxL0 set equal to 0, and sbFlag set equal to 1 as inputs and the output being assigned to ctrMvL0 and ctrPredFlagL0.- When sh_slice_type is equal to B, the derivation process for collocated motion vectors specified in clause 8.5.2.12 is invoked with currCb, colCb, (xColCb, yColCb), refIdxL1 set equal to 0, and sbFlag set equal to 1 as inputs and the output being assigned to ctrMvL1 and ctrPredFlagL1.- Otherwise, the following applies:ctrPredFlagL0 = 0 (723)ctrPredFlagL1 = 0 (724).

[0151] Referring to Table 7, movement information of the current block can be derived from a call block at a specific location. Here, the specific location can be derived based on a storage unit variable (mvUnit) for movement information. For example, the specific location (xColCb, yColCb) may be a location shifted by mvUnit_x or mvUnit_y from a corresponding location within the call block that corresponds to a location derived based on the bottom-right surrounding sample location (or center sample location) of the current block. More specifically, xColCb may be a location shifted by the value of mvUnit_x from the corresponding location within the call block. xColCb may be a location shifted to the right by the value of mvUnit_x and then to the left by the value of mvUnit_x from the corresponding location within the call block. Additionally, yColCb may be a location shifted by the value of mvUnit_y from the corresponding location within the call block. yColCb may be the corresponding position within the call block, right-shifted by the value of mvUnit_y and left-shifted by the value of mvUnit_y.

[0152] The descriptions of mvUnit_x, mvUnit_y, mvCompFlagX, and mvCompFlagY are as referenced in Table 5.

[0153] Table 8 is an example of a method for deriving a time motion vector predictor for the lower right control point during affine prediction.

[0154] The fourth (collocated bottom-right) control point motion vector cpMvLXCorner[ 3 ], reference index refIdxLXCorner[ 3 ], prediction list utilization flag predFlagLXCorner[ 3 ] and the availability flag availableFlagCorner[ 3 ] with = 0..1 are derived as follows:- The reference indices for the temporal merging candidate, refIdxLXCorner[ 3 ], with X = 0..1, are set equal to 0.- For X = 0..1, the variables mvLXCol and availableFlagLXCol are derived as follows:- If ph_temporal_mvp_enabled_flag is equal to 0, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.- Otherwise (ph_temporal_mvp_enabled_flag is equal to 1), the following applies:xColBr = xCb + cbWidth (761)yColBr = yCb + cbHeight (762)rightBoundaryPos = sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] ? SubpicRightBoundaryPos :pps_pic_width_in_luma_samples - 1 (763)botBoundaryPos = sps_subpic_treated_as_pic_flag[ CurrSubpicIdx ] SubpicBotBoundaryPos :pps_pic_height_in_luma_samples - 1 (764)mvUnit_x= mvCompFlagX 2 : 3mvUnit_y= mvCompFlagY ? 2 : 3- If yCb >> CtbLog2SizeY is equal to yColBr >> CtbLog2SizeY, yColBr is less than or equal to botBoundaryPos and xColBr is less than or equal to rightBoundaryPos, the following applies:- The luma location ( xColCb, yColCb ) is set equal to ( ( xColBr >> mvUnit_x ) << mvUnit_x, ( yColBr >> mvUnit_y ) << mvUnit_y ).- The variable colCb specifies the luma coding block covering the location ( xColCb, yColCb ) inside the collocated picture specified by ColPic.- The derivation process for collocated motion vectors as specified in clause 8.5.2.12 is invoked with currCb, colCb, ( xColCb, yColCb ), refIdxLXCorner[ 3 ] and sbFlag set equal to 0 as inputs, and the output is assigned to mvLXCol and availableFlagLXCol.- Otherwise, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.- The variables availableFlagCorner[ 3 ], predFlagL0Corner[ 3 ], cpMvL0Corner[ 3 ] and predFlagL1Corner[ 3 ] are derived as follows:availableFlagCorner[ 3 ] = availableFlagL0Col (765)predFlagL0Corner[ 3 ] = availableFlagL0Col (766)cpMvL0Corner[ 3 ] = mvL0Col (767)predFlagL1Corner[ 3 ] = 0 (768)When sh_slice_type is equal to B, the variables availableFlagCorner[ 3 ], predFlagL1Corner[ 3 ] and cpMvL1Corner[ 3 ] are derived as follows:availableFlagCorner[ 3 ] = availableFlagL0Col | | availableFlagL1Col (769)predFlagL1Corner[ 3 ] = availableFlagL1Col (770)cpMvL1Corner[ 3 ] = mvL1Col (771).

[0155] Referring to Table 8, the location (xColCb, yColCb) (hereinafter referred to as the storage location) derived based on the storage unit can be derived based on at least one of the storage unit width variable (mvUnit_x) or the storage unit height variable (mvUnit_y). Movement information of the call block can be obtained based on the storage unit block containing the storage location (xColCb, yColCb). The movement information may include a movement vector. More specifically, xColCb may be a location obtained by shifting the reference position x-coordinate of the call block according to the storage unit corresponding to mvUnit_x, and yColCb may be a location obtained by shifting the reference position y-coordinate of the call block according to the storage unit corresponding to mvUnit_y. xColCb may be the x-coordinate of the reference position of the call block shifted to the right by the value of mvUnit_x and then shifted to the left by the value of mvUnit_x, and yColCb may be the y-coordinate of the reference position of the call block shifted to the right by the value of mvUnit_y and then shifted to the left by the value of mvUnit_y. The reference position coordinates of the call block may be the coordinates (xColBr, yColBr) of the sample at the position corresponding to the coordinates of the sample at the bottom-right position of the current block within the call picture.

[0156] The descriptions of mvUnit_x, mvUnit_y, mvCompFlagX, and mvCompFlagY are as referenced in Table 5.

[0157] Referring to FIG. 4, a time motion vector predictor of the current block can be derived based on the motion information of the corresponding position block (S420).

[0158] The time motion vector predictor (TMVP) of the current block can be derived based on motion information of the corresponding location block (call block). In this case, the motion information of the call block can be obtained based on the location derived based on the storage unit. The TMVP of the current block can be derived based on the motion vector included in the motion information of the call block. For example, the TMVP of the current block may be the motion vector of the call block.

[0159] Referring to FIG. 4, a prediction sample of the current block can be derived based on the time motion vector predictor of the current block (S430).

[0160] Predicted samples of the current block can be generated based on the motion vector of the current block derived from the time motion vector predictor of the current block.

[0161] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs a decoding method according to the present disclosure.

[0162] Referring to FIG. 5, the decoding device (300) may include a time motion vector predictor induction unit (500), a prediction sample induction unit (510), and a restoration unit (520). The time motion vector predictor induction unit (500) and the prediction sample induction unit (510) may be provided in the inter prediction unit (332) of FIG. 3.

[0163] The time motion vector predictor derivation unit (500) can perform a process of determining the corresponding position block of the current block according to S400. Additionally, the time motion vector predictor derivation unit (500) can perform a process of acquiring movement information of the corresponding position block according to S410. Additionally, the time motion vector predictor derivation unit (500) can perform a process of deriving the time motion vector predictor of the current block based on the movement information of the corresponding position block according to S420.

[0164] The prediction sample induction unit (510) can perform a process of inducing a prediction sample of the current block based on the time movement vector predictor of the current block according to S430. The restoration unit (520) can perform a restoration process of the current block based on the prediction sample derived by the prediction sample induction unit (510).

[0165] FIG. 6 illustrates an encoding method performed by an encoding device (200) as an embodiment according to the present disclosure.

[0166] The corresponding location block of the current block can be determined (S600). The method for determining the corresponding location block of the current block is as described with reference to FIG. 4.

[0167] Movement information of the corresponding position block can be obtained (S610). The method for obtaining movement information of the corresponding position block is as described with reference to FIG. 4.

[0168] A time motion vector predictor for the current block can be derived based on the motion information of the corresponding position block (S620). The method for deriving the time motion vector predictor for the current block based on the motion information of the corresponding position block is as described with reference to FIG. 4.

[0169] A prediction sample of the current block can be derived based on the time motion vector predictor of the current block (S630). A method for deriving the time motion vector predictor of the current block based on at least one of the basic motion vector of the current block or the motion information of the corresponding position block of the current block is as described with reference to FIG. 4.

[0170] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs an encoding method according to the present disclosure.

[0171] Referring to FIG. 7, the encoding device (200) may include a time motion vector predictor derivation unit (700), a prediction sample derivation unit (710), a residual sample derivation unit (720), a transformation coefficient derivation unit (730), and a residual information encoding unit (740).

[0172] The time motion vector predictor derivation unit (700) and the prediction sample derivation unit (710) may be provided in the inter prediction unit (221) of FIG. 2. The residual sample derivation unit (720) and the transformation coefficient derivation unit (730) may be provided in the residual processing unit (230) of FIG. 2. The residual information encoding unit (740) may be provided in the entropy encoding unit (240).

[0173] The time motion vector predictor derivation unit (700) can perform a process of determining the corresponding position block of the current block according to S600. Additionally, the time motion vector predictor derivation unit (700) can perform a process of obtaining movement information of the corresponding position block according to S610. Additionally, the time motion vector predictor derivation unit (700) can perform a process of deriving the time motion vector predictor of the current block based on the movement information of the corresponding position block according to S620.

[0174] The prediction sample derivation unit (710) can perform the process of deriving a prediction sample of the current block based on the time movement vector predictor of the current block according to S630. The residual sample derivation unit (720) can perform the process of deriving a residual sample based on the prediction sample derived by the prediction sample derivation unit (710). The transformation coefficient derivation unit (730) can perform the process of deriving a transformation coefficient based on the residual sample derived by the residual sample derivation unit (720). The residual information encoding unit (740) can perform the process of encoding residual information based on the transformation coefficient derived by the transformation coefficient derivation unit (730).

[0175] In the embodiments described above, methods are described based on flowcharts as a series of steps or blocks; however, the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps as described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and other steps may be included, or one or more steps of the flowcharts may be omitted without affecting the scope of the embodiments of this document.

[0176] The method according to the embodiments of the present document described above may be implemented in the form of software, and the encoding device and / or decoding device according to the present document may be included in a device that performs image processing, such as a TV, computer, smartphone, set-top box, display device, etc.

[0177] When the embodiments described in this document are implemented in software, the method described above may be implemented as a module (process, function, etc.) that performs the function described above. The module may be stored in memory and executed by a processor. The memory may be located inside or outside the processor and may be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.

[0178] In addition, the decoding device and encoding device to which the embodiment(s) of the present specification are applied may be included in multimedia broadcasting transmission and reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, Video on Demand (VoD) service providers, Over-the-top video (OTT) devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, video phone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video devices, and may be used to process video signals or data signals. For example, Over-the-top video (OTT) devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, Digital Video Recorders (DVRs), etc.

[0179] Additionally, the processing method to which the embodiment(s) of this specification are applied may be produced in the form of a program that is executed by a computer and may be stored on a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of this specification may also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Additionally, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission over the Internet). Additionally, a bitstream generated by an encoding method may be stored on a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0180] Additionally, the embodiments of this specification may be implemented as a computer program product by program code, and said program code may be executed on a computer by the embodiments of this specification. said program code may be stored on a computer-readable carrier.

[0181] FIG. 8 shows an example of a content streaming system to which embodiments of the present disclosure can be applied.

[0182] Referring to FIG. 8, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0183] The above encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server may be omitted.

[0184] The bitstream above may be generated by an encoding method or a bitstream generation method to which the embodiment(s) of the present specification are applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0185] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server plays the role of controlling commands and responses between each device within the content streaming system.

[0186] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, to provide a seamless streaming service, the streaming server may store the bitstream for a certain period of time.

[0187] Examples of the above user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.

[0188] Each server within the above-mentioned content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.

[0189] The claims described in this specification may be combined in various ways. For example, the technical features of the method claims in this specification may be combined to be implemented as a device, and the technical features of the device claims in this specification may be combined to be implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a device, and the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a method.

Claims

1. A step of determining the corresponding location block of the current block; A step of obtaining movement information of the above-mentioned corresponding position block; A step of deriving a time motion vector predictor of the current block based on the motion information of the corresponding position block; and The method includes the step of deriving a prediction sample of the current block based on a time motion vector predictor of the current block, wherein An image decoding method in which movement information of the above-mentioned corresponding position block is obtained based on a position induced based on a storage unit.

2. In Paragraph 1, The above storage unit is determined based on whether specific conditions are met, and The above specific condition is a video decoding method in which the result of comparing the number of samples of the current picture with a predetermined number.

3. In Paragraph 2, A video decoding method in which the above storage unit is determined as a 4x4 unit when the number of samples in the current picture is less than a predetermined number, and as an 8x8 unit when the number of samples in the current picture is greater than or equal to the predetermined number.

4. In Paragraph 2, The position derived based on the above storage unit is a coordinate derived by performing a shift operation on the reference position coordinates of the above-mentioned corresponding position block based on the storage unit variable, and A video decoding method in which the reference position coordinates of the corresponding position block are either the coordinates of a sample at a position corresponding to the coordinates of a sample at the lower right position of the current block within the corresponding position picture of the current picture, or the coordinates of a sample at a position corresponding to the coordinates of a sample at a center position of the current block within the corresponding position picture.

5. In Paragraph 4, An image decoding method in which a position derived based on the above storage unit is a coordinate obtained by shifting the reference position coordinate of the corresponding position block by the value of the above storage unit variable.

6. In Paragraph 4, The above storage unit variable is determined based on the value of a condition fulfillment flag indicating whether the above specific condition is satisfied, and If the above condition satisfaction flag is 1, the above storage unit variable is determined to be 2, and A video decoding method in which, when the above condition satisfaction flag is 0, the above storage unit variable is determined to be 3.

7. In Paragraph 6, If the number of samples in the current picture is less than 2,073,600, the condition fulfillment flag is set to 1, and A video decoding method in which, if the number of samples of the current picture is greater than or equal to 2,073,600, the condition fulfillment flag is set to 0.

8. In Paragraph 1, The width of the above storage unit is determined based on whether a specific width condition is satisfied, and The height of the above storage unit is determined based on whether a specific height condition is satisfied, and The above specific width condition relates to the result of comparing the width of the current picture with a predetermined width, and The above-mentioned specific height condition is a video decoding method in which the result of comparing the height of the current picture with a predetermined height.

9. In Paragraph 8, The position derived based on the above storage unit is a coordinate derived by performing a shift operation on the reference position coordinate of the corresponding position block based on at least one of the storage unit width variable or the storage unit height variable, and A video decoding method in which the reference position coordinates of the corresponding position block are either the coordinates of a sample at a position corresponding to the coordinates of a sample at the lower right position of the current block within the corresponding position picture of the current picture, or the coordinates of a sample at a position corresponding to the coordinates of a sample at a center position of the current block within the corresponding position picture.

10. In Paragraph 9, The x-coordinate of the position derived based on the above storage unit is a coordinate obtained by shifting the x-coordinate of the reference position of the above-mentioned position block by the value of the above-mentioned storage unit width variable, and An image decoding method in which the y-coordinate of a position derived based on the above storage unit is a coordinate obtained by shifting the y-coordinate of the reference position of the above-mentioned position block by the value of the above-mentioned storage unit height variable.

11. In Paragraph 9, The above storage unit width variable is determined based on the value of a width condition satisfaction flag indicating whether the above specific width condition is satisfied, and An image decoding method in which the above storage unit height variable is determined based on the value of a height condition fulfillment flag indicating whether the above specific height condition is satisfied.

12. In Paragraph 11, If the width of the current picture above is less than 1920, the width condition satisfaction flag is set to 1, and If the width of the current picture is greater than or equal to 1920, the width condition satisfaction flag is set to 0, and If the height of the current picture above is less than 1080, the height condition satisfaction flag is set to 1, and A video decoding method in which, if the height of the current picture is greater than or equal to 1080, the height condition satisfaction flag is set to 0.

13. A step of determining the corresponding location block of the current block; A step of obtaining movement information of the above-mentioned corresponding position block; A step of deriving a time motion vector predictor of the current block based on the motion information of the corresponding position block; and The method includes the step of deriving a prediction sample of the current block based on a time motion vector predictor of the current block, wherein An image encoding method in which movement information of the above-mentioned corresponding position block is obtained based on a position induced based on a storage unit.

14. A computer-readable storage medium for storing a bitstream generated by the method according to paragraph 13.

15. A step of acquiring a bitstream for image information; wherein the bitstream is generated based on a step of determining a corresponding position block of a current block, a step of acquiring motion information of the corresponding position block, a step of deriving a time motion vector predictor of the current block based on the motion information of the corresponding position block, and a step of deriving a prediction sample of the current block based on the time motion vector predictor of the current block, and The method includes the step of transmitting data including the above bitstream, A method in which movement information of the above-mentioned corresponding position block is obtained based on a position induced based on a storage unit.