Image encoding / decoding method and device, and recording medium for storing bitstream

WO2026197743A1PCT designated stage Publication Date: 2026-09-24LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/004260
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-16
Filing Date
2026-03-16
Publication Date
2026-09-24

Smart Images

  • Figure KR2026004260_24092026_PF_FP_ABST
    Figure KR2026004260_24092026_PF_FP_ABST
Patent Text Reader

Abstract

An image decoding method and device according to the present disclosure may: derive a basic motion vector of the current block; determine a corresponding position picture of the current picture including the current block; determine a corresponding position block of the current block in the corresponding position picture; acquire motion information about the corresponding position block; derive a temporal motion vector predictor of the current block on the basis of at least one among the basic motion vector and the motion information about the corresponding position block; and derive a prediction sample of the current block on the basis of at least one among the temporal motion vector predictor of the current block and a temporal motion vector predictor (TMVP) reference picture. Here, the position of an upper left sample in the corresponding position block may be derived on the basis of the basic motion vector, and the motion information about the corresponding position block may include information about the motion vector of the corresponding position block.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding / decoding method and device, and a recording medium storing a bitstream

[0001] The present invention relates to a video encoding / decoding method and apparatus, and a recording medium storing a bitstream.

[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition) video, has been increasing across various application fields, and accordingly, high-efficiency video compression technologies are being discussed.

[0003] Various image compression technologies exist, such as inter-prediction methods that predict pixel values ​​in the current picture from previous or subsequent pictures, intra-prediction methods that predict pixel values ​​in the current picture using pixel information within the current picture, and entropy coding techniques that assign short codes to values ​​with high frequency and long codes to values ​​with low frequency; by utilizing these image compression technologies, image data can be effectively compressed for transmission or storage.

[0004] The present disclosure provides a method and apparatus for deriving a temporal motion vector predictor (TMVP) of a current block.

[0005] The image decoding method and apparatus according to the present disclosure can derive a basic motion vector of a current block, determine a corresponding position picture of a current picture containing the current block, determine a corresponding position block of the current block within the corresponding position picture, obtain motion information of the corresponding position block, derive a time motion vector predictor of the current block based on at least one of the basic motion vector or the motion information of the corresponding position block, and derive a prediction sample of the current block based on at least one of the time motion vector predictor of the current block or a TMVP (Temporal Motion Vector Predictor) reference picture.

[0006] In the image decoding method and apparatus according to the present disclosure, the position of the upper-left sample within the corresponding position block may be derived based on the basic motion vector, and the motion information of the corresponding position block may include information regarding the motion vector of the corresponding position block.

[0007] In the image decoding method and apparatus according to the present disclosure, when the resolution between the corresponding position picture and the current picture is different, the position of the upper-left sample within the corresponding position block can be derived based on a first scaling factor, and the first scaling factor can be calculated based on the ratio of the resolution between the corresponding position picture and the current picture.

[0008] In the image decoding method and apparatus according to the present disclosure, when the resolution between the corresponding position picture and the current picture is different, a picture having the same resolution as the current picture within the reference picture list of the current picture may be set as the corresponding position picture.

[0009] In the image decoding method and apparatus according to the present disclosure, when there are multiple pictures having the same resolution as the current picture within the reference picture list, the picture with the smallest absolute value of the difference in POC (Picture Order Count) with the current picture can be set as the corresponding position picture.

[0010] In the image decoding method and apparatus according to the present disclosure, when the resolution between the TMVP reference picture of the current block and the corresponding position picture is different, the time motion vector predictor of the current block may be derived based on a second scaling factor, and the second scaling factor may be calculated based on the ratio of the resolution between the TMVP reference picture of the current block and the corresponding position picture.

[0011] In the image decoding method and apparatus according to the present disclosure, the time motion vector predictor of the current block can be derived by multiplying the motion vector of the corresponding position block by the second scaling factor.

[0012] In the image decoding method and apparatus according to the present disclosure, when the resolution between the TMVP reference picture of the current block and the corresponding position picture and the resolution between the corresponding position picture and the current picture are different, the time motion vector predictor of the current block may be derived based on at least one of a first scaling factor or a second scaling factor, the first scaling factor may be calculated based on the ratio of the resolution between the corresponding position picture and the current picture, and the second scaling factor may be calculated based on the ratio of the resolution between the TMVP reference picture of the current block and the corresponding position picture.

[0013] In the image decoding method and apparatus according to the present disclosure, the time motion vector predictor of the current block may be derived by multiplying the value obtained by multiplying the basic motion vector by a first scaling factor and the sum of the motion vector of the corresponding position block by the second scaling factor.

[0014] In the image decoding method and apparatus according to the present disclosure, when the resolution between the TMVP reference picture of the current block and the corresponding position picture is different, the motion vector of the corresponding position block can be set to (0, 0).

[0015] In the image decoding method and apparatus according to the present disclosure, when the resolution between the TMVP reference picture of the current block and the corresponding position picture is different, a picture having the same resolution as the corresponding position picture within the reference picture list of the current picture can be set as the TMVP reference picture of the current block.

[0016] In the image decoding method and apparatus according to the present disclosure, when there are multiple pictures having the same resolution as the corresponding position picture within the reference picture list, the picture with the smallest absolute value of the difference in POC (Picture Order Count) with the current picture can be set as the TMVP reference picture of the current block.

[0017] In the image decoding method and apparatus according to the present disclosure, when the motion information of the corresponding position block includes a plurality of motion vector candidates, a time motion vector predictor of the current block can be derived based on motion vector candidates that reference a picture of the same resolution as the corresponding position picture.

[0018] The image encoding method and apparatus according to the present disclosure may derive a basic motion vector of a current block, determine a corresponding position picture of a current picture containing the current block, determine a corresponding position block of the current block within the corresponding position picture, obtain motion information of the corresponding position block, derive a time motion vector predictor of the current block based on at least one of the basic motion vector or the motion information of the corresponding position block, and derive a predicted sample of the current block based on at least one of the time motion vector predictor of the current block or a TMVP (Temporal Motion Vector Predictor) reference picture. Here, the position of the top-left sample within the corresponding position block may be derived based on the basic motion vector, and the motion information of the corresponding position block may include information regarding the motion vector of the corresponding position block.

[0019] A computer-readable digital storage medium is provided that stores encoded video / image information that causes an image decoding method to be performed by a decoding device according to the present disclosure.

[0020] A computer-readable digital storage medium is provided that stores video / image information generated according to the image encoding method according to the present disclosure.

[0021] A method and apparatus for transmitting video / image information generated according to the image encoding method according to the present disclosure are provided.

[0022] According to the present disclosure, when the resolution between a reference picture and a current picture is different, a more accurate prediction can be made by improving the derivation method of a time motion vector predictor, thereby improving the coding efficiency of inter-prediction.

[0023] FIG. 1 illustrates a video / image coding system according to the present disclosure.

[0024] FIG. 2 shows a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and to which encoding of a video / image signal is performed.

[0025] FIG. 3 shows a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and to which decoding of a video / image signal is performed.

[0026] FIG. 4 illustrates a decoding method performed by a decoding device (300) as an embodiment according to the present disclosure.

[0027] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs a decoding method according to the present disclosure.

[0028] FIG. 6 illustrates an encoding method performed by an encoding device (200) as an embodiment according to the present disclosure.

[0029] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs an encoding method according to the present disclosure.

[0030] FIG. 8 shows an example of a content streaming system to which embodiments of the present disclosure can be applied.

[0031] The present disclosure is susceptible to various modifications and may have various embodiments; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the present disclosure. Similar reference numerals have been used for similar components in the description of each drawing.

[0032] Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present disclosure, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.

[0033] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.

[0034] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as “comprising” or “having” are intended to specify the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0035] The present disclosure relates to video / video coding. For example, the methods / embodiments disclosed herein may be applied to methods disclosed in the VVC (versatile video coding) standard. Additionally, the methods / embodiments disclosed herein may be applied to methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVN2 (2nd generation of audio video coding standard), or next-generation video / video coding standards (e.g., H.267 or H.268).

[0036] This specification presents various embodiments regarding video / image coding, and unless otherwise noted, said embodiments may be performed in combination with one another.

[0037] In this specification, "video" may refer to a set of images over time. "Picture" generally refers to a unit representing a single image of a specific time period, and "slice" or "tile" is a unit that constitutes a part of a picture in coding. A slice or tile may contain one or more coding tree units (CTUs). A picture may consist of one or more slices or tiles. A tile is a rectangular area composed of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of ​​CTUs having a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of ​​CTUs having a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged continuously according to the CTU raster scan, whereas tiles within a picture may be arranged continuously according to the tile's raster scan. A single slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that can be exclusively contained in a single NAL unit. Meanwhile, a single picture may be divided into two or more subpictures. A subpicture may be a rectangular area of ​​one or more slices within a picture.

[0038] A pixel, or pel, can refer to the smallest unit that constitutes a picture (or image). Additionally, the term 'sample' may be used as a counterpart to pixel. A sample generally represents a pixel or its value, and it may represent only the pixel / pixel value of the luminance (luma) component or only the pixel / pixel value of the chroma component.

[0039] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of ​​a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.

[0040] In this specification, "A or B" may mean "only A," "only B," or "both A and B." Alternatively, in this specification, "A or B" may be interpreted as "A and / or B." For example, in this specification, "A, B or C" may mean "only A," "only B," "only C," or "any combination of A, B and C."

[0041] A slash ( / ) or a comma used in this specification may mean "and / or." For example, "A / B" may mean "A and / or B." Accordingly, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B or C."

[0042] In this specification, "at least one of A and B" may mean "only A," "only B," or "both A and B." Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted as synonymous with "at least one of A and B."

[0043] Additionally, in this specification, "at least one of A, B and C" may mean "only A," "only B," "only C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."

[0044] Additionally, parentheses used in this specification may mean "for example." Specifically, where indicated as "prediction (intra-prediction)," "intra-prediction" may be proposed as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be proposed as an example of "prediction." Furthermore, even where indicated as "prediction (i.e., intra-prediction)," "intra-prediction" may be proposed as an example of "prediction."

[0045] Technical features described individually within a single drawing in this specification may be implemented individually or simultaneously.

[0046] FIG. 1 illustrates a video / image coding system according to the present disclosure.

[0047] Referring to FIG. 1, the video / image coding system may include a first device (source device) (10) and a second device (receiving device) (20).

[0048] A source device (10) can transmit encoded video / image information or data in the form of a file or streaming to a receiving device (20) via a digital storage medium or network. The source device (10) may include a video source (11), an encoding device (12), and a transmission unit (13). The receiving device (20) may include a receiving unit (21), a decoding device (22), and a renderer (23). The encoding device (12) may be called a video / image encoding device, and the decoding device (22) may be called a video / image decoding device. A transmitter may be included in the encoding device. A receiver may be included in the decoding device. The renderer (23) may include a display unit, and the display unit may be composed of a separate device or an external component.

[0049] A video source (11) can acquire video / image through a process of capturing, synthesizing, or generating video / image. The video source (11) may include a video / image capture device and / or a video / image generation device. The video / image capture device may include one or more cameras, a video / image archive containing previously captured video / image, etc. The video / image generation device may include a computer, a tablet, a smartphone, etc., and can generate video / image (electronically). For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.

[0050] The encoding device (12) can encode an input video / image. The encoding device (12) can perform a series of procedures such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0051] The transmission unit (13) can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit (21) of the receiving device (20) via a digital storage medium or network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (13) may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit (21) can receive / extract the bitstream and transmit it to a decoding device (22).

[0052] The decoding device (22) can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device (12).

[0053] The renderer (23) can render the decoded video / image. The rendered video / image can be displayed through the display unit.

[0054] FIG. 2 shows a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and to which encoding of a video / image signal is performed.

[0055] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-predictor (221) and an intra-predictor (222). The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit (250) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to the embodiment. Additionally, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.

[0056] The image segmentation unit (210) can divide an input image (or picture, frame) input to an encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure.

[0057] For example, a single coding unit may be divided into multiple coding units with a deeper depth based on a quad tree structure, a binary tree structure, and / or a terrestrial structure. In this case, for example, the quad tree structure may be applied first and the binary tree structure and / or terrestrial structure may be applied later. Alternatively, the binary tree structure may be applied before the quad tree structure. A coding procedure according to the present specification may be performed based on a final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units of a lower depth so that a coding unit of the optimal size may be used as the final coding unit. Here, the term "coding procedure" may include procedures such as prediction, transformation, and restoration described below.

[0058] As another example, the processing unit may further include a Prediction Unit (PU) or a Transform Unit (TU). In this case, the Prediction Unit and the Transform Unit may each be divided or partitioned from the aforementioned final coding unit. The Prediction Unit may be a unit for sample prediction, and the Transform Unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from transformation coefficients.

[0059] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. A sample may be used to refer to a single picture (or image) as a term corresponding to a pixel or pel.

[0060] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231).

[0061] The prediction unit (220) performs a prediction for a block to be processed (hereinafter referred to as the current block) and can generate a predicted block containing prediction samples for the current block. The prediction unit (220) can determine whether intra prediction is applied or inter prediction is applied at the current block or CU level. The prediction unit (220) can generate various information regarding the prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit it to the entropy encoding unit (240). The information regarding the prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.

[0062] The intra prediction unit (222) can predict the current block by referencing samples within the current picture. The referenced samples may be located near the current block or at a certain distance from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one DC mode or a planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of fineness of the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional modes may be used. The intra prediction unit (222) may determine the prediction mode applied to the current block by using the prediction mode applied to the template area.

[0063] The inter prediction unit (221) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between the template area and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the template area may include a spatial template area (spatial neighboring block) existing within the current picture and a temporal template area (temporal neighboring block) existing in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal template area may be the same or different. The above temporal template area may be referred to by names such as collocated reference block, collocated CU (colCU), etc., and the reference picture containing the above temporal template area may be referred to as a collocated picture (colPic). For example, the inter prediction unit (221) may construct a list of motion information candidates based on the template areas and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of the template area as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the template area is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0064] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for prediction of a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called a combined inter and intra prediction (CIIP) mode. Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in screen content coding (SCC) for games. IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this specification. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values ​​within the picture can be signaled based on information regarding the palette table and palette index. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or to generate a residual signal.

[0065] The transformation unit (232) can generate transform coefficients by applying a transformation technique to a residual signal. For example, the transformation technique may include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on a prediction signal generated using all previously restored pixels. Additionally, the transformation process may be applied to a pixel block of the same size in a square, or to a block of variable size that is not square.

[0066] The quantization unit (233) quantizes the transformation coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (233) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients.

[0067] The entropy encoding unit (240) can perform various encoding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (240) may encode information required for video / image restoration (e.g., values ​​of syntax elements, etc.) together or separately, in addition to quantized transform coefficients.

[0068] Encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream at the level of a Network Abstraction Layer (NAL) unit. The video / image information may further include information regarding various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Additionally, the video / image information may further include general constraint information. In this specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream. The bitstream may be transmitted over a network or stored on a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) that transmits the signal output from the entropy encoding unit (240) and / or a storage unit (not shown) that stores it may be configured as internal / external elements of the encoding device (200), or the transmission unit may be included in the entropy encoding unit (240).

[0069] Quantized transformation coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transformation coefficients through the inverse quantization unit (234) and the inverse transformation unit (235). An adder (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter-prediction unit (221) or the intra-prediction unit (222). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (250) may be called a reconstruction unit or a reconstruction block generation unit. The generated restoration signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.

[0070] The filtering unit (260) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (260) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (270), specifically in the DPB of memory (270). The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0071] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter-prediction unit (221). Through this, when inter-prediction is applied, the encoding device can avoid prediction mismatches between the encoding device (200) and the decoding device, and can also improve encoding efficiency.

[0072] The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter-prediction unit (221). The memory (270) can store motion information of blocks from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (221) to be used as motion information in a spatial template area or motion information in a temporal template area. The memory (270) can store restoration samples of blocks restored within the current picture and transmit them to the intra-prediction unit (222).

[0073] FIG. 3 shows a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and to which decoding of a video / image signal is performed.

[0074] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-predictor (332) and an intra-predictor (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321).

[0075] The aforementioned entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) may be configured by a single hardware component (e.g., a decoding device chipset or processor) according to an embodiment. Additionally, the memory (360) may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.

[0076] When a bitstream containing video / image information is input, the decoding device (300) can restore the image in correspondence with the process in which the video / image information is processed by the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied by the encoding device. Accordingly, the processing unit for decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a binary tree structure. One or more conversion units may be derived from the coding unit. And, the restored image signal decoded and output through the decoding device (300) can be played back through a playback device.

[0077] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through an entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The decoding device can decode the picture based on information regarding the parameter sets and / or the general constraint information. The signaling / receiving information and / or syntax elements described below in this specification may be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information within the bitstream based on coding methods such as exponential chord coding, CAVLC, or CABAC, and output the values ​​of syntax elements required for image restoration and the quantized values ​​of transformation coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, determine a context model using information on the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information on symbols / bins decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Among the information decoded in the entropy decoding unit (310), information regarding prediction is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and the residual value for which entropy decoding was performed in the entropy decoding unit (310), i.e., quantized transformation coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive residual signals (residual blocks, residual samples, residual sample array). Additionally, among the information decoded in the entropy decoding unit (310), information regarding filtering can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310).

[0078] Meanwhile, the decoding device according to the present specification may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoding device (video / image / picture information decoding device) and a sample decoding device (video / image / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), inverse transform unit (322), adder (340), filtering unit (350), memory (360), inter prediction unit (332), and intra prediction unit (331).

[0079] In the inverse quantization unit (321), the quantized transformation coefficients can be inversely quantized to output transformation coefficients. The inverse quantization unit (321) can rearrange the quantized transformation coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain transformation coefficients.

[0080] In the inverse conversion unit (322), the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0081] The prediction unit (320) can perform a prediction for the current block and generate a predicted block containing prediction samples for the current block. The prediction unit (320) can determine whether an intra prediction or an inter prediction is applied to the current block based on information regarding the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter prediction mode.

[0082] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) may apply intra prediction or inter prediction for prediction of a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called a combined inter and intra prediction (CIIP) mode. Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, such as SCC (screen content coding). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this specification. The palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information regarding the palette table and palette index can be included in the above video / image information and signaled.

[0083] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or at a certain distance from the current block depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and a plurality of directional modes. The intra prediction unit (331) may determine the prediction mode applied to the current block by using the prediction mode applied to the template area.

[0084] The inter prediction unit (332) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between the template area and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the template area may include a spatial template area (spatial neighboring block) existing within the current picture and a temporal template area (temporal neighboring block) existing in the reference picture. For example, the inter prediction unit (332) may construct a motion information candidate list based on the template areas and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-prediction can be performed based on various prediction modes, and information regarding the prediction may include information indicating the inter-prediction mode for the current block.

[0085] The adder (340) can generate a restoration signal (restoration picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter prediction unit (332) and / or the intra prediction unit (331)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the prediction block can be used as the restoration block.

[0086] The addition unit (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra-predicting the next block to be processed within the current picture, may be output after filtering as described below, or may be used for inter-predicting the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.

[0087] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can generate a modified restored picture by applying various filtering methods to the restored picture, and can transmit the modified restored picture to memory (360), specifically to the DPB of memory (360). The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0088] The (modified) restored picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter prediction unit (332) to be used as motion information in a spatial template area or motion information in a temporal template area. The memory (360) can store restoration samples of blocks restored within the current picture and transmit them to the intra prediction unit (331).

[0089] In this specification, the embodiments described in the filtering unit (260), inter prediction unit (221), and intra prediction unit (222) of the encoding device (200) may be applied to the filtering unit (350), inter prediction unit (332), and intra prediction unit (331) of the decoding device (300) in the same or corresponding manner.

[0090] The temporal motion vector predictor (TMVP) of the current block can be derived based on motion information of the corresponding collocated block (hereinafter, collocated block) of the current block. The collocated block may be included in the corresponding collocated picture (hereinafter, collocated picture) of the current picture containing the current block. The collocated picture may be a picture corresponding to the basic motion vector of the current block, or a picture determined based on signaled information. The collocated block may be determined based on the basic motion vector of the current block. The basic motion vector of the current block may be derived based on at least one of the motion vector of the surrounding reference blocks of the current block or pre-defined offset information. A specific method for deriving the temporal motion vector predictor of the current block will be examined in detail with reference to FIG. 4.

[0091] FIG. 4 illustrates an image decoding method performed by a decoding device (300) as an embodiment according to the present disclosure.

[0092] Referring to Fig. 4, the base motion vector of the current block can be derived (S400). The base motion vector of the current block (baseMV) can be used to determine the call block of the current block.

[0093] For example, the basic motion vector of the current block may be the motion vector of the current block's reference block. In this case, the reference block may be a reference block adjacent to the current block, or a reference block not adjacent to the current block. A reference block adjacent to the current block may include at least one of a reference block adjacent to the top-left corner of the current block, a reference block adjacent to the top, a reference block adjacent to the top-right corner, a reference block adjacent to the left, or a reference block adjacent to the bottom-left corner.

[0094] As another example, the basic motion vector of the current block may be derived based on the motion vectors of at least two reference blocks at predefined locations relative to the current block. The basic motion vector of the current block may be the average value of the motion vectors of at least two reference blocks at predefined locations relative to the current block. Alternatively, the basic motion vector of the current block may be a weighted sum of the motion vectors of at least two reference blocks at predefined locations relative to the current block. The at least two reference blocks at predefined locations relative to the current block may be reference blocks adjacent to the current block. Alternatively, the at least two reference blocks may be reference blocks not adjacent to the current block. Alternatively, at least one of the at least two reference blocks may be a reference block adjacent to the current block, and at least one of the remaining reference blocks may be a reference block not adjacent to the current block.

[0095] As another example, the basic motion vector of the current block may be (0, 0). This may be done to reduce the computational complexity of the encoding or decoding device.

[0096] Referring to FIG. 4, a corresponding position picture of the current picture containing the current block can be determined (S410). The call block of the current block may be a block included in the corresponding position picture of the current picture (hereinafter, call picture).

[0097] For example, the call picture of the current picture can be determined on a picture sequence or slice sequence basis. In this case, the call picture of the current picture can be determined based on at least one of the call picture direction flag (ph_collocated_from_l0_flag) or the call picture index information (ph_collocated_ref_idx).

[0098] ph_collocated_from_l0_flag may be a flag indicating whether the call picture of the current picture is included in the L0 reference picture list. For example, if ph_collocated_from_l0_flag is 1, it may indicate that the call picture of the current picture is derived from the L0 reference picture list (Reference Picture List 0, RPL 0), and if ph_collocated_from_l0_flag is 0, it may indicate that the call picture of the current picture is derived from the L1 reference picture list (Reference Picture List 1, RPL 1). If the temporal motion vector predictor enable flag (ph_temporal_mvp_enabled_flag), which indicates whether the temporal motion vector predictor (TMVP) is enabled for the current picture, is 1, the reference picture list location flag (pps_rpl_info_in_ph_flag), which indicates the location where information about the reference picture list exists, is 1, and the L1 reference picture list count information (num_ref_entries[1][RplsIdx[1]]), which indicates the number of reference pictures in the L1 reference picture list, is 0, then ph_collocated_from_l0_flag can be considered as 1.

[0099] ph_collocated_ref_idx may be information representing the reference index of the call picture of the current picture. If ph_collocated_from_l0_flag is 1, ph_collocated_ref_idx may be information representing a reference picture entry within RPL 0. In this case, ph_collocated_ref_idx must be a value between 0 and num_ref_entries[0][RplsIdx[0]] - 1. num_ref_entries[0][RplsIdx[0]] may be L0 reference picture list count information representing the number of reference pictures within the L0 reference picture list. Or, if ph_collocated_from_l0_flag is 0, ph_collocated_ref_idx may be information representing a reference picture entry within RPL 1. In this case, ph_collocated_ref_idx must be a value between 0 and num_ref_entries[1][RplsIdx[1]] - 1. If ph_collocated_ref_idx does not exist, the value of ph_collocated_ref_idx may be considered 0.

[0100] As another example, the call picture of the current picture can be determined based on the basic motion vector of the current block. More specifically, the call picture of the current picture can be determined based on reference picture information containing motion information of a reference block referenced during the process of deriving the basic motion vector of the current block. In this case, if the basic motion vector of the current block is derived based on multiple reference blocks, one reference picture can be determined as the call picture of the current picture by a predefined method. For example, the call picture of the current picture can be determined based on reference picture information containing motion information of the first reference block referenced among multiple reference blocks. Alternatively, the basic motion vector of the current block can be derived based only on reference blocks containing the same reference picture information among multiple reference block candidates, and the call picture of the current picture can be determined based on said same reference picture information.

[0101] If the call picture of the current picture determined based on the basic motion vector of the current block is different from the call picture of the current picture determined in units of a picture sequence or slice sequence, the basic motion vector of the current block may be scaled to reference the call picture of the current picture determined in units of a picture sequence or slice sequence.

[0102] As another example, the call picture of the current picture may be determined as the reference picture with an index value of 0 within the L0 reference picture list of the current picture. Alternatively, the call picture of the current picture may be determined as the reference picture within the L0 reference picture list of the current picture that has the smallest absolute difference in Picture Order Count (POC) with the current picture. In this case, the L1 reference picture list may be used instead of the L0 reference picture list. For example, a reference picture list containing the reference picture with the smallest absolute difference in POC with the current picture, among the L0 reference picture list or the L1 reference picture list, may be used to determine the call picture of the current picture. Alternatively, a reference picture list according to the predicted direction of the basic motion vector of the current block, among the L0 reference picture list or the L1 reference picture list, may be used to determine the call picture of the current picture.

[0103] Referring to FIG. 4, the corresponding position block of the current block within the corresponding position picture of the current picture can be determined (S420). The call block of the current block can be determined based on at least one of the basic motion vector of the current block or the call picture of the current picture. More specifically, the call block can be determined as a block containing a sample corresponding to coordinates derived based on the basic motion vector within the call picture.

[0104] For example, a call block can be determined as a block containing a sample corresponding to a coordinate within a call picture pointed to by a base motion vector. For example, if the coordinates of the top-left sample of the current block are (CbX, CbY) and the base motion vector is (x, y), the call block can be determined as a block containing a sample at the coordinates (CbX+x, CbY+y) within the call picture. In this case, the sample at the coordinates (CbX+x, CbY+y) may be the top-left sample of the call block.

[0105] As another example, a call block may be determined as a block containing a sample corresponding to the coordinates of a position moved a predetermined distance relative to the base motion vector within the call picture. For example, if the coordinates of the top-left sample of the current block are (CbX, CbY) and the base motion vector is (x, y), the call block may be determined as a block containing a sample at the coordinates (CbX+x+m, CbY+y+n) within the call picture. In this case, the sample at the coordinates (CbX+x+m, CbY+y+n) may be the top-left sample of the call block. Here, m and n may be the width and height of the current block, respectively. Alternatively, m and n may be half the width and half the height of the current block, respectively. Alternatively, m and n may be predetermined values ​​defined identically in the encoding device and the decoding device.

[0106] The location of a call block (e.g., the location of the top-left sample within the call block) can be derived based on a basic motion vector. The call block can have the same size as the current block. Motion information can be obtained from the call block. The call block may also be divided into multiple sub-blocks. In this case, motion information can be obtained for each of the multiple sub-blocks.

[0107] The call block of the current block may be a block containing samples corresponding to coordinates derived based on basic motion vectors within the call picture as described above, and at least one of the size or position of the block may change according to the unit in which the motion vectors are stored in the picture buffer.

[0108] Referring to FIG. 4, movement information of the corresponding position block of the current block can be obtained (S430). At least one movement vector can be obtained based on the call block of the current block. At this time, movement information can be obtained for each sub-block unit within the call block.

[0109] For example, the call block of the current block may be a block containing samples of locations corresponding to coordinates derived based on basic motion vectors within the call picture, and motion information of said call block may be obtained. The motion information may include at least one of the motion vector of said call block or reference picture-related information.

[0110] As another example, a sample at a position corresponding to coordinates derived based on the basic motion vector within the call picture may be designated as the top-left sample, and a block of the same size as the current block may be determined as the call block. Motion information may be obtained for each sub-block unit within the call block. The size of the sub-block may be 1x1 to 16x16. Alternatively, the size of the sub-block may be larger than 16x16.

[0111] Referring to FIG. 4, a time motion vector predictor of the current block can be derived based on at least one of the basic motion vector of the current block or the motion information of the corresponding position block of the current block (S440).

[0112] The current block may perform predictions based on at least one of the time motion vector predictor (TMVP) or the TMVP reference picture. Here, the TMVP reference picture may mean the picture referenced when the current block uses TMVP.

[0113] The TMVP of the current block can be derived based on the motion information of the call block. For example, the motion vector of the call block can be determined as the TMVP of the current block. The TMVP reference picture of the current block can be determined based on the reference picture index included in the motion information of the call block.

[0114] As another example, if the motion information of a call block includes motion vectors in different directions, a motion vector in the same direction as the basic motion vector of the current block may be determined as the TMVP of the current block. Alternatively, all of the motion vectors in different directions may be determined as the TMVP of the current block. Or, if the absolute value of the difference in POC between the reference picture corresponding to a motion vector in one direction and the current picture is greater than a predetermined threshold, a motion vector in the other direction may be determined as the TMVP of the current block.

[0115] As another example, the TMVP of the current block can be derived by scaling the motion vector included in the motion information of the call block. If the reference picture of the call block and the TMVP reference picture of the current block differ according to the reference picture index included in the motion information of the call block, the motion vector of the call block can be scaled. More specifically, a scaling factor can be derived based on the ratio of the distance between the current picture and the TMVP reference picture of the current block and the distance between the call picture and the reference picture of the call block. The TMVP of the current block can be determined as a value obtained by scaling the motion vector of the call block based on the scaling factor. For example, the TMVP of the current block can be expressed as the product of the motion vector of the call block and the scaling factor.

[0116] The current block's time motion vector predictor (TMVP) can be derived based on the motion information of the call block and the current block's base motion vector. For example, the current block's TMVP can be derived based on the current block's base motion vector and the motion vector included in the call block's motion information. For example, the current block's TMVP can be expressed as the sum of the base motion vector and the call block's motion vector.

[0117] Referring to FIG. 4, a prediction sample of the current block can be derived based on at least one of the time motion vector predictor of the current block or the TMVP reference picture (S450). A prediction sample of the current block can be derived by performing inter-prediction based on the time motion vector predictor (TMVP) and the TMVP reference picture of the current block.

[0118] The TMVP reference picture of the current block can be determined based on the reference picture index included in the movement information of the call block. For example, the TMVP reference picture of the current block may be the same as the reference picture of the call block used to derive the TMVP of the current block.

[0119] If the reference picture of the above call block does not exist in the reference picture list of the current picture, the TMVP reference picture of the current block may be determined as a reference picture other than the reference picture of the above call block. For example, among the reference pictures in the reference picture list of the current picture, the reference picture with the smallest absolute value of the difference between the reference picture of the above call block and the POC may be determined as the TMVP reference picture of the current block. Alternatively, the TMVP reference picture of the current block may be derived based on the reference picture of the call block according to the reference picture index corresponding to a motion vector in a direction different from the motion vector of the call block used to derive the TMVP of the current block. Alternatively, if there are multiple call block candidates, the TMVP and TMVP reference picture of the current block may be derived based on other call block candidates.

[0120] The method for deriving the TMVP of the current block illustrated in FIG. 4 can be applied even when the size of the input image changes. As an example of a case where the size of the input image changes, a Reference Picture Resampling (RPR) condition may be considered. The RPR condition may be satisfied in at least one of the following cases: when the resolution between the first picture and the second picture is different, when the picture crop position between the first picture and the second picture is different, when the number of subpictures between the first picture and the second picture is different, or when the number of Coding Tree Units (CTU) between the first picture and the second picture is different. The first picture and the second picture may each be the current picture within the input image, the picture at the corresponding position of the current picture (call picture), or the TMVP reference picture of the current block. Additionally, the first picture and the second picture may be different pictures. However, in the present invention, cases where the size of the image changes are not limited to cases where the RPR condition is satisfied, and even in cases where the RPR condition is not satisfied, if the size of the image changes variably, the TMVP induction method of the current block according to the present invention may be applied.

[0121] Below, we will look at a method for inducing the TMVP of the current block when the resolution between the current picture's corresponding position picture (call picture) and the current picture is different.

[0122] For example, if the size (resolution) of the call picture is smaller than the size (resolution) of the current picture, the coordinates of the corresponding location block (call block) of the current block can be derived by performing scaling based on the ratio of the resolutions between the current picture and the call picture.

[0123] If the basic motion vectors used to derive the corresponding position block (call block) of the current block within the call picture and the coordinates of the current block are not scaled, it may not be suitable for deriving the coordinates of the call block within a call picture of a different size from the current picture. Accordingly, the coordinates of the call block within a call picture of a different size from the current picture can be derived by considering at least one of the ratio of resolution between the current picture and the call picture or the crop position.

[0124] The first scaling factor (scalingRatio1(v, u)) can be derived based on the ratio of the size of the current picture to the size of the call picture. The above v may be the value obtained by dividing the width of the call picture by the width of the current picture, and the above u may be the value obtained by dividing the height of the call picture by the height of the current picture. The coordinates ((colX, colY)) of the top-left sample of the call block can be derived as shown in the following Equation 1.

[0125]

[0126] Here, (CbX, CbY) may represent the coordinates of the top-left sample of the current block, and (x, y) may represent the basic motion vector of the current block.

[0127] The coordinates ((colX, colY)) of the top-left sample of the call block can also be derived as shown in the following mathematical equation 2.

[0128]

[0129] Due to differences in resolution between the pictures, offset information resulting from differences in coordinate standards between the pictures, or rounding values ​​during the scaling process, the coordinates of the top-left samples of the call blocks derived by Equation 1 and Equation 2 may differ from each other.

[0130] Although the above-described method was explained using the example where the call picture (resolution) is smaller than the current picture size (resolution), it can be applied even when the RPR condition is satisfied, and it is obvious that it can also be applied even when the sizes of the current picture and the call picture are different, even if the RPR condition is not satisfied.

[0131] As another example, if the resolution of the current picture and the resolution of the call picture are different, the default motion vector of the current block can be set to (0, 0). When the default motion vector is (0, 0), scaling to determine the position of the top-left sample of the call block may not be applied, so the computational complexity in the encoding or decoding device may be reduced.

[0132] As another example, if the resolution of the current picture differs from the resolution of the call picture, a picture with the same resolution as the current picture within the current picture's reference picture list can be set as the call picture. That is, a picture that does not satisfy the above RPR condition can be set as the call picture. In this case, if there are multiple pictures with the same resolution as the current picture within the current picture's reference picture list, the picture with the smallest absolute difference in POC with the current picture can be determined as the call picture. Alternatively, if there are multiple pictures with the same resolution as the current picture within the current picture's reference picture list, the picture with the smallest absolute difference in POC with an initial call picture having a different resolution from the current picture can be determined as the call picture.

[0133] As another example, if the resolution of the current picture and the resolution of the call picture are different, the time motion vector predictor for the current block may not be derived.

[0134] As another example, if there are multiple default motion vector candidates for the current block, only the default motion vector candidate referencing a picture of the same resolution as the current picture may be used. Alternatively, the default motion vector candidate referencing a picture of the same resolution as the current picture may be used preferentially. Or, only the default motion vector candidate referencing a picture of the same resolution as the pre-defined call picture may be used or preferentially used.

[0135] The above-described embodiments may be applied individually, or two or more embodiments may be combined and applied.

[0136] Below, we will look at how to derive the TMVP of the current block when the resolution between the TMVP reference picture of the current block and the current picture is different, or when the resolution between the TMVP reference picture of the current block and the call picture is different.

[0137] For example, if the resolution between the TMVP reference picture and the call picture of the current block is different, the TMVP of the current block can be derived based on the ratio of the resolutions between the TMVP reference picture and the call picture of the current block. The second scaling factor (scalingRatio2(a, b)) can be derived based on the ratio of the size of the TMVP reference picture of the current block to the size of the call picture. a may be a value obtained by dividing the width of the TMVP reference picture of the current block by the width of the call picture, and b may be a value obtained by dividing the height of the TMVP reference picture of the current block by the height of the call picture. The TMVP of the current block can be derived by scaling the motion vector of the call block within the call picture based on the second scaling factor. For example, the TMVP of the current block can be derived by multiplying the motion vector of the call block by the second scaling factor. In this case, if the resolution between the current picture and the call picture is different, the position of the top-left sample of the call block can be determined by the basic motion vector of the current block scaled based on the first scaling factor.

[0138] As another example, under at least one of the conditions where the resolution between the TMVP reference picture of the current block and the call picture is different, or where the resolution between the call picture and the current picture is different, the TMVP of the current block can be derived based on at least one of the first scaling factor or the second scaling factor. The TMVP of the current block can be derived as in Equation 3 or Equation 4.

[0139]

[0140]

[0141] Here, baseMV may be the base motion vector of the current block, and colMV may be the motion vector of the call block. Also, scalingRatio1 may be the first scaling factor, and scalingRatio2 may be the second scaling factor.

[0142] The TMVP of the current block may indicate the location of the prediction block referenced to perform inter-prediction of the current block. The TMVP of the current block may be utilized for prediction of blocks encoded or decoded after the current block, and accordingly, the TMVP of the current block may be scaled based on the ratio of resolution between the current picture and the TMVP reference picture of the current block and stored in memory. A third scaling factor (scalingRatio3(c, d)) may be derived based on the ratio of the size of the current picture to the size of the TMVP reference picture of the current block. c may be a value obtained by dividing the width of the current picture by the width of the TMVP reference picture of the current block, and d may be a value obtained by dividing the height of the current picture by the height of the TMVP reference picture of the current block. The value derived by scaling the TMVP of the current block based on the third scaling factor may be stored in memory. For example, the value obtained by multiplying the TMVP of the current block by the third scaling factor may be stored in memory.

[0143] As another example, if the resolution between the TMVP reference picture of the current block and the call picture is different, the motion vector (colMV) of the call block can be set to (0, 0). When the call motion vector is (0, 0), scaling may not be applied during the TMVP derivation process of the current block, so the computational complexity in the encoding or decoding unit may be reduced.

[0144] As another example, if the resolution between the current block's TMVP reference picture and the call picture is different, a picture within the current picture's reference picture list that has the same resolution as the call picture can be set as the current block's TMVP reference picture. That is, a picture that does not satisfy the RPR condition can be set as the current block's TMVP reference picture. In this case, if there are multiple pictures within the current picture's reference picture list that have the same resolution as the call picture, the picture with the smallest absolute difference in POC with the current picture can be set as the current block's TMVP reference picture. Alternatively, if there are multiple pictures within the current picture's reference picture list that have the same resolution as the call picture, the picture with the smallest absolute difference in POC with the initial current block's TMVP reference picture that has a resolution different from the call picture can be determined as the current block's TMVP reference picture.

[0145] As another example, if the resolution between the TMVP reference picture and the call picture of the current block is different, the time motion vector predictor of the current block may not be derived.

[0146] As another example, if there are multiple motion vector candidates for a call block, the TMVP reference picture of the current block can be determined based on motion vector candidates that reference a picture of the same resolution as the call picture, and only said motion vector candidates can be used. Alternatively, a basic motion vector candidate that references a picture of the same resolution as the call picture can be used in priority.

[0147] The motion vector candidates of the plurality of call blocks can be derived based on motion information of a specific sample location included within the call block. For example, if the location of the top-left sample of the call block is (0, 0) and the width and height of the call block are width and height, respectively, the specific sample location may include at least one of (0, 0), (width - 1, height - 1), (width >> 1, height >> 1), ((width >> 1) - 1, (height >> 1) - 1), (0, height - 1), and (width - 1, 0). Additionally, the motion vector candidates of the plurality of call blocks can be derived based on motion information of a reference block surrounding the call block. The reference block surrounding the call block may include a sample corresponding to at least one location among (width, height), (width << 1, height << 1), (0, height << 1), and (width << 1, 0). Additionally, the motion vector candidates of the plurality of call blocks may include motion vector candidates in at least one of the L0 direction or the L1 direction. Furthermore, the motion vector candidates of the plurality of call blocks may be motion vectors within an area in which the call blocks are divided into sub-block units.

[0148] The above-described embodiments may be applied individually, or two or more embodiments may be combined and applied.

[0149] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs a decoding method according to the present disclosure.

[0150] Referring to FIG. 5, the decoding device (300) may include a time motion vector predictor induction unit (500), a prediction sample induction unit (510), and a restoration unit (520). The time motion vector predictor induction unit (500) and the prediction sample induction unit (510) may be provided in the inter prediction unit (332) of FIG. 3.

[0151] The time motion vector predictor derivation unit (500) can perform a process of deriving a basic motion vector of the current block according to S400. Additionally, the time motion vector predictor derivation unit (500) can perform a process of determining a corresponding position picture of the current picture containing the current block according to S410. Additionally, the time motion vector predictor derivation unit (500) can perform a process of determining a corresponding position block of the current block within the corresponding position picture of the current picture according to S420. Additionally, the time motion vector predictor derivation unit (500) can perform a process of obtaining motion information of the corresponding position block of the current block according to S430. Additionally, the time motion vector predictor derivation unit (500) can perform a process of deriving a time motion vector predictor of the current block based on at least one of the basic motion vector of the current block or the motion information of the corresponding position block of the current block according to S440.

[0152] The prediction sample induction unit (510) can perform a process of inducing a prediction sample of the current block based on at least one of the time motion vector predictor of the current block according to S450 or the TMVP reference picture. The restoration unit (520) can perform a restoration process of the current block based on the prediction sample derived by the prediction sample induction unit (510).

[0153] FIG. 6 illustrates an encoding method performed by an encoding device (200) as an embodiment according to the present disclosure.

[0154] The basic motion vector of the current block can be derived (S600). The method for deriving the basic motion vector of the current block is as described with reference to FIG. 4.

[0155] The corresponding position picture of the current picture containing the current block can be determined (S610). The method for determining the corresponding position picture of the current picture containing the current block is as described with reference to FIG. 4.

[0156] The corresponding position block of the current block within the corresponding position picture of the current picture can be determined (S620). The method for determining the corresponding position block of the current block within the corresponding position picture of the current picture is as described with reference to FIG. 4.

[0157] Movement information of the corresponding position block of the current block can be obtained (S630). The method for obtaining movement information of the corresponding position block of the current block is as described with reference to FIG. 4.

[0158] A time motion vector predictor for the current block can be derived based on at least one of the basic motion vector of the current block or the motion information of the corresponding position block of the current block (S640). A method for deriving a time motion vector predictor for the current block based on at least one of the basic motion vector of the current block or the motion information of the corresponding position block of the current block is as described with reference to FIG. 4.

[0159] Predicted samples of the current block can be derived based on at least one of the time motion vector predictor of the current block or the TMVP reference picture (S640). The method of deriving predicted samples of the current block based on at least one of the time motion vector predictor of the current block or the TMVP reference picture is as described with reference to FIG. 4.

[0160] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs an encoding method according to the present disclosure.

[0161] Referring to FIG. 7, the encoding device (200) may include a time motion vector predictor derivation unit (700), a prediction sample derivation unit (710), a residual sample derivation unit (720), a transformation coefficient derivation unit (730), and a residual information encoding unit (740).

[0162] The time motion vector predictor derivation unit (700) and the prediction sample derivation unit (710) may be provided in the inter prediction unit (221) of FIG. 2. The residual sample derivation unit (720) and the transformation coefficient derivation unit (730) may be provided in the residual processing unit (230) of FIG. 2. The residual information encoding unit (740) may be provided in the entropy encoding unit (240).

[0163] The time motion vector predictor derivation unit (700) can perform a process of deriving a basic motion vector of the current block according to S600. Additionally, the time motion vector predictor derivation unit (700) can perform a process of determining a corresponding position picture of the current picture containing the current block according to S610. Additionally, the time motion vector predictor derivation unit (700) can perform a process of determining a corresponding position block of the current block within the corresponding position picture of the current picture according to S620. Additionally, the time motion vector predictor derivation unit (700) can perform a process of obtaining motion information of the corresponding position block of the current block according to S630. Additionally, the time motion vector predictor derivation unit (700) can perform a process of deriving a time motion vector predictor of the current block based on at least one of the basic motion vector of the current block or the motion information of the corresponding position block of the current block according to S640.

[0164] The prediction sample derivation unit (710) can perform a process of deriving a prediction sample of the current block based on at least one of the time motion vector predictor of the current block according to S650 or the TMVP reference picture. The residual sample derivation unit (720) can perform a process of deriving a residual sample based on the prediction sample derived by the prediction sample derivation unit (710). The transformation coefficient derivation unit (730) can perform a process of deriving a transformation coefficient based on the residual sample derived by the residual sample derivation unit (720). The residual information encoding unit (740) can perform a process of encoding residual information based on the transformation coefficient derived by the transformation coefficient derivation unit (730).

[0165] In the embodiments described above, methods are described based on flowcharts as a series of steps or blocks; however, the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps as described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and other steps may be included, or one or more steps of the flowcharts may be omitted without affecting the scope of the embodiments of this document.

[0166] The method according to the embodiments of the present document described above may be implemented in the form of software, and the encoding device and / or decoding device according to the present document may be included in a device that performs image processing, such as a TV, computer, smartphone, set-top box, display device, etc.

[0167] When the embodiments described in this document are implemented in software, the method described above may be implemented as a module (process, function, etc.) that performs the function described above. The module may be stored in memory and executed by a processor. The memory may be located inside or outside the processor and may be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.

[0168] In addition, the decoding device and encoding device to which the embodiment(s) of the present specification are applied may be included in multimedia broadcasting transmission and reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, Video on Demand (VoD) service providers, Over-the-top video (OTT) devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, video phone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video devices, and may be used to process video signals or data signals. For example, Over-the-top video (OTT) devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, Digital Video Recorders (DVRs), etc.

[0169] Additionally, the processing method to which the embodiment(s) of this specification are applied may be produced in the form of a program that is executed by a computer and may be stored on a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of this specification may also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Additionally, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission over the Internet). Additionally, a bitstream generated by an encoding method may be stored on a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0170] Additionally, the embodiments of this specification may be implemented as a computer program product by program code, and said program code may be executed on a computer by the embodiments of this specification. said program code may be stored on a computer-readable carrier.

[0171] FIG. 8 shows an example of a content streaming system to which embodiments of the present disclosure can be applied.

[0172] Referring to FIG. 8, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0173] The above encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server may be omitted.

[0174] The bitstream above may be generated by an encoding method or a bitstream generation method to which the embodiment(s) of the present specification are applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0175] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server plays the role of controlling commands and responses between each device within the content streaming system.

[0176] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, to provide a seamless streaming service, the streaming server may store the bitstream for a certain period of time.

[0177] Examples of the above user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.

[0178] Each server within the above-mentioned content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.

[0179] The claims described in this specification may be combined in various ways. For example, the technical features of the method claims in this specification may be combined to be implemented as a device, and the technical features of the device claims in this specification may be combined to be implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a device, and the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a method.

Claims

1. Step to derive the basic motion vector of the current block; A step of determining a corresponding position picture of the current picture containing the above current block; A step of determining the corresponding position block of the current block within the corresponding position picture; A step of obtaining movement information of the above-mentioned corresponding position block; A step of deriving a time motion vector predictor of the current block based on at least one of the above basic motion vector or the motion information of the above corresponding position block; and The method includes the step of deriving a prediction sample of the current block based on at least one of the time motion vector predictor of the current block or the TMVP (Temporal Motion Vector Predictor) reference picture. The position of the top-left sample within the above-mentioned corresponding position block is derived based on the above-mentioned basic motion vector, and An image decoding method in which the motion information of the corresponding position block includes information regarding the motion vector of the corresponding position block.

2. In Paragraph 1, When the resolution between the above-mentioned corresponding position picture and the above-mentioned current picture is different, the position of the top-left sample within the above-mentioned corresponding position block is derived based on a first scaling coefficient, and An image decoding method in which the first scaling factor is calculated based on the ratio of the resolution between the corresponding position picture and the current picture.

3. In Paragraph 1, A video decoding method that, when the resolution between the corresponding position picture and the current picture is different, sets a picture having the same resolution as the current picture within the reference picture list of the current picture as the corresponding position picture.

4. In Paragraph 3, A video decoding method in which, when there are multiple pictures having the same resolution as the current picture within the reference picture list, the picture with the smallest absolute value of the difference in POC (Picture Order Count) with the current picture is set as the corresponding position picture.

5. In Paragraph 1, If the resolution between the TMVP reference picture of the current block and the corresponding position picture is different, the time motion vector predictor of the current block is derived based on a second scaling factor, and An image decoding method in which the second scaling factor is calculated based on the ratio of the resolution between the TMVP reference picture of the current block and the corresponding position picture.

6. In Paragraph 5, An image decoding method in which the time motion vector predictor of the current block is derived by multiplying the motion vector of the corresponding position block by the second scaling factor.

7. In Paragraph 1, When the resolution between the TMVP reference picture of the current block and the corresponding position picture and the resolution between the corresponding position picture and the current picture are different, the time motion vector predictor of the current block is derived based on at least one of a first scaling factor or a second scaling factor, and The first scaling factor is calculated based on the ratio of the resolution between the corresponding position picture and the current picture, and An image decoding method in which the second scaling factor is calculated based on the ratio of the resolution between the TMVP reference picture of the current block and the corresponding position picture.

8. In Paragraph 7, An image decoding method in which the time motion vector predictor of the current block is derived by multiplying the value obtained by multiplying the basic motion vector by a first scaling factor and the sum of the motion vector of the corresponding position block by the second scaling factor.

9. In Paragraph 1, A video decoding method that sets the motion vector of the corresponding position block to (0, 0) when the resolution between the TMVP reference picture of the current block and the corresponding position picture is different.

10. In Paragraph 1, A video decoding method that, when the resolution between the TMVP reference picture of the current block and the corresponding position picture is different, sets a picture having the same resolution as the corresponding position picture within the reference picture list of the current picture as the TMVP reference picture of the current block.

11. In Paragraph 10, A video decoding method that, when there are multiple pictures having the same resolution as the corresponding position picture within the reference picture list, sets the picture with the smallest absolute value of the difference in POC (Picture Order Count) with the current picture as the TMVP reference picture of the current block.

12. In Paragraph 1, A video decoding method that, when the motion information of the corresponding position block includes a plurality of motion vector candidates, derives a time motion vector predictor of the current block based on motion vector candidates that reference a picture of the same resolution as the corresponding position picture.

13. Step to derive the basic motion vector of the current block; A step of determining a corresponding position picture of the current picture containing the above current block; A step of determining the corresponding position block of the current block within the corresponding position picture; A step of obtaining movement information of the above-mentioned corresponding position block; A step of deriving a time motion vector predictor of the current block based on at least one of the above basic motion vector or the motion information of the above corresponding position block; and The method includes the step of deriving a prediction sample of the current block based on at least one of the time motion vector predictor of the current block or the TMVP (Temporal Motion Vector Predictor) reference picture. The position of the top-left sample within the above-mentioned corresponding position block is derived based on the above-mentioned basic motion vector, and An image encoding method wherein the motion information of the corresponding position block includes information regarding the motion vector of the corresponding position block.

14. A computer-readable storage medium for storing a bitstream generated by the method according to paragraph 13.

15. A step of acquiring a bitstream for image information; wherein the bitstream is generated based on the steps of deriving a basic motion vector of a current block, determining a corresponding position picture of a current picture containing the current block, determining a corresponding position block of the current block within the corresponding position picture, acquiring motion information of the corresponding position block, deriving a time motion vector predictor of the current block based on at least one of the basic motion vector or the motion information of the corresponding position block, and deriving a prediction sample of the current block based on at least one of the time motion vector predictor of the current block or a TMVP (Temporal Motion Vector Predictor) reference picture, and The method includes the step of transmitting data including the above bitstream, The position of the top-left sample within the above-mentioned corresponding position block is derived based on the above-mentioned basic motion vector, and A method in which the movement information of the corresponding position block includes information regarding the movement vector of the corresponding position block.