Video Encoding / Decoding Method and Apparatus, and Recording Medium in which Bitstream is Stored

The SbTMVP method addresses inefficiencies in existing video compression by utilizing a diverse set of candidates for motion vector prediction, enhancing accuracy and efficiency in high-resolution video encoding and decoding.

JP2025522749APending Publication Date: 2025-07-17LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024575484
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-01
Filing Date
2023-07-03
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing video compression technologies face challenges in achieving high accuracy and efficiency, particularly in handling high-resolution and high-quality videos, due to limitations in inter prediction methods.

Method used

The method and apparatus utilize a subblock-based temporal motion vector predictor (SbTMVP) that considers a diverse set of candidates, including adjacent and non-adjacent spatial neighboring blocks, to derive motion vectors in sub-block units, using template matching to select the optimal predictor.

Benefits of technology

This approach enhances prediction accuracy and compression performance by considering more various candidates, improving the efficiency of video encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522749000001_ABST
    Figure 2025522749000001_ABST
Patent Text Reader

Abstract

The video decoding / encoding method and apparatus according to the present disclosure can derive a temporal vector of a current block based on a motion vector of one candidate among a group of candidates including a plurality of candidates, determine a collocated block of the current block in a collocated picture based on the temporal vector, derive a motion vector of the current block in sub-block units based on the motion vector of the collocated block, and perform an inter prediction on the current block based on the motion vector of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video encoding / decoding method and apparatus, and a recording medium storing a bitstream.

Background Art

[0002] In recent years, the demand for high-resolution and high-quality videos such as HD (High Definition) videos and UHD (Ultra High Definition) videos has been increasing in various application fields, and thus highly efficient video compression technologies have been discussed.

[0003] As video compression technologies, there are various technologies such as an inter prediction technology that predicts pixel values included in a current picture from pictures before or after the current picture, an intra prediction technology that predicts pixel values included in the current picture using pixel information within the current picture, and an entropy coding technology that assigns short codes to frequently occurring values and long codes to less frequently occurring values. Using such video compression technologies, video data can be effectively compressed and transmitted or stored.

Summary of the Invention

Problems to be Solved by the Invention

[0004] The present disclosure intends to provide a method and apparatus for performing inter prediction using a subblock-based temporal motion vector predictor (SbTMVP).

[0005] The present disclosure intends to provide a method and apparatus for considering more various candidates when deriving a subblock-based temporal motion vector predictor.

Means for Solving the Problems

[0006] The video decoding method and apparatus according to the present disclosure can derive a temporal vector of a current block based on a motion vector of one candidate among a group of candidates including a plurality of candidates, determine a collocated block of the current block in a collocated picture based on the temporal vector, derive a motion vector of the current block in sub-block units based on a motion vector of the collocated block, and perform an inter prediction on the current block based on the motion vector of the current block.

[0007] In the video decoding method and apparatus according to the present disclosure, the group of candidates may include adjacent spatial neighboring blocks of the current block and non-adjacent spatial neighboring blocks of the current block as candidates.

[0008] In the video decoding method and apparatus according to the present disclosure, the non-adjacent spatial neighboring blocks may include a first block including samples that are separated by a value obtained by multiplying the height of the current block by 2 from a left upper end sample adjacent to the upper left corner of the current block in a left sample line adjacent to the current block, and a second block including samples that are separated by a value obtained by multiplying the width of the current block by 2 from the left upper end sample in an upper sample line adjacent to the current block.

[0009] In the video decoding method and apparatus according to the present disclosure, the one candidate may be selected from the plurality of candidates included in the group of candidates based on template matching.

[0010] In the video decoding method and apparatus according to the present disclosure, the cost of the template matching may be calculated based on the sum of absolute differences (SAD) or the mean-removed sum of absolute differences (MRSAD) between the template region of the current block and the template region of the block specified by the motion vectors of the plurality of candidates included in the candidate group.

[0011] In the video decoding method and apparatus according to the present disclosure, the cost of the template matching may be calculated based on the SAD (sum of absolute differences) or MRSAD (mean-removed sum of absolute differences) between the template region of the current block and the template region of the block or sub-block specified by the motion vector of the block induced based on the plurality of candidates included in the candidate group.

[0012] The video decoding method and apparatus according to the present disclosure can configure the candidate group including the plurality of candidates based on the surrounding blocks of the current block.

[0013] In the video decoding method and apparatus according to the present disclosure, the candidate group may be configured by adding the surrounding blocks at specific positions to the candidate group in a predefined order.

[0014] The video decoding method and apparatus according to the present disclosure can check whether the motion vector of the surrounding block at the specific position overlaps with the motion vectors of the candidates previously included in the candidate group.

[0015] In the video decoding method and apparatus according to the present disclosure, whether there is an overlap may be determined based on whether the difference between the motion vector of the candidate previously included in the candidate group and the motion vector of the surrounding block at the specific position is smaller than a predefined threshold.

[0016] In the video decoding method and apparatus according to the present disclosure, the collocated picture may be determined based on a syntax element signaled by at least one syntax of a picture header or a slice header.

[0017] In the video decoding method and apparatus according to the present disclosure, the syntax element may be signaled separately from a syntax element indicating a collocated picture for a temporal motion vector predictor.

[0018] The video encoding method and apparatus according to the present disclosure determine a temporal vector of a current block based on a motion vector of a candidate among a group of candidates including a plurality of candidates, determine a collocated block of the current block in a collocated picture based on the temporal vector, determine the motion vector of the current block in sub-block units based on the motion vector of the collocated block, and can perform inter prediction on the current block based on the motion vector of the current block.

[0019] In the video encoding method and apparatus according to the present disclosure, the group of candidates may include adjacent spatial neighboring blocks of the current block and non-adjacent spatial neighboring blocks of the current block as candidates.

[0020] In the video encoding method and apparatus according to the present disclosure, the non-adjacent spatial neighboring block may include a first block including samples that are separated from the upper left sample adjacent to the upper left corner of the current block by a value obtained by multiplying the height of the current block by 2 in the left sample line adjacent to the current block, and a second block including samples that are separated from the upper left sample by a value obtained by multiplying the width of the current block by 2 in the upper sample line adjacent to the current block.

[0021] In the video encoding method and apparatus according to the present disclosure, the one candidate may be selected from the plurality of candidates included in the candidate group based on template matching.

[0022] In the video encoding method and apparatus according to the present disclosure, the cost of the template matching may be calculated based on the sum of absolute differences (SAD) or the mean-removed sum of absolute differences (MRSAD) between the template area of the current block and the template area of the block specified by the motion vectors of the plurality of candidates included in the candidate group.

[0023] In the video encoding method and apparatus according to the present disclosure, the cost of the template matching may be calculated based on the SAD or MRSAD between the template area of the current block and the template area of the block or sub-block specified by the motion vector of the block induced based on the plurality of candidates included in the candidate group.

[0024] The video encoding method and apparatus according to the present disclosure can constitute the candidate group including the plurality of candidates based on the peripheral blocks of the current block.

[0025] In the video encoding method and apparatus according to the present disclosure, the candidate group may be configured by adding peripheral blocks at a specific position to the candidate group in a predefined order.

[0026] The video encoding method and apparatus according to the present disclosure can check whether the motion vector of the peripheral block at the specific position overlaps with the motion vectors of the candidates previously included in the candidate group.

[0027] In the video encoding method and apparatus according to the present disclosure, whether there is an overlap may be determined based on whether the difference between the motion vector of the candidate previously included in the candidate group and the motion vector of the peripheral block at the specific position is smaller than a predefined threshold.

[0028] In the video encoding method and apparatus according to the present disclosure, the collocated picture may be determined based on a syntax element signaled by at least one syntax of a picture header or a slice header.

[0029] In the video encoding method and apparatus according to the present disclosure, the syntax element may be signaled separately from the syntax element indicating the collocated picture for a temporal motion vector predictor.

[0030] There is provided a computer-readable digital storage medium storing encoded video / video information for causing a decoding apparatus according to the present disclosure to perform a video decoding method.

[0031] There is provided a computer-readable digital storage medium storing video / video information generated by the video encoding method according to the present disclosure.

[0032] A method and apparatus for transmitting video / video information generated by a video encoding method according to the present disclosure are provided.

Advantages of the Invention

[0033] In the present disclosure, when inducing a sub-block-based temporal motion vector predictor, by considering more various candidates in prediction, the accuracy of prediction can be improved and the compression performance can be enhanced.

Brief Description of the Drawings

[0034]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

[0035] The present disclosure is capable of various modifications and can have various embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood that the present disclosure includes all modifications, equivalents, or alternatives included in the spirit and technical scope of the present disclosure. In the description of each figure, similar reference numerals are used for similar components.

[0036] Terms such as first, second, etc. can be used to describe various components, but the components should not be limited by these terms. These terms are only used for the purpose of distinguishing one component from another. For example, without departing from the scope of the present disclosure, the first component may be named the second component, and similarly, the second component may also be named the first component. The term "and / or" includes any combination of a plurality of related listed items or any one of a plurality of related listed items.

[0037] When a component is referred to as being "coupled" or "connected" to another component, it should be understood that it may be directly coupled or connected to the other component, or there may be another component in between. On the other hand, when a component is referred to as being "directly coupled" or "directly connected" to another component, it should be understood that there is no other component in between.

[0038] The terms used in this application are only for the purpose of explaining specific embodiments and are not intended to limit the present disclosure. Singular expressions include plural expressions as well, unless otherwise specified in the context. In this application, terms such as "comprising" or "having" are used to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should not be construed as precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0039] This disclosure relates to video / video coding. For example, the methods / examples disclosed in this specification may be applied to the methods disclosed in the VVC (versatile video coding) standard. Also, the methods / examples disclosed in this specification may be applied to the methods disclosed in the EVC (essential video coding) standard, the AV1 (AOMedia Video1) standard, the AVS2 (2nd generation of audio video coding standard), or next-generation video / video coding standards (e.g., H.267 or H.268, etc.).

[0040] This specification presents various examples related to video / video coding, and unless otherwise specifically mentioned, these examples may be carried out in combination with each other.

[0041] In this specification, "video" can mean a collection of a series of images over time. "Picture" generally means a unit representing one image in a specific time period, and "slice" / "tile" is a unit that constitutes a part of a picture in coding. A slice / tile may include one or more CTUs (Coding Tree Units). One picture may be composed of one or more slices / tiles. One tile is a rectangular area composed of a plurality of CTUs within a specific tile column and a specific tile row of one picture. A tile column is a rectangular area of CTUs having the same height as the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs having a height specified by the picture parameter set and the same width as the width of the picture. The CTUs within one tile are arranged continuously by a CTU raster scan, but the tiles within one picture may be arranged continuously by a tile raster scan. One slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that can be exclusively included in a single NAL unit. On the other hand, one picture may be divided into two or more sub-pictures. A sub-picture may be a rectangular area of one or more slices within a picture.

[0042] A pixel, pel, or sample can mean the smallest unit that constitutes one picture (or video). Also, the term "sample" may be used as a term corresponding to a pixel. A sample can generally represent a pixel or the value of a pixel, and may represent only the pixel / pixel value of the luma component, or may represent only the pixel / pixel value of the chroma component.

[0043] A unit can represent the basic unit of video processing. A unit may include at least one of a specific area of a picture and information related to the area. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. A unit may, in some cases, be used with the same meaning as terms such as "block" or "area". In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.

[0044] In this specification, "A or B" can mean "only A", "only B", or "both A and B". In other words, "A or B" in this specification can be interpreted as "A and / or B". For example, "A, B or C" in this specification can mean "only A", "only B", "only C", or "any combination of A, B and C".

[0045] The slashes ( / ) and commas used in this specification can mean "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B or C".

[0046] In this specification, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, in this specification, expressions such as "at least one of A or B" and "at least one of A and / or B" may be interpreted identically to "at least one of A and B".

[0047] Also, in this specification, "at least one of A, B and C" can mean "only A", "only B", "only C", or "any combination of A, B and C". Also, "at least one of A, B or C" and "at least one of A, B and / or C" can mean "at least one of A, B and C".

[0048] Also, the parentheses used in this specification can mean "for example". Specifically, when it is displayed as "prediction (intra prediction)", "intra prediction" may be proposed as an example of "prediction". In other words, "prediction" in this specification is not limited to "intra prediction", and "intra prediction" may be proposed as an example of "prediction". Also, when it is displayed as "prediction (that is, intra prediction)", "intra prediction" may be proposed as an example of "prediction".

[0049] In this specification, the technical features separately described in one drawing may be embodied separately or simultaneously.

[0050] FIG. 1 is a diagram showing a video / video coding system according to the present disclosure.

[0051] Referring to FIG. 1, the video / video coding system may include a first device (source device) and a second device (receiving device).

[0052] The source device can transmit encoded video / image information or data in the form of a file or a stream to the receiving device via a digital storage medium or a network. The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitting unit may be included in the encoding device. The receiving unit may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component.

[0053] The video source can obtain video / images through processes such as video / image capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images may be generated by a computer or the like, and in this case, the video / image capture process may be replaced by the process of generating related data.

[0054] The encoding device can encode the input video / image. The encoding device can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) may be output in the form of a bitstream.

[0055] The transmitting unit can transmit the encoded video / video information or data output in the form of a bitstream to the receiving unit of the receiving device in the form of a file or streaming via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit may include elements for generating a media file according to a predefined file format and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.

[0056] The decoding device can decode the video / video by performing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.

[0057] The renderer can render the decoded video / video. The rendered video / video may be displayed on the display unit.

[0058] FIG. 2 is a block diagram schematically showing an encoding device to which the application of the embodiment of the present disclosure is applicable and where video / video signal encoding is performed.

[0059] Referring to FIG. 2, the encoding device 200 may include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor 220 may include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 may include a transformer (232), a quantizer 233, a dequantizer (234), and an inverse transformer (235). The residual processor 230 may further include a subtractor (231). The adder 250 may also be referred to as a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by one or more hardware components (e.g., an encoding device chipset or a processor) according to an embodiment. Also, the memory 270 may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory 270 as an internal / external component.

[0060] The video segmentation unit 210 can divide the input video (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be referred to as a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure.

[0061] For example, one coding unit may be divided into a plurality of coding units with deeper depths based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. In this case, for example, the quad-tree structure may be applied first, and the binary-tree structure and / or the ternary-tree structure may be applied later. Or, the binary-tree structure may be applied earlier than the quad-tree structure. Based on the final coding unit that is no longer divided, the coding procedure according to this specification may be performed. In this case, based on the coding efficiency according to the video characteristics, etc., the largest coding unit may be immediately used as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units with deeper depths, and the coding unit having an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described later.

[0062] As another example, the processing unit may further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit may each be divided or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit of deriving transform coefficients and / or a unit of deriving a residual signal from the transform coefficients.

[0063] The unit may, in some cases, be used in the same sense as terms such as a block or an area. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can also represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A sample may be used as a term corresponding to one picture (or video) for a pixel or a pel.

[0064] The encoding device 200 can subtract a prediction signal (prediction block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input video signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, in the encoding device 200, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input video signal (original block, original sample array) may be called the subtraction unit 231.

[0065] The prediction unit 220 can perform a prediction on a processing target block (hereinafter referred to as the current block), and generate a predicted block that includes a prediction sample for the current block. The prediction unit 220 can determine whether intra prediction or inter prediction is applied in units of the current block or CU. As will be described later in the description of each prediction mode, the prediction unit 220 can generate various pieces of information related to prediction, such as prediction mode information, and transmit them to the entropy encoding unit 240. The information related to prediction may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0066] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to may be located in the neighborhood of the current block or at a certain distance from the current block depending on the prediction mode. In intra prediction, the prediction mode may include one or more non-directional modes and a plurality of directional modes. The non-directional mode may include at least one of the DC mode or the Planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of fineness of the prediction direction. However, this is an example, and a greater or smaller number of directional modes may be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.

[0067] The inter prediction unit 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (such as L0 prediction, L1 prediction, Bi prediction, etc.). In inter prediction, the neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block can also be called a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes. For example, in the skip mode and the merge mode, the inter prediction unit 221 can use the motion information of neighboring blocks as the motion information of the current block. In the skip mode, different from the merge mode, the residual signal does not need to be transmitted. In the motion vector prediction (MVP) mode, the motion vector of a neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling a motion vector difference.

[0068] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply intra prediction or inter prediction for predicting one block, and can also apply intra prediction and inter prediction simultaneously. This can be called CIIP (combined inter and intra prediction). Further, the prediction unit can also proceed based on the intra block copy (IBC) prediction mode or the palette mode for predicting a block. The IBC prediction mode or the palette mode may be used for coding content video / motion video such as games like SCC (screen content coding). IBC basically performs prediction within the current picture, but may be performed similarly to inter prediction in that it induces a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this specification. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information regarding the palette table and the palette index. The prediction signal generated by the prediction unit 220 may be used to generate a restored signal or may be used to generate a residual signal.

[0069] The conversion unit 232 can generate transform coefficients by applying a conversion method to the residual signal. For example, the conversion method may include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, when the GBT represents the relationship information between pixels as a graph, it means the conversion obtained from this graph. The CNT means generating a prediction signal using all previously restored pixels and the conversion obtained based on that. Also, the conversion process may be applied to pixel blocks having the same size of a square, or may be applied to blocks of variable sizes other than squares.

[0070] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240, and the entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be called residual information. The quantization unit 233 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.

[0071] The entropy encoding unit 240 can perform various encoding methods such as exponential Golomb, CAVLC (context - adaptive variable length coding), CABAC (context - adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements) in addition to the quantized transform coefficients.

[0072] The encoded information (e.g., encoded video / video information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / video information may further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information may further include general constraint information. In this specification, the information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / video information. The video / video information may be encoded by the above - described encoding procedure and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu - ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 240 may be configured as internal / external elements of the encoding device 200, or the transmission unit may be included in the entropy encoding unit 240.

[0073] The quantized transform coefficients output from the quantization unit 233 may be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients using the inverse quantization unit 234 and the inverse transform unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 250 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block may be used as the reconstructed block. The addition unit 250 may be referred to as a restoration unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next block to be processed within the current picture, and as will be described later, may also be used for inter prediction of the next picture after passing through filtering. On the other hand, LMCS (luma mapping with chroma scaling) may be applied in the picture encoding and / or restoration process.

[0074] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240. The information related to filtering may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0075] The modified restored picture transmitted to the memory 270 may be used as a reference picture in the inter prediction unit 221. By doing so, when inter prediction is applied, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and can also improve the encoding efficiency.

[0076] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the block in which the current picture intra motion information is derived (or encoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.

[0077] FIG. 3 is a block diagram schematically showing a decoding apparatus to which an embodiment of the present disclosure is applicable and in which decoding of a video / video signal is performed.

[0078] Referring to FIG. 3, the decoding apparatus 300 may include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filtering unit (filter, 350), and a memory (memory, 360). The predictor 330 may include an inter-prediction unit 332 and an intra-prediction unit 331. The residual processor 320 may include an inverse quantization unit (dequantizer, 321) and an inverse transform unit (inverse transformer, 321).

[0079] The entropy decoding unit 310, the residual processing unit 320, the prediction unit 330, the addition unit 340, and the filtering unit 350 described above may be configured by one hardware component (for example, a decoding chipset or a processor) according to an embodiment. The memory 360 may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.

[0080] When a bitstream including video / video information is input, the decoding device 300 can restore the video corresponding to the process in which the video / video information is processed by the encoding device in FIG. 2. For example, the decoding device 300 can derive units / blocks based on the block division related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing unit applied in the encoding device. Therefore, the processing unit for decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit having a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more conversion units may be derived from the coding unit. Then, the restored video signal decoded and output by the decoding device 300 may be played back by a playback device.

[0081] The decoding device 300 can receive the signal output from the encoding device in FIG. 2 in the form of a bitstream, and the received signal may be decoded by the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream and derive information (e.g., video / video information) necessary for video restoration (or, picture restoration). The video / video information may further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information may further include general constraint information. The decoding device can decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this specification may be decoded by the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the conversion coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element from the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding and the block to be decoded, or the information of the symbol / bin decoded in the previous stage, predicts the occurrence probability of the bin by the determined context model, performs arithmetic decoding of the bin, and can generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded for the context model of the next symbol / bin after determining the context model.Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit 330 (inter prediction unit 332 and intra prediction unit 331), and the residual value obtained by performing entropy decoding in the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information, may be input to the residual processing unit 320. The residual processing unit 320 can induce a residual signal (residual block, residual sample, residual sample array). Also, the information related to filtering among the information decoded by the entropy decoding unit 310 may be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310.

[0082] On the other hand, the decoding device according to the present specification can be called a video / video / picture decoding device, and the decoding device can be distinguished into an information decoder (video / video / picture information decoder) and a sample decoding device (video / video / picture sample decoding device). The information decoding device may include the entropy decoding unit 310, and the sample decoding device may include at least one of the inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.

[0083] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output the transform coefficients. The inverse quantization unit 321 can reorder the quantized transform coefficients in the form of a two-dimensional block. In this case, the reordering can be performed based on the coefficient scan order performed in the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using quantization parameters (for example, quantization step size information) to obtain the transform coefficients.

[0084] In the inverse transform unit 322, the transform coefficients are inverse transformed to obtain a residual signal (residual block, residual sample array).

[0085] The prediction unit 320 can perform a prediction on the current block and generate a predicted block including the predicted samples for the current block. The prediction unit 320 can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.

[0086] The prediction unit 320 can generate a prediction signal based on various prediction methods described below. For example, the prediction unit 320 can apply intra prediction or inter prediction for the prediction of one block, and can also apply intra prediction and inter prediction simultaneously. This can be called the CIIP (combined inter and intra prediction) mode. Also, the prediction unit may be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode may be used for the coding of content video / moving video such as games like SCC (screen content coding). IBC basically performs prediction within the current picture, but may be performed similarly to inter prediction in terms of inducing a reference block within the current picture. That is, IBC can use at least one of the inter prediction methods described in this specification. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index may be included in and signaled in the video / video information.

[0087] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples referred to may be located in the neighborhood of the current block or at a certain distance from the current block depending on the prediction mode. In intra prediction, the prediction mode may include one or more non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.

[0088] The inter prediction unit 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (such as L0 prediction, L1 prediction, Bi prediction, etc.). In inter prediction, the peripheral blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 332 can configure a motion information candidate list based on the peripheral blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction may be performed based on various prediction modes, and the information regarding the prediction may include information indicating the inter prediction mode for the current block.

[0089] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the block to be processed, as in the case where the skip mode is applied, the prediction block may be used as the restored block.

[0090] The adder 340 can be referred to as a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, and as will be described later, it may be output after filtering, or may be used for inter prediction of the next picture. On the other hand, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.

[0091] The filtering unit 350 can apply filtering to the restoration signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically to the DPB of the memory 360. The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like.

[0092] The (modified) restored picture stored in the DPB of the memory 360 may be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the block where the motion information within the current picture is derived (or decoded), and / or the motion information of the block within the already restored picture. The stored motion information can be transmitted to the inter prediction unit 332 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 360 can store the restored samples of the restored blocks within the current picture and can transmit them to the intra prediction unit 331.

[0093] In this specification, the examples described with respect to the filtering unit 260, the inter prediction unit 221, and the intra prediction unit 222 of the encoding device 200 may be applied identically or correspondingly to the filtering unit 350, the inter prediction unit 332, and the intra prediction unit 331 of the decoding device 300, respectively.

[0094] FIG. 4 is a diagram showing an inter prediction method performed by a decoding device 300 according to an embodiment of the present disclosure.

[0095] In an embodiment of the present disclosure, in performing inter prediction using a subblock-based temporal motion vector predictor (SbTMVP), a method of improving the prediction accuracy by considering more various candidates is proposed. In the present disclosure, the subblock-based temporal motion vector predictor represents a predictor induced based on motion information of a temporal neighboring block in subblock units, and it goes without saying that the name is not limited thereto. In the present disclosure, the subblock-based temporal motion vector predictor may be abbreviated as SbTMVP for convenience of explanation.

[0096] Referring to FIG. 4, the decoding device can derive a temporal vector of a current block based on a motion vector of one candidate among a group of candidates including a plurality of candidates (S400).

[0097] Unlike the TMVP of the sub-block merge mode (SbTMVP), which derives the motion information of a block at a predefined position (i.e., the lower right corner position or the center position of the current block) within a predefined reference picture (or a collocated picture, colPic) as the motion vector predictor (MVP), SbTMVP induces the MVP on a sub-block basis by using the motion information of a collocated block (or a col block) derived from the motion information of an adjacent block (e.g., the left adjacent block of the current block) as the MVP of the current block. In an embodiment of the present disclosure, the higher the accuracy of the motion information of the adjacent block and the more various candidates are considered, the more the performance of SbTMVP can be improved. Therefore, a method therefor is proposed.

[0098] In other words, a method for improving the SbTMVP induction process considering various temporal vectors for identifying collocated blocks is proposed. The proposed method may be applied in substantially the same way not only to the sub-block based MVP induction process but also to the MVP induction process in the MERGE / AMVP process.

[0099] According to an embodiment of the present disclosure, in order to improve the performance of the existing SbTMVP that uses the motion information of a collocated block derived from the motion information of the left adjacent block as the MVP, an optimal MVP can be induced by considering various motion information derived from a number of adjacent blocks. As an example, the candidate group may include adjacent spatial adjacent blocks of the current block and non-adjacent spatial adjacent blocks of the current block as candidates. Details will be described below with reference to the drawings.

[0100] The decoding device can determine (or identify) the collocated block of the current block within the collocated picture based on the induced time vector (S410). As an example, the collocated picture may be defined identically for the encoding device and the decoding device. Or, as an example, the collocated picture may be signaled by a higher-level syntax. The collocated block of the current block may be determined by the time vector within the collocated picture.

[0101] The decoding device can induce the motion vector of the current block in sub-block units based on the motion vector of the collocated block (S420). That is, the decoding device can induce (or determine) the motion vector of the sub-block within the collocated block specified by the time vector as the motion vector of the corresponding sub-block within the current block.

[0102] The decoding device can perform an inter prediction for the current block based on the motion vector of the current block (S430). The decoding device can generate a predicted block of the current block by performing an inter prediction based on the motion information induced in sub-block units within the current block. As an example, the predicted block of the current block may be generated in sub-block units by the motion information induced in sub-block units.

[0103] FIG. 5 is a flowchart showing a process of inducing a sub-block-based motion vector predictor to which the embodiment of the present disclosure is applicable.

[0104] In FIG. 5, it is assumed that only the motion vector of the left adjacent block (that is, the block adjacent to the left of the left lower end sample of the current block, A1) is used as the time vector as the conventional SbTMVP.

[0105] Referring to FIG. 5, when SbTMVP is used, the collocated block may be induced using the motion information of the left adjacent block as a time vector. First, it may be confirmed whether the motion information of the left adjacent block is available. If the motion information of the left adjacent block is available, the time vector may be set with the motion information of the left adjacent block as the time vector. If the motion information of the left adjacent block is not available, the time vector may be set to a zero vector.

[0106] In the collocated picture, the collocated block may be specified by the time vector. If the default motion vector of the central position of the collocated block is available, the motion information for each sub-block unit can be saved as MVP. That is, if the default motion vector is available, the motion information for each sub-block unit may be induced. The availability of the default motion vector may be a condition for inducing the motion information for each sub-block unit.

[0107] The motion information of sub-blocks that are not available at the stage of inducing the motion information for each sub-block unit may be replaced with the default motion vector. As an example, the collocated block may be in inter mode and may be determined to be available when not in intra block copy (IBC) mode.

[0108] FIG. 6 is a flowchart showing the process of inducing a sub-block-based motion vector predictor according to an embodiment of the present disclosure.

[0109] Referring to FIG. 6, in order to consider various motion information according to the embodiment of the present disclosure, the SbTMVP process described in FIG. 5 above may be applied with modifications as shown in FIG. 6.

[0110] According to the embodiment of the present disclosure, a time vector can be induced using adjacent blocks at various predefined positions. That is, adjacent blocks at various predefined positions can be used as SbTMVP candidates.

[0111] As an example, a time vector may be derived using a plurality of neighboring blocks {A1, B1, B0, A0, B2, D0, D1}. The time vector may be derived using the motion vector of one candidate specified from a group of candidates including a plurality of candidates. That is, the group of candidates may be {A1, B1, B0, A0, B2, D0, D1} including a plurality of neighboring blocks.

[0112] Referring to FIG. 6, when the motion vector of the center position of the collocated block (i.e., the default motion vector) derived from the time vector in order within the group of candidates is available, it may be considered as a candidate for SbTMVP. The time vector derivation process is as shown in FIG. 6 or as described in FIG. 5 above, and related overlapping descriptions are omitted. This process may be repeated until MAX_NUM candidates are available or until all candidates have been checked. Here, MAX_NUM may represent the maximum number of SbTMVP candidates. MAX_NUM may be a predefined value. Alternatively, MAX_NUM may be a value explicitly signaled or a value implicitly derived based on coded information.

[0113] If all checks have been performed on all specified neighboring blocks but less than MAX_NUM candidates are satisfied, a zero vector may be considered as the time vector. Thereafter, a candidate having the lowest cost may be determined based on template matching (or template matching cost) for the SbTMVP candidates obtained from the above process. The determined candidate can be used to finally derive the MVP for each sub-block unit.

[0114] In the present disclosure, the group of candidates used for time vector derivation may represent a group of candidates including a plurality of predefined neighboring blocks as candidates. Alternatively, as in the embodiment shown in FIG. 6, the group of candidates used for time vector derivation may represent a group of candidates including candidates equal to or less than the maximum number (MAX_NUM) for which availability has been confirmed among a plurality of predefined neighboring blocks.

[0115] FIG. 7 is a diagram illustrating peripheral blocks to which an embodiment of the present disclosure is applicable.

[0116] FIG. 7 shows various reference blocks that can be considered for time vector derivation. That is, in deriving a time vector for identifying collocated blocks, blocks adjacent to the current block and blocks not adjacent to the current block may be used. In the present disclosure, they can be referred to as adjacent blocks and non-adjacent blocks, respectively.

[0117] As an example, collocated blocks may be derived using candidates adjacent to the current block. Examples of adjacent blocks such as {A1, B1, B0, A0, B2, A2, B3} shown in FIG. 7 may be considered. This is just an example, and the order considered may be changed, and some of the listed candidates may be added or omitted. For example, the adjacent blocks are not limited to the above example, and all blocks adjacent to the current block may be targeted. The adjacent blocks may include blocks at specified positions among the blocks adjacent to the current block. As an example, it may be determined by checking by scanning the entire 4x4 block adjacent to the left side / upper side.

[0118] Specifically, the blocks existing on the left side can be scanned from the lower side to the upper side. That is, it can be scanned from the lower left adjacent block of the current block to the upper left adjacent block of the current block. The blocks existing at the upper end can be scanned from the right side to the left side. That is, it can be scanned from the upper right adjacent block of the current block to the upper left adjacent block of the current block.

[0119] Alternatively, the blocks existing on the left side can be scanned from the upper side to the lower side. That is, it is possible to scan from the upper left adjacent block of the current block to the lower left adjacent block of the current block. The blocks existing at the upper end can be scanned from the left side to the right side. That is, it is possible to scan from the upper left adjacent block of the current block to the upper right adjacent block of the current block. Note that this is just an example and modifications may be applied.

[0120] Also, as an example, in the process of checking the left and upper blocks, the number of blocks on the left / upper side may be defined. As an example, when scanning the left blocks in a specified order, if there are K available blocks, the scanning of the left blocks can be terminated and the upper blocks can be scanned. If there are L available blocks on the upper side, the scanning of the upper blocks may be terminated. At this time, the values of K and L may be predefined values or values variably determined according to the form of the blocks.

[0121] According to an embodiment of the present disclosure, in order to consider various reference blocks, collocated blocks can be induced using candidates that are not adjacent to the current block. Due to the order of video encoding / decoding, only the left and upper blocks of the current block are valid, so the process of inducing the MVP also considers only the motion vectors of the adjacent left and upper blocks.

[0122] To overcome such disadvantages, the TMVP technology refers to the motion information of the lower end and the right side position of the current block. At this time, for the collocated block, the lower right end position or the central position in the collocated picture is considered. In the embodiments of the present disclosure, similar to the intention of preferentially considering the lower right end position in the above-described TMVP, various SbTMVP candidates may be used.

[0123] As an example, the non-adjacent spatial adjacent block may include a first block including samples separated by a value obtained by multiplying the height of the current block by 2 from the upper left end sample adjacent to the upper left corner of the current block within the left sample line adjacent to the current block, and a second block including samples separated by a value obtained by multiplying the width of the current block by 2 from the upper left end sample within the upper end sample line adjacent to the current block.

[0124] As an example, for considering the movement of the lower end and / or the right side position, as shown in FIG. 7, blocks D0 and D1 may be used as non-adjacent blocks. This is one example, and the positions of blocks D0 and D1 may be applied to one or more of the candidates according to the following Equation 1.

[0125] [Equation 1] D0(x, y) = {(-1, 2xH), (-1, 2xH - 1), (-1, 4xH), (-1, 4xH - 1), (-1, W+H), (-1, W+H - 1)} D1(x, y) = {(2xW, -1), (2xW-1, -1), (4xW, - 1), (4xW-1, -1), (W+H, -1), (W+H-1, -1)}

[0126] In Equation 1, at this time, W and H represent the width and height of the current block, respectively. In Equation 1, alternative values may be applied to W and H. For example, W and H may be defined as values of 2, 4, 8, 16, 32.

[0127] Also, the non-adjacent block candidates may be determined according to the form of the current block. As an example, when W>=H, the non-adjacent block may be determined as in the following Equation 2.

[0128] [Equation 2] D0(x, y) = {(-1, 2xW), (-1, 2xW - 1), (-1, 4xW), (-1, 4xW - 1)} D1(x, y) = {(2xW, - 1), (2xW - 1, - 1), (4xW, - 1), (4xW - 1, - 1)}

[0129] Referring to Equation 2, when W >= H, it may be adjusted to a position different from the position of the lower left block induced by Equation 1. Also, as an example, when W < H, the non-adjacent blocks may be determined as in the following Equation 3.

[0130] [Equation 3] D0(x, y) = {(-1, 2xH), (-1, 2xH - 1), (-1, 4xH), (-1, 4xH - 1)} D1(x, y) = {(2xH, - 1), (2xH - 1, - 1), (4xH, - 1), (4xH - 1, - 1)}

[0131] Referring to Equation 3, when W < H, it may be adjusted to a position different from the position of the upper right block induced by Equation 1.

[0132] Also, a zero vector may be considered to account for various reference blocks. The existing SbTMVP replaces with a zero vector when the left block is not available, while the proposed method does not use the zero vector as an alternative candidate for other candidates and can be used when the number of candidates is not satisfied.

[0133] According to the embodiments of the present disclosure, by using the motion vectors of adjacent blocks, the motion vectors of non-adjacent blocks, and the zero vector, up to a maximum of MAX_NUM candidates can be considered as candidates for SbTMVP, thereby improving the prediction accuracy of SbTMVP and the compression performance.

[0134] Hereinafter, when inducing SbTMVP using a plurality of motion vectors (or temporal vectors) as in the above-described embodiments, a method will be described to enable consideration of more various candidates by checking for the presence or absence of effective duplication. The embodiments described below can apply the duplication check to the plurality of candidates described above.

[0135] For example, the presence or absence of duplication can be checked using a temporal vector derived from an adjacent block. That is, it may be confirmed whether there is duplication between the temporal vectors.

[0136] Or, for example, the presence or absence of duplication can be checked using the motion vector (DefaultMV) of the central position of the collocated block induced using the temporal vector.

[0137] Or, for example, the presence or absence of duplication can be checked using the motion vector for each sub-block unit of the collocated block induced using the temporal vector. At this time, the duplication check may be performed on the motion vectors of all sub-blocks within the collocated block. Or, the duplication check may be performed on the sub-blocks at specific positions within the collocated block.

[0138] The above-described duplication check method can be applied to the motion vectors of the adjacent blocks, non-adjacent blocks, and zero vectors described above. Also, in order to reduce the complexity of the duplication check, it can also be applied to only some candidates. That is, the duplication check may be performed between specific candidates defined in advance.

[0139] As an example, since the A0 block and the A1 block in FIG. 7 are very adjacent, the duplication check can be applied. Since the A0 block and the B0 block are not adjacent to each other, it is inferred that they have different motion information, and the duplication check does not need to be applied. Or, only adjacent blocks to each other may be the target of the duplication check, and the duplication check does not need to be applied to non-adjacent blocks and zero vectors.

[0140] FIG. 8 is a diagram for explaining a method for checking overlap between motion vectors according to an embodiment of the present disclosure.

[0141] Referring to FIG. 8, the overlap check can be determined based on whether there are the same reference picture and the same motion vector. As an embodiment, when the difference from the motion vector is smaller than a specific threshold, it can be determined as an overlap candidate.

[0142] Specifically, as shown in FIG. 8, the motion information in the collocated picture may be stored in specific units. For example, the motion information in the collocated picture may be stored in 8x8 units. Thus, when performing an overlap check based on the temporal vector, if the difference value between the motion vectors of each candidate is smaller than a specific value, they can exist within the same unit, and the motion information may overlap.

[0143] Therefore, according to an embodiment, by determining the presence or absence of overlap based on a threshold, more various candidates can be considered. As an example, the threshold may be defined in advance and may be defined as 1-pel, 1 / 4pel, 1 / 16pel, etc.

[0144] Hereinafter, a method for determining an optimal candidate from among a plurality of SbTMVP candidates induced by the method described above will be described.

[0145] Inter prediction may include a sub-block merge mode among general merge modes. The sub-block merge mode may include an SbTMVP mode and an affine merge mode. When the sub-block merge mode is used, a sub-block merge candidate list including N candidates may be configured. An index for specifying a candidate within the candidate list may be signaled. In one embodiment, the candidate list may include one SbTMVP candidate and N-1 affine merge candidates.

[0146] According to an embodiment of the present disclosure, even if the candidates for SbTMVP increase, finally, by selecting one SbTMVP candidate and constructing a candidate list as before, the performance of the affine merge candidates can be maintained. For this purpose, a template matching method may be applied. That is, among a plurality of SbTMVP candidates, the candidate with the smallest template matching cost may be selected as the SbTMVP.

[0147] The template matching cost may be calculated in at least one of the following ways. Adjustments may be applied considering the trade-off between accuracy and complexity.

[0148] - The SAD (sum of absolute differences) or MRSAD (mean-removed sum of absolute differences) of the template area can be calculated using the time vectors derived based on each adjacent / non-adjacent / zero vector. Since it is based on the similarity of the template between the current block and the collocated blocks derived from the time vector, a highly reliable time vector can be found. There is an advantage in that it can be calculated without an increase in complexity compared to calculating the DefaultMv and sub-block MV of each candidate.

[0149] - The SAD or MRSAD of the template area between the reference block indicated by the MV (DefaultMV) at the center position of the collocated block derived using the time vector and the current block can be calculated. There is an advantage in that the accuracy as an MVP can be increased compared to using the time vector in that the motion vector for the representative reference block actually applied (i.e., used as a prediction sample) is utilized.

[0150] - The SAD or MRSAD of the template region can be calculated using each sub-block unit MV within the collocated block induced using the time vector. This is the most accurate as an MVP in terms of utilizing the motion vectors actually applied in each sub-block unit, but the complexity may increase relatively. At this time, since not all sub-blocks can have a template region depending on the position of the sub-blocks within the current block, in one embodiment, the template matching cost can be calculated only for the sub-blocks adjacent to the left and upper sides of the current block.

[0151] When attempting to calculate the template matching cost in CU units (i.e., the current block), the template region may be determined as WxN and NxH. W and H represent the width and height of the CU, respectively. N may be determined as 1, 2, 4, etc. considering the complexity of the calculation and the size of the block. When attempting to calculate the template matching cost in sub-block units, all or part of WxM and MxH may be used for the template region. M may be determined as 1, 2, 4, etc. considering the complexity of the calculation and the size of the block.

[0152] Also, the method described above may be applied with the following modifications. When the number of sub-block merge candidates increases and the performance improves as various candidates are considered, a plurality of the SbTMVP candidates configured in the above embodiment can be used as the final sub-block merge candidates. As an example, a sub-block merge candidate list can be configured with two SbTMVP candidates. At this time, among the SbTMVP candidates induced as two SbTMVP candidates, the candidate with the smallest template matching cost and the next smallest candidate may be selected.

[0153] Specifically, as follows, the signaling method for SbTMVP may be applied with modifications. For example, a sub-block merge candidate list may be configured as shown in Table 1 below.

[0154]

Table 1

[0155] Referring to Table 1, as exemplified by Method 1, a specific number (e.g., 2) of SbTMVP candidates may be added to the sub-block merge candidate list, and then affine merge candidates may be added.

[0156] Also, a flag indicating whether the SbTMVP candidate is included in the sub-block merge candidate list may be signaled. That is, as exemplified by Method 2, the sub-block merge candidate list configuration may change depending on the SbTMVP presence / absence flag (SbTMVPFlag) that signals the presence or absence of SbTMVP in the sub-block merge mode. When SbTMVPFlag is 1, an SbTMVP candidate list may be configured, and when SbTMVPFlag is 0, an affine candidate list may be configured.

[0157] In Method 1 of Table 1, when there is no SbTMVP candidate or only one candidate exists, the affine merge candidate can be added to the sub-block merge candidate list, so waste can be prevented from the perspective of signaling bit efficiency. On the other hand, Method 2 has the advantage that while the bit usage for the affine candidates is maintained the same instead of adding SbTMVPFlag, various candidates for the SbTMVP mode can be considered. In one embodiment, in the case of Method 2, the number of SbTMVP candidates can be set to be generated up to the maximum number of candidates in the SbTMVP candidate list in order to prevent unnecessary transmission of SbTMVPFlag.

[0158] FIG. 9 is a flowchart illustrating a method for configuring a merge candidate list according to an embodiment of the present disclosure.

[0159] SbTMVP operates as a part of the sub-block merge mode among the general merge modes in terms of predicting motion information in sub-block units. Different from TMVP, SbTMVP does not obtain motion information at a defined position, so more accurate motion prediction is possible. Therefore, according to an embodiment of the present disclosure, a method of applying SbTMVP, which is the same as the aforementioned SbTMVP but induces motion information in coding block units rather than sub-block units, as one of the candidates for the regular merge mode is proposed. In the present disclosure, SbTMVP excluding the induction method in sub-block units is referred to as M-TMVP (Moved-TMVP). The name is not limited to this, and M-TMVP may simply be called TMVP.

[0160] Referring to FIG. 9, a merge candidate list may be configured by the regular merge mode. The merge candidate list may include spatial candidates, temporal candidates, non-adjacent candidates, HMVP (history-based motion vector predictor), and zero candidates. As an example, the spatial candidates, temporal candidates, non-adjacent candidates, HMVP, and zero candidates may be inserted into the merge candidate list in this order.

[0161] At this time, the temporal candidates may include M-TMVP according to this embodiment together with the existing TMVP. As an example, variations such as TMVP being replaced by M-TMVP, or the candidate with higher compression efficiency being selected from TMVP and M-TMVP are possible. At this time, template matching may be used to determine the candidate with higher compression efficiency. Also, M-TMVP may be added to the merge candidate list in an order different from that in FIG. 9.

[0162] For M-TMVP derivation, the methods described with reference to FIGS. 5 and 6 above may be equally applicable. As described above, M-TMVP can derive motion information in block units rather than sub-block units. At this time, the block unit may be a coding block unit. Therefore, in the methods described with reference to FIGS. 5 and 6 above, the sub-block unit motion vector derivation stage may be omitted or replaced by the block unit motion vector derivation stage. In the methods described with reference to FIGS. 5 and 6 above, operations other than the sub-block unit motion vector derivation stage may be equally applicable to M-TMVP, and related overlapping descriptions are omitted.

[0163] FIG. 10 is a diagram illustrating a method for deriving a temporal merge candidate according to an embodiment of the present disclosure.

[0164] Referring to FIG. 10, for M-TMVP that does not derive motion information in sub-block units, it can be changed so as to collect various motion information different from existing motion vectors. FIG. 10 assumes a case where the motion vector of the left adjacent block adjacent to the left side of the left lower end sample of the current block is used as the temporal vector to identify the collocated block. However, in addition to the left adjacent block, adjacent blocks at various positions described with reference to FIGS. 6 and 7 above may be used as temporal vector candidates.

[0165] When M-TMVP is used, in order to use more various motion information for prediction, the position of the collocated block can be changed and applied as follows. That is, in the process of deriving the collocated block using the motion vector and temporal vector of the adjacent block, the motion vector at the right lower end position of the block can be determined as ColMv (the motion vector of the collocated block). Or, if it is not available after checking the right lower end position, the motion vector at the center position can be determined as ColMv.

[0166] The M-TMVP proposed in this embodiment may be applied not only to the regular merge mode but also to the AMVP mode, and of course may be applied to various inter-prediction modes constituting the MVP.

[0167] FIG. 11 is a diagram illustrating a method for deriving motion information from a reference block according to an embodiment of the present disclosure.

[0168] In the derivation process of SbTMVP, the motion vectors (or motion information) of each sub-block existing in the collocated block may exist in various types. For example, a uni-directional prediction block and a bi-directional prediction block may exist, and they may have motion vectors in the same direction as the collocated picture or in the opposite direction of the collocated picture.

[0169] Referring to FIG. 11, when each sub-block existing in the collocated block is a uni-directional prediction block (i.e., only mvColL0 exists) and it is in the RA (random access) condition, the motion vector of the current block may be derived from the collocated block as shown in FIG. 11.

[0170] When the current picture is in the RA condition, colMvL0, which is the MVP of the current block, can be derived using the motion vector mvColL0 of the collocated block existing in the collocated picture. As shown in FIG. 11, the final colMvL0 can be derived using the ratio of the distance between the current picture (currPic) and the reference picture (refPic(L0)) and the distance between the collocated picture (colPic) and the reference picture of the collocated picture (colRefPic(L0)).

[0171] FIG. 12 is a diagram illustrating a method for deriving motion information from a reference block according to an embodiment of the present disclosure.

[0172] Referring to FIG. 12, when each sub-block existing in the collocated block is a uni-directional prediction block (i.e., only mvColL0 exists) and it is in the LD (low delay) condition, the motion vector of the current block may be derived from the collocated block as shown in FIG. 12.

[0173] When the current picture is in the LD condition, colMvLX, where X = 0, 1, can be derived as shown in FIG. 12 using the motion vector mvColL0 of the collocated block existing in the collocated picture. The final colMvLX can be derived using the ratio of the distance between the current picture (currPic) and each reference picture (refPic(LX), where X is 0, 1) and the distance between the collocated picture (colPic) and the reference picture (colRefPic(LX)) of each collocated picture.

[0174] Generally, in a video codec, bidirectional prediction provides a smoothing effect using the average between reference blocks and improves the compression performance by enhancing the prediction accuracy using two pieces of motion information. Therefore, in this embodiment, a method for improving the compression performance by changing the above-described process for deriving colMv so that bidirectional prediction blocks can be generated will be described with reference to FIGS. 13 and 14.

[0175] FIG. 13 is a diagram illustrating a method for deriving motion information from a reference block according to an embodiment of the present disclosure.

[0176] The method for deriving colMv in the RA condition described above with reference to FIG. 11 may be applied with modifications as shown in FIG. 13.

[0177] Referring to FIG. 13, in addition to colMvL0, colMvL1 can also be derived using mvColL0. Similarly, colMvL1 may be derived using the distance ratio and directionality with respect to the reference picture. At this time, colMvL1 may have a direction opposite to that of colMvL0 and may be derived by scaling the absolute value of colMvL0 according to the ratio of the distances between the reference pictures.

[0178] Also, as an example, when DefaultMV includes bidirectional motion information, the motion vector can be set for the sub-blocks that are L(1-X) directionally predicted in the LX direction using the motion vector of DefaultMv. Thereby, a bidirectional prediction effect can be expected.

[0179] FIG. 14 is a diagram illustrating a method for deriving motion information from a reference block according to an embodiment of the present disclosure.

[0180] The method for deriving colMv under the LD condition described above with reference to FIG. 12 may be applied with modifications as shown in FIG. 14.

[0181] Referring to FIG. 14, in the LD condition, the unidirectional prediction block can be set to have only unidirectional motion without being forcibly converted to the bidirectional prediction mode. Thereby, the computational complexity can be reduced and the consistency with RA can be maintained.

[0182] Alternatively, as an example, it may be applied without considering the RA / LD condition as follows. For example, when applying the motion information of the block to be predicted unidirectionally to the current block, the motion vector can be scaled and applied to the reference picture with a short distance between the current picture and the reference picture.

[0183] Also, according to an embodiment of the present disclosure, information for specifying a collocated picture may be signaled based on a high-level syntax. Regarding the prior art, for TMVP and SbTMVP, the collocated picture is determined based on the information parsed from the picture header and slice header as shown in Tables 2 and 3 below. The signaled information is applied uniformly to all temporal candidate derivation processes. Table 2 shows the picture header syntax, and Table 3 shows the slice header syntax.

[0184] [Table 2]

[0185] [Table 3]

[0186] Referring to Table 2 and Table 3, according to the prior art, since TMVP and SbTMVP induce time candidates using only the same collocated picture despite their different characteristics, there is a problem that the different characteristics between TMVP and SbTMVP cannot be effectively reflected.

[0187] Therefore, in one embodiment of the present disclosure, the collocated picture of SbTMVP can be signaled separately from TMVP so that SbTMVP can have a separate collocated picture. As an example, as shown in Table 4 and Table 5 below, the collocated picture of SbTMVP can be signaled separately from TMVP.

[0188] [Table 4]

[0189] [Table 5]

[0190] Table 4 shows the picture header syntax, and Table 5 shows the slice header syntax. Referring to Table 4, the syntax elements ph_sb_collocated_from_l0_flag and ph_sb_collocated_ref_idx indicating the collocated picture for SbTMVP may be defined / signaled. Referring to Table 5, the syntax elements sh_sb_collocated_from_l0_flag and sh_sb_collocated_ref_idx indicating the collocated picture for SbTMVP may be defined / signaled.

[0191] FIG. 15 is a diagram showing a schematic configuration of an inter prediction unit 332 that performs an inter prediction method according to the present disclosure.

[0192] Referring to FIG. 15, the intra prediction unit 331 may include a time vector induction unit 1500, a collocated block determination unit 1510, a motion vector induction unit 1520, and a prediction sample generation unit 1530.

[0193] The temporal vector derivation unit 1500 can derive the temporal vector of the current block based on the motion vector of one candidate among a group of candidates including a plurality of candidates.

[0194] The SbTMVP in the sub-block merge mode is different from the TMVP that derives the motion information of a block at a predefined position (i.e., the lower right end position or the center position of the current block) within a predefined reference picture (or a collocated picture, colPic) as a motion vector predictor (MVP). Instead, it uses the motion information of a collocated block (or col block) derived from the motion information of an adjacent block (e.g., the left adjacent block of the current block) as the MVP of the current block, and can derive the MVP in sub-block units.

[0195] An improved method may be applied in the SbTMVP derivation process that considers various temporal vectors for identifying collocated blocks. The proposed method may be applied in substantially the same way to the MVP derivation process in the MERGE / AMVP process in addition to the sub-block based MVP derivation process.

[0196] As described above, in order to improve the performance of the existing SbTMVP that uses the motion information of a collocated block derived from the motion information of the left adjacent block as the MVP, by considering various motion information derived from a plurality of adjacent blocks, an optimal MVP can be derived. At this time, the methods described in FIGS. 6 and 7 may be applied. Related overlapping explanations are omitted.

[0197] Also, as described above, when adding a candidate group to candidates based on a plurality of adjacent blocks, a duplication check may be performed to determine whether the candidate group has a motion vector that duplicates a candidate previously included in the candidate group. As an example, as described above with reference to FIG. 8, the presence or absence of duplication may be determined based on a comparison between the difference between motion vectors and a threshold value.

[0198] The collocated block determination unit 1510 can determine (or identify) the collocated block of the current block in the collocated picture based on the derived temporal vector. As an example, the collocated picture may be defined identically in the encoding device and the decoding device. Or, as an example, the collocated picture may be signaled by a higher-level syntax. The collocated block of the current block may be determined by the temporal vector within the collocated picture.

[0199] The motion vector derivation unit 1520 can derive the motion vector of the current block in sub-block units based on the motion vector of the collocated block. That is, the motion vector derivation unit 1520 can derive (or determine) the motion vector of the sub-block within the collocated block specified by the temporal vector as the motion vector of the corresponding sub-block within the current block.

[0200] The predictive sample generation unit 1530 can perform inter prediction on the current block based on the motion vector of the current block. The predictive sample generation unit 1530 can generate a predicted block of the current block by performing inter prediction based on the motion information derived in sub-block units within the current block. As an example, the predicted block of the current block may be generated in sub-block units by the motion information derived in sub-block units.

[0201] FIG. 16 is a diagram showing an inter prediction method performed by an encoding device 200 according to an embodiment of the present disclosure.

[0202] In one embodiment of the present disclosure, an inter prediction method performed by an encoding device will be described. Prior to performing inter prediction using SbTMVP, the embodiments described with reference to FIGS. 4 to 15 may be applied in substantially the same manner, and duplicate descriptions will be omitted here.

[0203] Referring to FIG. 16, the encoding device can determine a temporal vector of a current block based on a motion vector of one candidate among a group of candidates including a plurality of candidates (S1601).

[0204] SbTMVP in sub-block merge mode is different from TMVP that derives motion information of a block at a predefined position (i.e., the lower right end position or the central position of the current block) within a predefined reference picture (or, collocated picture, colPic) as a motion vector predictor (MVP). Instead, it uses the motion information of a collocated block (or, col block) derived from the motion information of an adjacent block (e.g., the left adjacent block of the current block) as the MVP of the current block, and can derive the MVP in units of sub-blocks.

[0205] An improved method in the SbTMVP derivation process considering various temporal vectors for identifying collocated blocks may be applied. The proposed method may be applied in substantially the same manner to the MVP derivation process in the MERGE / AMVP process in addition to the sub-block based MVP derivation process.

[0206] As described above, in order to improve the performance of the existing SbTMVP that uses the motion information of a collocated block derived from the motion information of the left adjacent block as the MVP, by considering various motion information derived from a plurality of adjacent blocks, an optimal MVP can be derived. At this time, the methods described with reference to FIGS. 6 and 7 may be applied. Related duplicate descriptions will be omitted.

[0207] Also, as described above, when inducing the optimal MVP, a template matching method may be applied.

[0208] Also, as described above, when adding a candidate group to candidates based on a plurality of adjacent blocks, a duplication check as to whether or not there is a motion vector that duplicates a candidate previously included in the candidate group may be performed. As an example, as described above with reference to FIG. 8, the presence or absence of duplication may be determined based on a comparison between the difference between motion vectors and a threshold value.

[0209] The encoding device can determine (or specify) a collocated block of the current block in the collocated picture based on the induced time vector (S1610). As an example, the collocated picture may be defined identically in the encoding device and the decoding device. Or, as an example, the collocated picture may be signaled by a higher-level syntax. The collocated block of the current block may be determined by the time vector within the collocated picture.

[0210] The encoding device can induce the motion vector of the current block in units of sub-blocks based on the motion vector of the collocated block (S1620). That is, the encoding device can induce (or determine) the motion vector of a sub-block within the collocated block specified by the time vector as the motion vector of the corresponding sub-block within the current block.

[0211] The encoding device can perform an inter prediction for the current block based on the motion vector of the current block (S1630). The encoding device can generate a predicted block of the current block by performing an inter prediction based on the motion information induced in units of sub-blocks within the current block. As an example, the predicted block of the current block may be generated in units of sub-blocks by the motion information induced in units of sub-blocks.

[0212] FIG. 17 is a diagram showing a schematic configuration of an inter prediction unit 221 that performs an inter prediction method according to the present disclosure.

[0213] Referring to FIG. 17, the inter prediction unit 221 may include a temporal vector determination unit 1700, a collocated block determination unit 1710, a motion vector determination unit 1720, and a prediction sample generation unit 1730.

[0214] Specifically, the temporal vector determination unit 1700 can determine a temporal vector of a current block based on a motion vector of one candidate among a group of candidates including a plurality of candidates.

[0215] The SbTMVP in the sub-block merge mode is different from the TMVP that derives the motion information of a block at a position defined in advance (i.e., the lower right end position or the center position of the current block) in a reference picture (or a collocated picture, colPic) defined in advance as a motion vector predictor (MVP). It uses the motion information of a collocated block (or a col block) derived from the motion information of an adjacent block (e.g., the left adjacent block of the current block) as the MVP of the current block and can derive the MVP in units of sub-blocks.

[0216] An improvement method in the SbTMVP derivation process considering various temporal vectors for specifying a collocated block may be applied. The proposed method may be applied in substantially the same manner to the MVP derivation process in the MERGE / AMVP process in addition to the sub-block based MVP derivation process.

[0217] As described above, in order to improve the performance of the existing SbTMVP that uses the motion information of the collocated block derived from the motion information of the left adjacent block as the MVP, by considering various motion information derived from a plurality of adjacent blocks, an optimal MVP can be derived. At this time, the methods described with reference to FIGS. 6 and 7 may be applied. Related overlapping explanations are omitted.

[0218] Also, as described above, when deriving an optimal MVP, a template matching method may be applied.

[0219] Also, as described above, when adding a candidate group to the candidates based on a plurality of adjacent blocks, a duplication check may be performed to determine whether the candidate group has a motion vector that duplicates a candidate previously included in the candidate group. As an example, as described above with reference to FIG. 8, the presence or absence of duplication may be determined based on a comparison between the difference between the motion vectors and a threshold value.

[0220] The collocated block determination unit 1710 can determine (or identify) the collocated block of the current block in the collocated picture based on the derived time vector. As an example, the collocated picture may be defined identically in the encoding device and the decoding device. Or, as an example, the collocated picture may be signaled by a higher-level syntax. The collocated block of the current block may be determined by the time vector within the collocated picture.

[0221] The motion vector determination unit 1720 can derive the motion vector of the current block in units of sub-blocks based on the motion vector of the collocated block. That is, the motion vector determination unit 1720 can derive (or determine) the motion vector of the sub-block in the collocated block specified by the time vector as the motion vector of the corresponding sub-block in the current block.

[0222] The prediction sample generation unit 1730 can perform inter prediction on the current block based on the motion vector of the current block. The prediction sample generation unit 1730 can generate a predicted block of the current block by performing inter prediction based on the motion information induced in units of sub-blocks within the current block. As an example, the predicted block of the current block may be generated in units of sub-blocks by the motion information induced in units of sub-blocks.

[0223] In the above-described embodiments, the method is described based on a flowchart as a series of steps or blocks, but the embodiments are not limited to the order of the steps, and a certain step may occur in a different order and at a different time or simultaneously from those described above. Also, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and other steps may be included or one or more steps of the flowchart may be deleted without affecting the scope of the embodiments of this document.

[0224] The method according to the embodiments of this document described above may be embodied in software form, and the encoding device and / or decoding device according to this document may be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.

[0225] When an embodiment is implemented as software in this document, the above-described method may be implemented as a module (process, function, etc.) that performs the above-described functions. The module may be stored in a memory and executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. The processor may include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory may include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document may be implemented and performed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each figure may be implemented and performed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm may be stored in a digital storage medium.

[0226] In addition, the decoding device and encoding device to which the embodiments of this specification are applied may be included in a multimedia broadcast transmission / reception device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (argumente reality) device, an image phone video device, a transportation means terminal (for example, a vehicle terminal (including an autonomous driving vehicle), an airplane terminal, a ship terminal, etc.) and a medical video device, etc., and may be used to process a video signal or a data signal. For example, as an OTT video (Over the top video) device, a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc. may be included.

[0227] In addition, the processing method to which the embodiments of this specification are applied may be produced in the form of a program executed by a computer and may be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiments of this specification may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (for example, transmission via the Internet). Also, a bitstream generated by an encoding method may be stored in a computer-readable recording medium or may be transmitted through a wired or wireless communication network.

[0228] Also, the embodiments of this specification may be embodied as a computer program product by program code, and the program code may be executed by a computer according to the embodiments of this specification. The program code may be stored on a computer-readable carrier.

[0229] FIG. 18 shows an example of a content streaming system to which the embodiments of the present disclosure are applicable.

[0230] Referring to FIG. 18, a content streaming system to which the embodiments of this specification are applied may generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0231] The encoding server compresses content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a camcorder directly generates a bitstream, the encoding server may be omitted.

[0232] The bitstream may be generated by an encoding method or a bitstream generation method to which the embodiments of this specification are applied, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0233] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as a medium to inform the user of what services are available. If a user requests a service desired by the user to the web server, the web server transmits it to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server plays a role of controlling commands / responses between each device in the content streaming system.

[0234] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0235] Examples of the user device include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, an HMD (head mounted display)), a digital TV, a desktop computer, a digital signage, and the like.

[0236] Each server within the content streaming system may be operated as a distributed server, and in this case, the data received by each server may be processed distributively.

[0237] The claims described in this specification may be combined in various ways. For example, the technical features of the method claims in this specification may be combined and embodied as an apparatus, and the technical features of the apparatus claims in this specification may be combined and embodied as a method. Further, the technical features of the method claims in this specification and the technical features of the apparatus claims may be combined and embodied as an apparatus, and the technical features of the method claims in this specification and the technical features of the apparatus claims may be combined and embodied as a method.

Claims

Claim 1 Deriving a temporal vector of a current block based on a motion vector of one candidate among a group of candidates including a plurality of candidates; Determining a collocated block of the current block in a collocated picture based on the temporal vector; Deriving a motion vector of the current block in sub-block units based on a motion vector of the collocated block; Performing an inter prediction on the current block based on the motion vector of the current block. A video decoding method comprising: Claim 2 The video decoding method according to claim 1, wherein the group of candidates includes adjacent spatial neighboring blocks of the current block and non-adjacent spatial neighboring blocks of the current block as candidates. Claim 3 The video decoding method according to claim 2, wherein the non-adjacent spatial neighboring blocks include a first block including samples separated by a value obtained by multiplying the height of the current block by 2 from a left upper end sample adjacent to the upper left corner of the current block in a left sample line adjacent to the current block, and a second block including samples separated by a value obtained by multiplying the width of the current block by 2 from the left upper end sample in an upper sample line adjacent to the current block. Claim 4 The video decoding method according to claim 1, wherein the one candidate is selected from the plurality of candidates included in the group of candidates based on template matching. Claim 5 The video decoding method according to claim 4, wherein the cost of the template matching is calculated based on SAD (sum of absolute differences) or MRSAD (mean-removed sum of absolute differences) between a template area of the current block and a template area of a block specified by motion vectors of the plurality of candidates included in the group of candidates. Claim 6 The cost of the template matching is the SAD (sum of absolute differences) or MRSAD (mean-removed sum of absolute differences) between the template region of the current block and the template region of the block specified by the motion vector of the block or sub-block induced based on the plurality of candidates included in the candidate group, and is calculated based on the motion vector of the block or sub-block induced based on the template region of the current block and the plurality of candidates included in the candidate group. The video decoding method according to claim 4.

7. Further comprising a step of constructing the candidate group including the plurality of candidates based on the peripheral blocks of the current block, The candidate group is constructed by adding the peripheral blocks at specific positions to the candidate group in a predefined order. The video decoding method according to claim 1.

8. The step of constructing the candidate group is, The video decoding method according to claim 7, including a step of checking whether the motion vector of the peripheral block at the specific position overlaps with the motion vectors of the candidates previously included in the candidate group.

9. Whether there is an overlap is determined based on whether the difference between the motion vector of the candidate previously included in the candidate group and the motion vector of the peripheral block at the specific position is smaller than a predefined threshold. The video decoding method according to claim 8.

10. The collocated picture is determined based on a syntax element signaled by at least one syntax of a picture header or a slice header, The syntax element is signaled separately from the syntax element indicating the collocated picture for the temporal motion vector predictor. The video decoding method according to claim 1.

11. Determining a temporal vector of the current block based on the motion vector of one candidate among a candidate group including a plurality of candidates; and Determining a collocated block of the current block in the collocated picture based on the temporal vector. Determining the motion vector of the current block in sub-block units based on the motion vectors of the collocated blocks; Performing inter prediction on the current block based on the motion vector of the current block, the method for video encoding comprising the steps.

12. A computer-readable storage medium storing a bitstream generated by the video encoding method according to Claim 11.

13. Determining a temporal vector of a current block based on a motion vector of one candidate among a group of candidates including a plurality of candidates; Determining a collocated block of the current block in a collocated picture based on the temporal vector; Determining the motion vector of the current block in sub-block units based on the motion vectors of the collocated blocks; Generating a predicted sample of the current block by performing inter prediction on the current block based on the motion vector of the current block; Generating a bitstream by encoding the current block based on the predicted sample; Transmitting data including the bitstream, the method for data transmission for video information comprising the steps.