Image encoding / decoding method and device based on inter prediction, and recording medium for storing bit stream
Through temporal motion vector prediction and geometric segmentation methods, the problem of low inter-frame prediction efficiency for high-resolution and stereoscopic image content is solved, and the coding efficiency and prediction accuracy are improved.
Patent Information
- Application Number
- CN202480008437.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-09
- Filing Date
- 2024-01-17
- Publication Date
- 2025-09-12
AI Technical Summary
Existing video coding technologies have difficulty in effectively compressing high-resolution and stereoscopic image content, especially in terms of inter-frame prediction.
The method of temporal motion vector prediction and geometric segmentation is adopted to determine the same-position block, derive the block vector and generate the motion vector of the current block based on the block vector for inter-frame prediction. Template matching and scaling factor are used to improve the prediction accuracy.
The coding efficiency of inter-frame prediction and the coding efficiency of geometric segmentation information are improved, and the accuracy of prediction is enhanced.
Smart Images

Figure CN120642335A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to encoders and decoders, and more particularly to methods and apparatus for video encoding and decoding. Background Art
[0002] As market demand for high-resolution video increases, technologies that can efficiently compress high-resolution images are becoming increasingly necessary. In response to this market demand, ISO / IEC's MPEG (Moving Picture Experts Group) and ITU-T's VCEG (Video Coding Experts Group) joined forces to form the Joint Collaboration Team on Video Coding (JCT-VC). In January 2013, they developed the HEVC (High Efficiency Video Coding) video compression standard and have been actively researching and developing the next-generation compression standard.
[0003] Video compression primarily involves intra-frame prediction, inter-frame prediction, transforms, quantization, entropy coding, and in-loop filters. As demand for high-resolution images increases, so does the demand for stereoscopic content as a new imaging service. Discussions are underway to develop video compression technologies that can efficiently deliver high-resolution and ultra-high-resolution stereoscopic content. Summary of the Invention
[0004] Technical issues
[0005] The present disclosure provides a method and apparatus for inter-frame prediction via temporal motion vector prediction.
[0006] The present disclosure provides a prediction method and apparatus based on geometric segmentation.
[0007] Technical Solutions
[0008] The decoding method and apparatus according to the present disclosure may: determine a co-located block for temporal motion vector prediction of a current block; derive a block vector from the co-located block; derive a motion vector of the current block based on the block vector; and generate a prediction signal for the current block by performing inter-frame prediction based on the motion vector of the current block.
[0009] In the decoding method and apparatus according to the present disclosure, the co-located block may be a block encoded in an intra block copy (IBC) mode.
[0010] In the decoding method and apparatus according to the present disclosure, the co-located block may be a block encoded in an intra mode based on template matching.
[0011] In the decoding method and apparatus according to the present disclosure, the block vector may be derived based on a position difference between the co-located block and a reference block used for template matching of the co-located block.
[0012] In the decoding method and apparatus according to the present disclosure, the reference picture of the current block may be replaced with the co-located picture to which the co-located block belongs.
[0013] In the decoding method and apparatus according to the present disclosure, a motion vector of a current block may be derived by applying a predetermined scaling factor to the block vector. Here, the scaling factor may be derived based on at least one of a first POC difference between a current picture to which the current block belongs and a reference picture for the current block, or a second POC difference between a co-located picture to which the co-located block belongs and a reference picture for the current block.
[0014] The encoding method and apparatus according to the present disclosure may: determine a co-located block for temporal motion vector prediction of a current block; derive a block vector from the co-located block; derive a motion vector of the current block based on the block vector; and generate a prediction signal for the current block by performing inter-frame prediction based on the motion vector of the current block.
[0015] A computer-readable digital storage medium is provided, in which encoded video / image information is stored, which enables a decoding device according to the present disclosure to perform an image decoding method.
[0016] A computer-readable digital storage medium is provided, in which video / image information generated according to the image encoding method according to the present disclosure is stored.
[0017] Provided are a method and apparatus for transmitting video / image information generated according to an image encoding method according to the present disclosure.
[0018] Technical Effects
[0019] The coding efficiency of inter-frame prediction can be improved through temporal motion vector prediction according to the present disclosure.
[0020] The coding efficiency of geometric partition information can be improved through the prediction of geometric partition information according to the present disclosure, and the accuracy of the prediction can be improved through more accurate geometric partitioning. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a block diagram showing an image encoding apparatus according to the present invention.
[0022] Figure 2 is a block diagram showing an image decoding apparatus according to the present invention.
[0023] Figure 3 A method for performing inter-prediction based on temporal motion vector prediction according to the present disclosure is shown.
[0024] Figure 4A method for performing inter prediction based on temporal motion vector prediction when the co-located block is a block coded in IBC mode is shown.
[0025] Figure 5 A method for performing inter prediction based on temporal motion vector prediction when the co-located block is a block coded in IBC mode is shown.
[0026] Figure 6 A prediction method based on geometric segmentation according to the present disclosure is shown.
[0027] Figure 7 A method for predicting geometric segmentation information according to the present disclosure is shown.
[0028] Figure 8 A method for predicting geometric segmentation information according to the present disclosure is shown.
[0029] Figure 9 A method for predicting geometric segmentation information according to the present disclosure is shown.
[0030] Figure 10 A method for deriving geometric segmentation information according to the present disclosure is shown.
[0031] Figure 11 A method for deriving geometric segmentation information according to the present disclosure is shown.
[0032] Figure 12 A method for predicting geometric segmentation information according to the present disclosure is shown. DETAILED DESCRIPTION
[0033] With reference to the accompanying drawings attached to this specification, embodiments of the present invention will be described in detail so that those skilled in the art can easily implement the embodiments in the technical field to which the present invention belongs. However, the present invention can be implemented in different forms and is not limited to the embodiments described herein. In addition, in order to clearly illustrate the present invention in the drawings, parts not related to the description are omitted, and like reference numerals are attached to like parts throughout the specification.
[0034] Throughout this specification, when a part is referred to as being “connected to” another part, it includes not only a case of being directly connected but also a case of being electrically connected with other elements therebetween.
[0035] In addition, throughout this specification, when a part is referred to as “comprising” components, it means that other components may also be included, rather than excluding other components, unless explicitly objected otherwise.
[0036] In addition, although the terms "first", "second", etc. may be used to describe various components, the components should not be limited by these terms. These terms are only used to distinguish one component from other components.
[0037] In addition, in the embodiments of the devices and methods described in this specification, some configurations of the devices or some steps of the methods may be omitted. In addition, the order of some configurations of the devices or some steps of the methods may be changed. In addition, another configuration may be inserted into some configurations of the devices, or another step may be inserted into some steps of the methods.
[0038] In addition, some configurations or some steps in the first embodiment of the present disclosure may be added to the second embodiment of the present disclosure, or may be replaced with some configurations or some steps in the second embodiment.
[0039] In addition, since the structural units shown in the embodiments of the present disclosure are shown independently to represent different characteristic functions, it does not mean that each structural unit is composed of a separate hardware or software structural unit. In other words, for the sake of ease of description, each structural unit is described by being listed as each structural unit, but at least two structural units in each structural unit can be combined to form a structural unit, or a structural unit can be divided into multiple structural units to perform functions. The integrated embodiment and individual embodiment of each structural unit are also included in the scope of the rights of the present disclosure, unless they exceed the essence of the present disclosure.
[0040] First, the terms used in this article are briefly described below.
[0041] The video decoding device described below may be a device included in a server terminal such as a civilian security camera device, a civilian security system, a military security camera device, a military security system, a personal computer (PC), a laptop computer, a portable multimedia player (PMP), a wireless communication terminal, a smart phone, a TV application server and a service server, etc., and may refer to various devices, including user terminals such as various equipment, communication devices for communicating with wired or wireless communication networks such as communication modems, etc., a memory for storing various programs and data for decoding images or performing inter-frame prediction or intra-frame prediction for decoding, and a microprocessor for executing programs and calculations and controlling programs.
[0042] In addition, the image encoded into a bit stream by the encoder can be sent to the image decoding device in real time or non-real time through a wired or wireless communication network such as the Internet, a wireless local area network, a wireless LAN network, a wireless broadband (WiBro) network, a mobile communication network, etc., or through various communication interfaces such as cables, universal serial buses (USB), etc., and can be decoded, reconstructed and played back as an image. Alternatively, the bit stream generated by the encoder can be stored in a memory. The memory can include both volatile memory and non-volatile memory. In this specification, the memory can be expressed as a recording medium for storing a bit stream.
[0043] Generally, a video can be composed of a series of pictures, and each picture can be divided into coding units such as blocks. In addition, those skilled in the art to which this embodiment belongs should understand that the term "picture" described below can be replaced by other terms with equivalent meanings, such as "image", "frame", etc. In addition, those skilled in the art to which this embodiment belongs should understand that the term "coding unit" can be replaced by other terms with equivalent meanings, such as "unit block", "block", etc.
[0044] Hereinafter, embodiments of the present invention will be described in more detail with reference to the accompanying drawings. In describing the present invention, repeated descriptions of the same components are omitted.
[0045] Figure 1 is a block diagram showing an image encoding apparatus according to the present invention.
[0046] Reference Figure 1 , a conventional image encoding device 100 may include a picture division unit 110, prediction units 120 and 125, a transform unit 130, a quantization unit 135, a rearrangement unit 160, an entropy encoding unit 165, an inverse quantization unit 140, an inverse transform unit 145, a filter unit 150, and a memory 155.
[0047] The picture division unit 110 may divide the input picture into at least one processing unit. In this case, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). Hereinafter, in the embodiments of the present disclosure, a coding unit may be used as a unit for performing encoding, or may be used as a unit for performing decoding.
[0048] The prediction unit may be divided into at least one square shape or rectangular shape having the same size within one coding unit, or may be divided so that any one of the prediction units divided within one coding unit has a different shape and / or size from another prediction unit. When a prediction unit for performing intra-frame prediction is generated based on a coding unit, if the prediction unit is not the minimum coding unit, intra-frame prediction may be performed without being divided into a plurality of prediction units N×N.
[0049] The prediction units 120 and 125 may include an inter-frame prediction unit 120 that performs inter-frame prediction or inter-picture prediction, and an intra-frame prediction unit 125 that performs intra-frame prediction or intra-picture prediction. Whether inter-frame prediction or intra-frame prediction is performed for a prediction unit may be determined, and specific information according to each prediction method (e.g., intra-frame prediction mode, motion vector, reference picture, etc.) may be determined. The residual value (residual block) between the generated prediction block and the original block may be input to the transform unit 130. Furthermore, prediction mode information, motion vector information, and the like used for prediction may be encoded using the residual values in the entropy coding unit 165 and transmitted to the decoder. However, when the decoder-side motion information derivation technique according to the present invention is applied, prediction mode information, motion vector information, and the like are not generated in the encoder, and therefore the corresponding information is not transmitted to the decoder. Alternatively, the encoder may signal and transmit information indicating that motion information is derived and used on the decoder side, as well as information regarding the technique used to derive motion information.
[0050] The inter-frame prediction unit 120 may predict a prediction unit based on information of at least one picture in a previous picture or a subsequent picture of the current picture, or in some cases may predict a prediction unit based on information of some regions encoded within the current picture. The inter-frame prediction unit 120 may include a reference picture interpolation unit, a motion prediction unit, and a motion compensation unit.
[0051] The reference picture interpolation unit can be provided with reference picture information from the memory 155, and can generate pixel information equal to or less than an integer pixel in the reference picture. For luma pixels, an 8-tap DCT-based interpolation filter with different filter coefficients can be used to generate pixel information equal to or less than an integer pixel in units of 1 / 4 pixels. For chroma signals, a 4-tap DCT-based interpolation filter with different filter coefficients can be used to generate pixel information equal to or less than an integer pixel in units of 1 / 8 pixels.
[0052] The motion prediction unit can perform motion prediction based on the reference picture interpolated by the reference picture interpolation unit. As a method for calculating the motion vector, various methods such as FBMA (block matching algorithm based on full search), TSS (three-step search), NTS (new three-step search algorithm), etc. can be used. Based on the interpolated pixels, the motion vector can have a motion vector value in units of 1 / 2 or 1 / 4 pixels. In the motion prediction unit, the current prediction unit can be predicted by using different motion prediction methods. As the motion prediction method, various methods such as skip mode, merge mode, advanced motion vector prediction (AMVP) mode, intra-frame block copy mode, etc. can be used. In addition, when applying the decoder-side motion information derivation technology according to the present invention, a bilateral matching method and a template matching method using motion trajectory can be applied as the method performed in the motion prediction unit. In this regard, the following will be described. Figure 3 The template matching method and the bilateral matching method are described in detail in .
[0053] The intra-frame prediction unit 125 can generate a prediction unit based on reference pixel information surrounding the current block, that is, pixel information within the current picture. When the reference pixel is a pixel for which inter-frame prediction is performed because the neighboring block in the current prediction unit is a block for which inter-frame prediction is performed, the reference pixel included in the block for which inter-frame prediction is performed can be used by replacing it with the reference pixel information of the neighboring block for which intra-frame prediction is performed. In other words, when a reference pixel is unavailable, the unavailable reference pixel information can be used by replacing it with at least one reference pixel from among the available reference pixels.
[0054] In addition, a residual block including residual value information that is a difference value between a prediction unit predicted based on the prediction units generated in the prediction units 120 and 125 and an original block in the prediction unit may be generated. The generated residual block may be input to the transform unit 130.
[0055] In the transform unit 130, the original block and the residual block including the residual value information in the prediction unit generated by the prediction units 120 and 125 may be transformed by using a transform method such as discrete cosine transform (DCT), discrete sine transform (DST), and KLT. Whether DCT, DST, or KLT is applied to transform the residual block may be determined based on intra prediction mode information in the prediction unit used to generate the residual block.
[0056] The quantization unit 135 may quantize the value transformed into the frequency domain in the transform unit 130. The quantization coefficient may be changed according to the importance of the block or the image. The value calculated in the quantization unit 135 may be provided to the inverse quantization unit 140 and the rearrangement unit 160.
[0057] The rearrangement unit 160 may perform rearrangement of coefficient values with respect to the quantized residual value.
[0058] The rearrangement unit 160 can change the coefficients in the two-dimensional block form into a one-dimensional vector form through a coefficient scanning method. For example, in the rearrangement unit 160, a zigzag scanning method can be used to scan the DC coefficient to the coefficient in the high frequency domain, and then change it into a one-dimensional vector form. Depending on the size of the transform unit and the intra-frame prediction mode, instead of the zigzag scanning, vertical scanning that scans the coefficients in the two-dimensional block form in the column direction or horizontal scanning that scans the coefficients in the two-dimensional block form in the row direction can be used. In other words, which scanning method to use among zigzag scanning, vertical scanning, and horizontal scanning can be determined based on the size of the transform unit and the intra-frame prediction mode.
[0059] The entropy coding unit 165 may perform entropy coding based on the value calculated by the rearrangement unit 160. For example, the entropy coding may use various coding methods such as Exponential Golomb coding, CAVLC (Context Adaptive Variable Length Coding), and CABAC (Context Adaptive Binary Arithmetic Coding). In this regard, the entropy coding unit 165 may encode residual value coefficient information in the coding units from the rearrangement unit 160 and the prediction units 120 and 125. In addition, according to the present invention, information indicating that motion information is derived and used on the decoder side and information about the technology used to derive the motion information may be signaled and transmitted.
[0060] In the inverse quantization unit 140 and the inverse transform unit 145, the value quantized in the quantization unit 135 is dequantized, and the value transformed in the transform unit 130 is inversely transformed. The residual value generated in the inverse quantization unit 140 and the inverse transform unit 145 can generate a reconstructed block by combining with the prediction unit predicted by the motion estimation unit, the motion compensation unit, and the intra prediction unit included in the prediction units 120 and 125.
[0061] The filter unit 150 may include at least one of a deblocking filter, an offset modification unit, or an adaptive loop filter (ALF). The deblocking filter may remove block distortion caused by the boundaries between blocks in the reconstructed image. The offset modification unit may modify the offset from the original image in units of pixels for the image to be deblocked. Offset modification may be performed for a specific image using a method of dividing the pixels included in the image into a certain number of regions, determining the regions to be offset, and applying the offset to the corresponding regions, or a method of applying the offset by considering the edge information of each pixel. Adaptive loop filtering (ALF) may be performed based on a value obtained by comparing the filtered reconstructed image with the original image. The pixels included in the image may be divided into predetermined groups, and a filter to be applied to the corresponding group may be determined to perform filtering differently for each group.
[0062] The memory 155 may store the reconstructed block or picture calculated by the filter unit 150 and may provide the stored reconstructed block or picture to the prediction units 120 and 125 when performing inter prediction.
[0063] Figure 2 is a block diagram showing an image decoding apparatus according to the present invention.
[0064] Reference Figure 2 , the image decoder 200 may include an entropy decoding unit 210 , a rearrangement unit 215 , an inverse quantization unit 220 , an inverse transform unit 225 , prediction units 230 and 235 , a filter unit 240 , and a memory 245 .
[0065] When an image bitstream is input from an image encoder, the input bitstream may be decoded in a process reverse to that of the image encoder.
[0066] The entropy decoding unit 210 may perform entropy decoding in a process opposite to the process of performing entropy encoding in the entropy encoding unit of the image encoder. For example, various methods such as Exponential Golomb coding, CAVLC (Context Adaptive Variable Length Coding), and CABAC (Context Adaptive Binary Arithmetic Coding) may be applied depending on the method performed in the image encoder.
[0067] In the entropy decoding unit 210, information related to intra prediction and inter prediction performed in the encoder may be decoded.
[0068] The rearrangement unit 215 may perform rearrangement based on a method of rearranging in the coding unit the bitstream entropy-decoded in the entropy decoding unit 210. Coefficients expressed in a one-dimensional vector form may be reconstructed and rearranged into coefficients in a two-dimensional block form.
[0069] The inverse quantization unit 220 may perform dequantization based on a quantization parameter provided in the encoder and a coefficient value of the rearranged block.
[0070] The inverse transform unit 225 can perform the transform performed in the transform unit on the result of the quantization performed in the image encoder, that is, the inverse transform for DCT, DST, and KLT, that is, inverse DCT, inverse DST, and inverse KLT. The inverse transform can be performed based on the transmission unit determined in the image encoder. In the inverse transform unit 225 of the image decoder, the transform technology (e.g., DCT, DST, KLT) can be selectively performed based on multiple information such as the prediction method, the size of the current block, the prediction direction, etc.
[0071] The prediction units 230 and 235 may generate a prediction block based on the information related to the prediction block generation provided in the entropy decoding unit 210 and the pre-decoded block or picture information provided in the memory 245 .
[0072] As described above, when the size of the prediction unit is the same as the size of the transform unit when performing intra prediction or intra-screen prediction in the same manner as the operation in the image encoder, intra prediction of the prediction unit is performed based on the pixel at the left position, the pixel at the upper left position, and the pixel at the top position of the prediction unit. However, when the size of the prediction unit is different from the size of the transform unit when performing intra prediction, intra prediction can be performed by using reference pixels based on the transform unit. In addition, intra prediction using N×N partitioning can be used only for the minimum coding unit.
[0073] The prediction units 230 and 235 may include a prediction unit determination unit, an inter-frame prediction unit, and an intra-frame prediction unit. The prediction unit determination unit may receive various information input from the entropy decoding unit 210, such as prediction unit information, prediction mode information of the intra-frame prediction method, and information related to motion prediction of the inter-frame prediction method, divide the prediction unit in the current coding unit, and determine whether the prediction unit performs inter-frame prediction or intra-frame prediction. On the other hand, when the encoder 100 transmits information indicating that motion information is derived and used on the decoder side and information about the technology used to derive motion information, but does not transmit motion prediction related information for inter-frame prediction, the prediction unit determination unit determines whether the inter-frame prediction unit 23 performs prediction based on the information transmitted from the encoder 100.
[0074] The inter-frame prediction unit 230 can perform inter-frame prediction on the current prediction unit by using information required for inter-frame prediction in the current prediction unit provided by the image encoder, based on information included in at least one picture of a previous picture or a subsequent picture of the current picture including the current prediction unit. To perform inter-frame prediction, it can be determined based on the coding unit whether the motion prediction method used in the prediction unit included in the corresponding coding unit is skip mode, merge mode, AMVP mode, or intra block copy mode. Alternatively, the inter-frame prediction unit 230 can perform inter-frame prediction by autonomously deriving motion information by deriving and using information on the motion information and information on the technique used to derive the motion information at the decoder side according to an instruction provided by the image encoder.
[0075] The intra prediction unit 235 can generate a prediction block based on pixel information within the current picture. When the prediction unit is a prediction unit that performs intra prediction, intra prediction can be performed based on the intra prediction mode information in the prediction unit provided from the image encoder. The intra prediction unit 235 may include an adaptive intra smoothing (AIS) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is a part for performing filtering on the reference pixels of the current block, and the AIS filter can be applied by determining whether to apply the filter according to the prediction mode in the current prediction unit. AIS filtering can be performed on the reference pixels of the current block by using the prediction mode and AIS filter information in the prediction unit provided from the image encoder. When the prediction mode of the current block is a mode in which AIS filtering is not performed, the AIS filter may not be applied.
[0076] When the prediction mode in the prediction unit is a prediction unit that performs intra-frame prediction based on pixel values interpolated from reference pixels, the reference pixel interpolation unit may interpolate the reference pixels to generate reference pixels in units of pixels equal to or less than an integer value. When the prediction mode in the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixels, the reference pixels may not be interpolated. When the prediction mode of the current block is the DC mode, the DC filter may generate the prediction block through filtering.
[0077] The reconstructed block or picture may be provided to the filter unit 240. The filter unit 240 may include a deblocking filter, an offset modification unit, and an ALF.
[0078] Information about whether a deblocking filter is applied to a corresponding block or picture and information about whether a strong filter or a weak filter is applied when applying the deblocking filter can be provided from the image encoder. The deblocking filter of the image decoder can receive the deblocking filter related information provided from the image encoder and perform deblocking filtering on the corresponding block in the image decoder.
[0079] The offset modification unit may perform offset modification on the reconstructed image based on the type of offset modification applied to the image during encoding, offset value information, etc. ALF may be applied to the coding unit based on information on whether ALF is applied, ALF coefficient information, etc. provided from the encoder. Such ALF information may be provided by being included in a specific parameter set.
[0080] The memory 245 may store and use the reconstructed picture or block as a reference picture or reference block, and also provide the reconstructed picture to the output unit.
[0081] In this disclosure, terms may be defined as follows.
[0082] The current block (CurrCb) may mean a block to be currently encoded / decoded.
[0083] The current picture (CurrPic) may mean a picture including the current block.
[0084] The motion vector (CurrMV) of the current block may specify a reference block of the current block within a reference picture (RefPic) of the current block.
[0085] A co-located block (ColCb) is a block belonging to a different picture from the current picture, and may refer to a block at the same position as the current block. A co-located block may refer to a block including sample coordinates adjacent to the lower right corner of the current block. Alternatively, a co-located block may refer to a block including the center sample coordinates of the current block. However, this is not limiting, and a co-located block may refer to a block including the upper left sample coordinates of the current block. Alternatively, a co-located block may refer to a block including sample coordinates shifted by a predetermined offset vector from the upper left sample coordinates of the current block. Here, the offset vector may be determined based on a motion vector of a spatially adjacent block (e.g., a left adjacent block, a top adjacent block) adjacent to the current block.
[0086] A co-located picture (ColPic) may refer to a picture that includes a co-located block. A co-located picture may be any of a plurality of reference pictures belonging to a reference picture list used for inter prediction of the current block. A co-located picture may be included in at least one of the reference picture list (List0) in the L0 direction or the reference picture list (List1) in the L1 direction.
[0087] The motion vector (ColMV) of the co-located block may specify a reference block of the co-located block within the reference picture (ColRefPic) of the co-located block. Alternatively, when the co-located block is a block encoded in the intra block copy mode (IBC mode), the motion vector (ColMV) of the co-located block may refer to a block vector specifying a reference block of the co-located block located within the co-located picture.
[0088] The motion information may include at least one of a motion vector, a block vector, a reference picture index, or prediction direction information.
[0089] Figure 3 A method for performing inter-prediction based on temporal motion vector prediction according to the present disclosure is shown.
[0090] Reference Figure 3 , the motion vector of the current block can be obtained based on the temporal motion vector prediction. Inter-frame prediction (or motion compensation) can be performed based on the motion vector of the current block to generate a prediction signal for the current block.
[0091] Temporal motion vector prediction according to the present disclosure is a method for deriving a motion vector of a current block by using a feature that motions at the same position are similar in different pictures.
[0092] When generating a prediction signal for a current block by inter-frame prediction, a candidate list may be generated based on spatial candidates and / or temporal candidates for the current block. Here, spatial candidates may be derived from neighboring blocks that are spatially adjacent to the current block (i.e., spatially adjacent blocks). Spatial candidates may have motion information of spatially adjacent blocks. Temporal candidates may be derived from neighboring blocks that are temporally adjacent to the current block (i.e., temporally adjacent blocks). Similarly, temporal candidates may have motion information of temporally adjacent blocks. Temporal candidates may refer to candidates used for temporal motion vector prediction. Temporally adjacent blocks may correspond to co-located blocks.
[0093] The motion information of the co-located block used for temporal motion vector prediction can be confirmed. When the co-located block is a block encoded in inter-frame mode (i.e., when the co-located block has motion information), the motion vector of the current block can be derived based on the motion vector of the corresponding co-located block. As an example, the motion vector of the co-located block can be configured as the motion vector of the current block. Alternatively, the motion vector of the co-located block can be modified based on at least one of the reference picture information of the co-located block or the reference picture information of the current block, and the modified motion vector can be configured as the motion vector of the current block. Here, the reference picture information may include at least one of a picture order count (POC) of the reference picture, a reference picture index, or whether the reference picture is a long-term reference picture. On the other hand, when the co-located block is not a block encoded in inter-frame mode, temporal motion vector prediction may not be used.
[0094] Alternatively, temporal motion vector prediction can also be used for subblock-based temporal motion vector prediction. Subblock-based temporal motion vector prediction may refer to a method for performing inter-frame prediction on a subblock basis by using a motion vector corresponding to the position of a reference block among the motion vectors stored in a motion information buffer of a reference picture. Here, a reference picture may refer to a co-located picture including a co-located block. The co-located block may be specified based on the motion vector of a spatially neighboring block adjacent to the current block.
[0095] The reference block can be divided into multiple sub-blocks, and the presence of a motion vector can be checked on a sub-block basis. When a motion vector exists in a sub-block within the reference block, inter-frame prediction can be performed based on the motion vector of the corresponding sub-block to generate a prediction signal for the sub-block corresponding to the corresponding sub-block within the current block. On the other hand, when a motion vector does not exist in a sub-block within the reference block, inter-frame prediction can be performed based on the motion vector of a spatially neighboring block adjacent to the current block to generate a prediction signal for the sub-block corresponding to the corresponding sub-block within the current block. Alternatively, when a motion vector does not exist in a sub-block within the reference block, inter-frame prediction can be performed based on the motion vector of another sub-block within the reference block. Here, the other sub-block may be a sub-block including the center sample coordinates within the reference block.
[0096] exist Figure 3 In
[15] , CurPic, ColPic, ColRefPic, and RefPic may respectively refer to the current picture, the co-located picture including the co-located block used for temporal motion vector prediction, the reference picture of the co-located block, and the reference picture of the current block. CurCb and ColCb may respectively refer to the current block and the co-located block. CurMV and ColMV may respectively refer to the motion vector of the current block and the motion vector of the co-located block. Here, the motion vector of the current block may be derived through temporal motion vector prediction.
[0097] Information about the co-located picture may be obtained from at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH). Based on the information about the co-located picture, any one of a plurality of reference pictures belonging to a reference picture list may be determined, and the determined reference picture may be used as the co-located picture. The information about the co-located picture may include at least one of a co-located picture index, a flag indicating whether the co-located picture belongs to a reference picture list in the L0 direction, or a flag indicating whether the co-located picture belongs to a reference picture list in the L1 direction.
[0098] The encoder may determine the picture P or B in the reference picture list as a co-located picture for temporal motion vector prediction, and encode information about the co-located picture to specify the co-located picture.
[0099] To reduce the size of the motion information buffer, the motion information of the co-located blocks can be sampled in units of a specific block size and stored in the motion information buffer. In this case, the specific block size can be a square block such as 4×4, 8×8, or 16×16. Alternatively, the specific block size can be a non-square block such as N×M. Here, N and M can be any square number of 2.
[0100] In addition, in order to further reduce the size of the motion information buffer, the motion vector can be compressed and stored in the motion information buffer. When the motion vector is compressed and stored, the motion vector can be reconstructed by performing an inverse process of the compression method on the stored motion vector.
[0101] The compression method of motion vector can represent motion vector with a smaller number of bits. As an example, when the x-axis size and y-axis size of the current motion vector are represented with 18 bits, the size of the x-axis motion vector and the y-axis motion vector can be represented with 10 bits. In this case, the motion vector represented with 10 bits can be represented as a fixed point (fixed-point). When the motion vector is represented as a fixed point, compression can be performed by shifting the motion vector to the right by the difference between the original number of bits and the compressed number of bits. Alternatively, the motion vector represented with 10 bits can be represented as a floating point (floating-point). When the motion vector is represented as a floating point, 6 bits in the 10 bits can mean the mantissa with a sign, and the remaining 4 bits in the 10 bits can mean the exponent. The number of bits used to represent the exponent can be less than the number of bits used to represent the mantissa.
[0102] The number of bits of the compressed motion vector can be determined based on the type of the current picture or slice. As an example, when the current picture or slice can use only intra-frame prediction, a block vector instead of a motion vector can be stored at the position of a block encoded in IBC mode. Since the block vector can be represented with a size smaller than the motion vector, the block vector can be represented with a smaller number of bits in compression. The number of bits used to represent the compressed block vector can be determined based on the size of the encoding / decoding unit. Here, the size can be defined as width, height, the product of width and height, the maximum / minimum value of width and height, etc. The encoding / decoding unit can be a picture, a tile, a slice, a coding tree unit (CTU), or a virtual pipeline data unit (VPDU). The IBC mode can be a method for generating a prediction signal by using a reference block that belongs to a pre-reconstructed area within the same picture as the current block and is specified by a block vector.
[0103] Alternatively, the block vectors for blocks encoded in IBC mode may be stored in a motion information (or motion field) buffer. As an example, the block vectors may be stored in the same motion information buffer as the motion vectors for blocks encoded in inter mode. Alternatively, the block vectors may be stored in a different motion information buffer than the motion vectors for blocks encoded in inter mode. In this case, the motion information buffer may be referred to as a block vector buffer.
[0104] The block vector can be compressed and stored in the motion information buffer. Alternatively, when the co-located block used for temporal motion vector prediction (or sub-block-based temporal motion vector prediction) is a block encoded in IBC mode, the corresponding co-located block can have a block vector. The block vector of the co-located block can be added to the candidate list and used to derive the motion vector of the current block. The block vector of the co-located block can be compressed and stored in the motion information buffer. As an example, the block vector can be compressed with a different precision than the motion vector used for other inter-frame modes. The precision of the block vector can be 1 / N, and N can be any one of 2, 4, 8 or 16.
[0105] In temporal motion vector prediction, information of a co-located block corresponding to the current block in a co-located picture may be first checked to predict the motion vector of the current block.
[0106] When the co-located block is encoded in intra mode, IBC mode, or palette mode, temporal motion vector prediction may not be performed based on the corresponding co-located block. On the other hand, when the co-located block is not encoded in intra mode, IBC mode, or palette mode, temporal motion vector prediction may be performed based on the corresponding co-located block, and the motion vector of the current block may be derived based on the motion vector of the co-located block.
[0107] When performing temporal motion vector prediction, the motion vector of the current block may be scaled based on at least one of a POC difference (curPocDiff) between the current picture and a reference picture of the current block or a POC difference (colPocDiff) between a co-located picture and a reference picture of the co-located block.
[0108] When the reference picture of the co-located block is a long-term reference picture, or when curPocDiff and colPocDiff are the same, scaling of the motion vector of the current block can be omitted.
[0109] Inter-frame prediction may be performed based on the motion vector of the current block to generate a prediction signal of the current block.
[0110] Figure 4 A method for performing inter prediction based on temporal motion vector prediction when the co-located block is a block coded in IBC mode is shown.
[0111] In the temporal motion vector prediction according to the present disclosure, when the co-located block is a block encoded in IBC mode, the block vector of the co-located block can be used for temporal motion vector prediction in the same manner as the motion vector. As an example, temporal motion vector prediction can be enabled in the next picture immediately following the reconstructed picture I. In this case, the motion information (or motion vector) of the co-located block or the co-located block can be added to the candidate list as a merge candidate or an AMVP candidate. Alternatively, the block vector of the co-located block can be used for temporal motion vector prediction in units of sub-blocks.
[0112] Not only the motion vectors of blocks coded in inter-frame mode, but also the block vectors of blocks coded in IBC mode can be stored in the motion information buffer (motion field). When the co-located block of the current block is not only a block coded in IBC mode but also a block coded in inter-frame mode, the co-located block or the motion information of the co-located block can be used as a temporal candidate.
[0113] The accuracy of the prediction can be improved by increasing the number of candidates included in the candidate list. In other words, the number of temporal candidates that can be added to the candidate list can also be increased. For blocks encoded in IBC mode and for blocks encoded in inter-frame mode, block vectors can be considered as temporal candidates, and such different temporal motion information can be used to increase the accuracy of the prediction and improve compression performance.
[0114] exist Figure 4 In
[15] , CurPic and ColPic may refer to the current picture and the co-located picture including the co-located block used for temporal motion vector prediction, respectively. CurCb and ColCb may refer to the current block and the co-located block, respectively. CurMV and ColMV may refer to the motion vector of the current block and the motion vector of the co-located block, respectively. Here, the motion vector of the current block may be derived through temporal motion vector prediction.
[0115] Information about the co-located picture may be obtained from at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH). Based on the information about the co-located picture, any one of a plurality of reference pictures belonging to a reference picture list may be determined, and the determined reference picture may be used as the co-located picture. The information about the co-located picture may include at least one of a co-located picture index, a flag indicating whether the co-located picture belongs to a reference picture list in the L0 direction, or a flag indicating whether the co-located picture belongs to a reference picture list in the L1 direction.
[0116] The encoder may determine the picture P or B in the reference picture list as a co-located picture for temporal motion vector prediction, and encode information about the co-located picture to specify the co-located picture.
[0117] In temporal motion vector prediction, information of a co-located block corresponding to the current block in a co-located picture may be first checked to predict the motion vector of the current block.
[0118] When the co-located block is encoded in intra mode or palette mode, temporal motion vector prediction may not be performed based on the corresponding co-located block. On the other hand, when the co-located block is encoded in IBC mode or inter mode, temporal motion vector prediction may be performed based on the corresponding co-located block, and the motion vector of the current block may be derived based on the block vector or motion vector of the co-located block.
[0119] Alternatively, when encoding a co-located block by intra prediction based on template matching, a vector representing the position difference between the co-located block and a reference block used for template matching can be derived as a block vector. The motion vector of the current block can be derived based on the derived block vector.
[0120] Template matching-based intra prediction refers to a prediction method that calculates the cost between template regions and specifies a reference block within a search range based on the cost. The search range can be limited to the reconstruction region within the current picture. In other words, the search range can be all or part of the reconstruction region within the current picture. The region with the minimum cost within the search range can be found, and the block with a corresponding region as the template region can be specified as the reference block.
[0121] When temporal motion vector prediction is enabled / applied to the current block, the reference picture of the current block may be replaced with a co-located picture. When temporal motion vector prediction is enabled / applied, this may mean deriving the motion vector of the current block based on the temporal candidate or adding the temporal candidate to the candidate list of the current block.
[0122] Inter-frame prediction may be performed based on the motion vector of the current block to generate a prediction signal of the current block.
[0123] Figure 5 A method for performing inter prediction based on temporal motion vector prediction when the co-located block is a block coded in IBC mode is shown.
[0124] Reference Figure 5 , when the co-located block is a block encoded in IBC mode, a motion vector of the current block may be derived by scaling a block vector of the corresponding block, and inter-frame prediction may be performed based on the derived motion vector.
[0125] exist Figure 5In
[15] , CurPic, ColPic, and RefPic may refer to the current picture, the co-located picture including the co-located block used for temporal motion vector prediction, and the reference picture of the current block, respectively. CurCb and ColCb may refer to the current block and the co-located block, respectively. CurMV and ColMV may refer to the motion vector of the current block and the motion vector of the co-located block, respectively. Here, the motion vector of the current block may be derived through temporal motion vector prediction.
[0126] Information about the co-located picture may be obtained from at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH). Based on the information about the co-located picture, any one of a plurality of reference pictures belonging to a reference picture list may be determined, and the determined reference picture may be used as the co-located picture. The information about the co-located picture may include at least one of a co-located picture index, a flag indicating whether the co-located picture belongs to a reference picture list in the L0 direction, or a flag indicating whether the co-located picture belongs to a reference picture list in the L1 direction.
[0127] The encoder may determine a picture I, P, or B in the reference picture list as a co-located picture for temporal motion vector prediction, and encode information about the co-located picture to specify the co-located picture.
[0128] In temporal motion vector prediction, information of a co-located block corresponding to the current block in a co-located picture may be first checked to predict the motion vector of the current block.
[0129] When encoding a co-located block in intra mode or palette mode, temporal motion vector prediction may not be performed based on the corresponding co-located block. On the other hand, when encoding a co-located block in IBC mode or inter mode, temporal motion vector prediction may be performed based on the corresponding co-located block. The motion vector of the current block may be derived based on the block vector or motion vector of the co-located block.
[0130] Alternatively, when encoding a co-located block by intra prediction based on template matching, a vector representing the position difference between the co-located block and a reference block used for template matching can be derived as a block vector. The motion vector of the current block can be derived based on the derived block vector.
[0131] In addition, the derived motion vector may be scaled based on a predetermined scaling factor. Here, the scaling factor may be derived based on at least one of a POC difference (curPocDiff) between the current picture and the reference picture of the current block, or a POC difference (colPocDiff) between the co-located picture and the reference picture of the current block. In this case, when curPocDiff and colPocDiff are the same, scaling of the motion vector may be omitted.
[0132] A prediction signal of the current block may be generated by performing inter prediction based on a motion vector derived or scaled for the current block.
[0133] When encoding the same block in inter-frame mode, the above-mentioned Figure 3 The method performs inter-frame prediction, and when encoding the same block in IBC mode, the above-mentioned Figure 4 or Figure 5 The method performs inter-frame prediction.
[0134] Figure 6 A prediction method based on geometric segmentation according to the present disclosure is shown.
[0135] The prediction method based on geometric partitioning may be any one of the predefined inter-frame prediction modes. According to the prediction method based on geometric partitioning, a prediction block may be generated from a reference picture of a current block, and a final prediction block may be generated by a weighted sum of the prediction blocks.
[0136] Based on predetermined geometric partitioning information, the current block can be partitioned into two partitions, namely a first partition and a second partition. A first prediction block can be generated for the first partition, and a second prediction block can be generated for the second partition. Here, when the first partition has unidirectional prediction information, the first prediction block can be a block generated by unidirectional prediction. Alternatively, when the first partition has bidirectional prediction information, the first prediction block can be a block generated by bidirectional prediction. Similarly, when the second partition has unidirectional prediction information, the second prediction block can be a block generated by unidirectional prediction. Alternatively, when the second partition has bidirectional prediction information, the second prediction block can be a block generated by bidirectional prediction. Alternatively, both the first partition and the second partition can be restricted to having only unidirectional prediction information.
[0137] The final prediction block for the current block can be generated by taking a weighted sum of the first prediction block and the second prediction block. In this case, a mask defining weights for each sample position within the prediction block can be used for the weighted sum. The mask can be determined for each of the two prediction blocks. The weights can be determined based on geometric partitioning information.
[0138] The geometric segmentation information according to the present disclosure may be information for segmenting the current block into two partitions by using a straight line. As an example, the geometric segmentation information may include at least one of distance information indicating the distance between the segmentation line and the center of the current block or angle information indicating the angle of the segmentation line. The weight of the weighted sum may be determined based on at least one of the angle information or the distance information. The angle information and the distance information may be defined in the form of an index. Figures 7 to 12 The method for predicting / deriving geometric segmentation information is described in detail.
[0139] Alternatively, the geometric partitioning information may be defined as an index. Masks predefined identically for the encoder and decoder may be used according to the corresponding index. The mask may be determined for each of the two prediction blocks.
[0140] The above-mentioned mask can be defined / determined for each of the segmentation lines available for geometric segmentation. Alternatively, a mask can be defined / determined for only a portion of the available segmentation lines, and the mask for the remaining segmentation lines can be derived based on the mask for the portion of the segmentation lines. As an example, the mask for any of the remaining segmentation lines can be derived by rotating any one of the masks for the portion of the segmentation lines by a predetermined angle (e.g., 45 degrees, 90 degrees, 135 degrees, or 180 degrees). Alternatively, the mask for any of the remaining segmentation lines can be derived by inverting any one of the masks for the portion of the segmentation lines.
[0141] A flag for using a rotated mask may be encoded in the encoding device and sent to the decoding device. A flag for using an inverted mask may be encoded in the encoding device and sent to the decoding device. The flag may be encoded per coding block. The flag may indicate whether a rotated mask is used. The flag may indicate whether an inverted mask is used.
[0142] Figure 7 A method for predicting geometric segmentation information according to the present disclosure is shown.
[0143] The geometric partition information of the current block may be predicted based on the geometric partition information of neighboring blocks adjacent to the current block.
[0144] exist Figure 7 In
[15] , CurCb may be a block for which current prediction is attempted (i.e., the current block), and NbLCb and NbACb may be blocks for which prediction based on geometric partitioning is performed in blocks adjacent to the left of the current block and blocks adjacent to the top of the current block, respectively. partCand0, as geometric partitioning information of NbACb, may be used as a candidate for geometric partitioning information of the current block. In addition, partCand1, as geometric partitioning information of NbLCb, may be used as a candidate for geometric partitioning information of the current block.
[0145] Blocks using geometric partitioning-based prediction may be searched for among neighboring blocks of the current block. When at least one neighboring block uses geometric partitioning-based prediction, geometric partitioning information of the current block may be predicted based on geometric partitioning information of the neighboring blocks.
[0146] One or any one of the plurality of geometric segmentation information candidates may be configured as the geometric segmentation information for the current block. To this end, an index specifying any one of the geometric segmentation information candidates may be used. The index may be signaled via the bitstream. The index may be signaled per block (e.g., coding block).
[0147] Figure 8 A method for predicting geometric segmentation information according to the present disclosure is shown.
[0148] Blocks using geometric partitioning-based prediction may be searched for among neighboring blocks of the current block. When at least one neighboring block uses geometric partitioning-based prediction, geometric partitioning information of the current block may be predicted based on geometric partitioning information of the neighboring blocks.
[0149] Specifically, the segmentation line of the neighboring block based on the geometric segmentation information of the neighboring block can be extended to the current block, and the geometric segmentation information corresponding to the extended segmentation line can be used as the geometric segmentation information candidate of the current block. Through this process, one or more geometric segmentation information candidates can be derived.
[0150] One or any one of the plurality of geometric segmentation information candidates may be configured as the geometric segmentation information for the current block. To this end, an index specifying any one of the geometric segmentation information candidates may be used. The index may be signaled via the bitstream. The index may be signaled per block (e.g., coding block).
[0151] Figure 9 A method for predicting geometric segmentation information according to the present disclosure is shown.
[0152] Figure 9 A method for predicting geometric partitioning information when angles of partitioning lines used for geometric partitioning of adjacent blocks are the same.
[0153] Blocks using geometric partitioning-based prediction may be searched for among neighboring blocks of the current block. When at least one neighboring block uses geometric partitioning-based prediction, geometric partitioning information of the current block may be predicted based on geometric partitioning information of the neighboring blocks.
[0154] Specifically, the segmentation line of the neighboring block based on the geometric segmentation information of the neighboring block can be extended to the current block, and the geometric segmentation information corresponding to the extended segmentation line can be used as the geometric segmentation information candidate of the current block. Through this process, one or more geometric segmentation information candidates can be derived.
[0155] like Figure 9As shown, when the angles of the segmentation lines according to at least two geometric segmentation information candidates are the same, the corresponding geometric segmentation information candidates can be used as the geometric segmentation information of the current block. In this case, the signaling of the index of any one of the geometric segmentation information candidates can be omitted.
[0156] Figure 10 and Figure 11 A method for deriving geometric segmentation information according to the present disclosure is shown.
[0157] exist Figure 10 and Figure 11 In the example, partCand0 may refer to the geometric partition information of the current block predicted based on the geometric partition information of the neighboring blocks. Here, the prediction of the geometric partition information is based on the geometric partition information of the current block. Figures 7 to 9 Any of the prediction methods described.
[0158] The predicted geometric partition information may be modified based on at least one of predetermined delta angle information (delta_angle) or delta distance information (delta_distance), and the modified geometric partition information may be configured as final geometric partition information of the current block.
[0159] The incremental angle information may include at least one of absolute value information of the incremental angle or sign information of the incremental angle. The absolute value information of the incremental angle may indicate the magnitude of the rotation angle of the predicted segmentation line based on the predicted geometric segmentation information. The sign information of the incremental angle may indicate the rotation direction of the predicted segmentation line. The incremental distance information may include at least one of absolute value information of the incremental distance or sign information of the incremental distance. The absolute value information of the incremental distance may indicate the magnitude of the movement distance of the predicted segmentation line. The sign information of the incremental distance may indicate the movement direction of the predicted segmentation line. At least one of the incremental angle information or the incremental distance information may be signaled via a bitstream.
[0160] The incremental angle information may be based on the angle of the predicted partition line. The encoding device may determine the incremental angle that generates a prediction block with the minimum error among various incremental angles, encode the incremental angle, and signal the incremental angle. Similarly, the incremental distance information may be based on the distance of the predicted partition line. The encoding device may determine the incremental distance that generates a prediction block with the minimum error among various incremental distances, encode the incremental distance, and signal the incremental distance.
[0161] The incremental angle information and / or incremental distance information may be quantized to a specific number of bits. In this case, the number of bits may be signaled from the encoding device via at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH).
[0162] Figure 12 A method for predicting geometric segmentation information according to the present disclosure is shown.
[0163] Blocks using geometric partitioning-based prediction may be searched for among neighboring blocks of the current block. When at least one neighboring block uses geometric partitioning-based prediction, geometric partitioning information of the current block may be predicted based on geometric partitioning information of the neighboring blocks.
[0164] Specifically, the segmentation line of the neighboring block based on the geometric segmentation information of the neighboring block can be extended to the current block, and the geometric segmentation information corresponding to the extended segmentation line can be used as the geometric segmentation information candidate of the current block. Through this process, one or more geometric segmentation information candidates can be derived.
[0165] like Figure 12 As shown, there may be a case where the segmentation lines according to multiple geometric segmentation information candidates intersect with each other. In this case, the geometric segmentation information of the current block can be derived based on the first position where the segmentation line according to the geometric segmentation information of NbACb intersects with the current block and the second position where the segmentation line according to the geometric segmentation information of NbLCb intersects with the current block.
[0166] The various embodiments of the present disclosure do not list all possible combinations but are intended to describe representative aspects of the present disclosure, and matters described in the various embodiments may be applied independently or in combination of at least two.
[0167] In addition, various embodiments of the present disclosure may be implemented through hardware, firmware, software, or a combination thereof. For implementations through hardware, the implementations may be implemented through one or more ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), general-purpose processors, controllers, microcontrollers, microprocessors, and the like.
[0168] The scope of the present disclosure includes software or machine-executable instructions (i.e., operating systems, applications, firmware, programs, etc.) that enable the operations of the methods according to various embodiments to be performed on a device or computer, as well as non-transitory computer-readable media that store such software or instructions and are executable on a device or computer.
Claims
1. An image decoding method, comprising: determining a co-located block for temporal motion vector prediction of a current block; deriving a block vector from the co-located block; deriving a motion vector for the current block based on the block vector; as well as A prediction signal of the current block is generated by performing inter-frame prediction based on the motion vector of the current block.
2. The method according to claim 1, wherein The co-located block is a block coded in an intra block copy (IBC) mode.
3. The method according to claim 1, wherein The co-located block is a block coded in an intra mode based on template matching.
4. The method according to claim 3, wherein: The block vector is derived based on a position difference between the co-located block and a reference block used for template matching of the co-located block.
5. The method according to claim 1, wherein The reference picture of the current block is replaced by the co-located picture to which the co-located block belongs.
6. The method according to claim 1, wherein The motion vector of the current block is derived by applying a predetermined scaling factor to the block vector, and The scaling factor is derived based on at least one of a first picture order count (POC) difference between a current picture to which the current block belongs and a reference picture of the current block, or a second POC difference between a co-located picture to which the co-located block belongs and a reference picture of the current block.
7. An image encoding method, comprising: determining a co-located block for temporal motion vector prediction of a current block; deriving a block vector from the co-located block; deriving a motion vector for the current block based on the block vector; as well as A prediction signal of the current block is generated by performing inter-frame prediction based on the motion vector of the current block.
8. A method for transmitting image data, comprising: determining a co-located block for temporal motion vector prediction of a current block; deriving a block vector from the co-located block; deriving a motion vector for the current block based on the block vector; generating a prediction signal of the current block by performing inter-frame prediction based on a motion vector of the current block; deriving a residual signal of the current block based on the prediction signal; generating a bitstream by encoding the residual signal; as well as The image data including the bitstream is transmitted.