Optical flow-based inter-prediction correction method, device, and recording medium

The optical flow-based method addresses inter-screen prediction inefficiencies by calculating motion vector corrections, enhancing data compression and transmission efficiency in video encoding/decoding processes.

WO2025226084A1PCT designated stage Publication Date: 2025-10-30KWANGWOON UNIVERSITY INDUSTRY ACADEMIC COLLABORATION FOUNDATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/005647
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-25
Filing Date
2025-04-25
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing video compression technologies face challenges in effectively handling inter-screen prediction, leading to inefficiencies in data size reduction and transmission.

Method used

An optical flow-based method is employed to derive and correct inter-screen prediction information by calculating motion vector correction values using pixel gradient analysis, enabling precise prediction block generation.

Benefits of technology

This approach enhances the accuracy and efficiency of inter-screen prediction, improving data compression and transmission by optimizing motion vector correction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025005647_30102025_PF_FP_ABST
    Figure KR2025005647_30102025_PF_FP_ABST
Patent Text Reader

Abstract

An image encoding / decoding method, device, and recording medium according to the present disclosure may comprise the steps of: configuring a reference sample on the basis of inter-prediction mode information about the current block; calculating a pixel gradient value of the reference sample on the basis of the current picture including the current block and a distance value between the current block and a reference picture; calculating a motion vector correction value of the current block based on the pixel gradient value; and generating a prediction block of the current block on the basis of the motion vector correction value and the inter- prediction mode information.
Need to check novelty before this filing date? Find Prior Art

Description

Optical flow-based inter-screen prediction correction method, device, and recording medium

[0001] The present disclosure may relate to an optical flow-based inter-screen prediction correction method, device, and recording medium.

[0002] Video data is a massive data volume, and various compression techniques are being studied to reduce data size for efficient storage and transmission. Statistical analysis of redundant signals and various techniques utilizing these characteristics are being applied to video compression technologies. Furthermore, effective methods are being developed for transmitting the information required for the decoder to restore signals removed during the encoding process.

[0003] The present disclosure relates to a method for deriving a correction value of inter-screen prediction information based on inter-screen prediction inter-optical flow, and for correcting inter-screen prediction information or prediction values ​​using the derived correction value, thereby proposing an effective inter-screen prediction method.

[0004] The video encoding / decoding method, device, and recording medium of the present disclosure can comprise the steps of configuring a reference sample based on a peripheral area of ​​a prediction reference block between a current block and a screen, calculating a pixel gradient value of the reference sample, calculating a motion vector correction value based on the pixel gradient value, and correcting a motion vector of a boundary area of ​​the current block based on the motion vector correction value.

[0005] The video encoding / decoding method, device, and recording medium of the present disclosure may include: a step of configuring a reference sample based on inter-screen prediction mode information of a current block; a step of calculating a pixel gradient value of the reference sample based on a current picture including the current block and a distance value between the current block and a reference picture; a step of calculating a motion vector correction value of the current block based on the pixel gradient value; and a step of generating a prediction block of the current block based on the motion vector correction value and the inter-screen prediction mode information.

[0006] In the video encoding / decoding method, device, and recording medium of the present disclosure, the reference sample includes a first reference block determined by a motion vector for a first reference picture of the current block and a second reference block determined by a motion vector for a second reference picture of the current block, and directions of the first reference picture and the second reference picture may be different from each other.

[0007] In the video encoding / decoding method, device, and recording medium of the present disclosure, the reference sample includes a first reference block determined by a motion vector for a first reference picture of the current block and a second reference block determined by a motion vector for a second reference picture of the current block, and the directions of the first reference picture and the second reference picture may be the same.

[0008] In the video encoding / decoding method, device and recording medium of the present disclosure, the reference sample is composed of a reference block, and a motion vector that determines the reference block can be determined based on pixel information in decimal pixel units for the reference block.

[0009] In the video encoding / decoding method, device and recording medium of the present disclosure, the pixel information per decimal pixel for the reference block can be obtained by applying interpolation filtering to the reference block.

[0010] In the video encoding / decoding method, device, and recording medium of the present disclosure, the size of the reference area used for the interpolation filtering can be determined based on the size of the current block and the number of taps of the interpolation filter.

[0011] In the video encoding / decoding method, device and recording medium of the present disclosure, the size of the determined reference area may be limited to a specific size or less.

[0012] In the video encoding / decoding method, device, and recording medium of the present disclosure, the distance value between the current block and the reference picture can be calculated as a difference value between the display order of the current picture and the display order of the reference picture.

[0013] In the image encoding / decoding method, device and recording medium of the present disclosure, the pixel gradient value may include a horizontal direction gradient value and a vertical direction gradient value.

[0014] In the video encoding / decoding method, device and recording medium of the present disclosure, the pixel gradient value can be calculated using a boundary detection filter.

[0015] In the video encoding / decoding method, device and recording medium of the present disclosure, the filter size and filter coefficient of the boundary detection filter can be determined based on the size of the current block and the inter-screen prediction mode information.

[0016] In the video encoding / decoding method, device and recording medium of the present disclosure, in response to the edge detection filter being applied to pixels located at the edge of the reference sample, pixels located outside the reference sample area of ​​the reference sample can be filled through a padding operation.

[0017] The present disclosure enables effective inter-screen prediction of signals.

[0018] Figure 1 shows a block diagram of a decoder.

[0019] Figure 2 illustrates a process of generating a prediction block using an optical flow-based inter-screen prediction correction method.

[0020] Figures 3 and 4 illustrate the distance values ​​between the current picture and the reference picture.

[0021] Figure 5 shows the filter coefficients of the horizontal filter and the vertical filter.

[0022] Figure 6 illustrates one embodiment of subsampling.

[0023] FIG. 7 illustrates a process of generating a prediction block by correcting inter-screen prediction information for one of two reference pictures using an optical flow-based inter-screen prediction correction method.

[0024] Figure 8 illustrates a process of correcting a motion vector of inter-screen prediction information using an optical flow-based inter-screen prediction correction method.

[0025] Figure 9 illustrates one embodiment of a reference sample group configuration.

[0026] Figure 10 shows the surrounding lines of the current block.

[0027] Figure 11 shows the boundary area of ​​the current block.

[0028] Figure 12 shows an example of correcting the motion vector of the boundary area of ​​the current block.

[0029] FIG. 13 illustrates a method for generating an inter-screen prediction block by applying an optical flow-based inter-screen prediction correction method when the current block is in a mode of performing inter-screen prediction using a temporal interpolation block and one reference block.

[0030] FIG. 14 illustrates a method for generating an inter-screen prediction block by applying an optical flow-based inter-screen prediction correction method when the current block is in a mode that performs inter-screen prediction using a geometric segmentation mode.

[0031] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention. Throughout the description of each drawing, similar reference numerals have been used to designate similar components.

[0032] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, a first component may be referred to as a "second component," and similarly, a second component may also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.

[0033] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.

[0034] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0035] Hereinafter, with reference to the attached drawings, preferred embodiments of the present invention will be described in more detail. Hereinafter, identical components in the drawings will be designated by the same reference numerals, and redundant descriptions of identical components will be omitted.

[0036]

[0037] Hereinafter, the terms "frame" and "picture" have the same meaning and may be used interchangeably. Hereinafter, the terms "optical flow" and "optical flow" have the same meaning and may be used interchangeably. In the present invention, the order of each process in each block diagram may be changed or omitted.

[0038]

[0039] Figure 1 shows a block diagram of a decoder.

[0040] FIG. 1 is a block diagram illustrating an embodiment of the present invention for determining prediction and inverse transformation units of a decoder, and performing prediction and inverse transformation for the determined units, modes, and methods to ultimately perform restoration. The prediction unit determination unit is a step for determining the size and shape of a block for performing prediction, and may include all methods based on direct or indirect information transmitted from the encoder. This may include not only a method for direct size and shape information of the current decoded block, but also a method for utilizing information that may affect the determination of size information of the current decoded block, such as the number of divisions, depth, shape of division, direction of division, size information for the minimum division block, division information of surrounding previously decoded neighboring blocks, and prediction mode. The prediction technique determination unit is a step for determining a prediction technique for each block. The prediction technique may mean an intra-screen prediction technique, an inter-screen prediction technique, an IBC (Intra block copy) technique, a palette mode technique, a technique combining intra-screen and inter-screen prediction methods, etc. The prediction mode determination unit is a step for determining a prediction mode according to the technique determined by the prediction technique determination unit, and is a step for determining a method for determining pixel values ​​of an actual prediction block, such as a direction of prediction within a screen, the number of reference pictures in inter-screen prediction, and an affine mode and a merge mode. According to an embodiment, the prediction technique determination unit and the prediction mode determination unit may be integrated into one step and operate in the same form as the prediction technique and mode determination unit. The prediction execution unit may generate a first prediction block or blocks through the determined prediction technique and prediction mode.

[0041] The inverse transform unit determination unit determines the inverse transform unit on which the inverse transform will be performed on the residual signal. The inverse transform unit may refer to a unit that determines whether to perform inverse transform on the residual signal and transmits information about the inverse transform. The inverse transform kernel determination unit determines the kernel used for the inverse transform. At this time, at least one inverse transform kernel may be determined for one inverse transform unit. The inverse transform performing unit may perform inverse transform on the residual signal through the determined inverse transform unit and kernel. The residual signal generated by the inverse transform performing unit may be added to the prediction signal generated by the prediction performing unit to generate a restored signal. According to an embodiment, the inverse transform may be performed for N times, and the sizes of the inverse transform unit and the inverse transform kernel may not be the same. That is, the inverse transform may be performed only on some of the residual signals or transformed signals input to the inverse transform performing unit.

[0042] According to an embodiment, the process of FIG. 1 may be performed independently for a luminance block and a chrominance block, respectively, and when the color format of the input video is a YUV format (YUV420, YUV411, YUV422, YUV444, etc.), it may be performed for the luminance block and then for the chrominance block. According to an example, when the color format of the input video is RGB, color conversion to YUV may be performed and then encoding may be performed.

[0043] The following describes the contents of the present invention in detail with reference to Fig. 1.

[0044] Prediction unit decision unit

[0045] According to one embodiment, in the encoder / decoder, a prediction unit is determined by a prediction unit determination unit, and according to the embodiment, the prediction unit may be a current block, one of the sub-blocks into which the current block is divided, a set of pixels, or the value of one pixel.

[0046] A prediction unit may include size and shape information for performing predictions on a luminance component and a chrominance component. The prediction unit may be determined dependently or independently for the luminance component and the chrominance component. Dependently means that the prediction units of the luminance component or chrominance components are not determined independently for each component, but rather that once the prediction unit of one component is determined, the units of other components or components are determined with corresponding sizes and shapes. In this case, one component may correspond to one or more of the luminance component and the chrominance component. That is, the luminance component may be determined based on information on the chrominance components. In another embodiment, the prediction unit of the chrominance component may have a size corresponding to the prediction unit of the luminance component, depending on the color format of the input video or the converted color format. In the case of dependent determination, information on the prediction unit of the other component corresponding to one component, i.e., the dependently determined component, may be omitted. Independent determination means that the prediction units of the luminance component and the chrominance component are determined separately, and in the case of independent determination, information on the prediction unit for each component may be transmitted separately.

[0047] Prediction Technology Decision Division

[0048] In one embodiment, the prediction technique of each prediction unit may be determined by the prediction technique determination unit. The prediction technique may be one of inter-screen prediction, intra-screen prediction, IBC technique, palette mode technique, or a technique combining intra-screen and inter-screen prediction methods.

[0049] In an embodiment, if the current block is not an intra-screen prediction technology, a 1-bit flag may be signaled / parsed, and if the flag is Skip, the inter-screen prediction merge mode or the IBC prediction merge mode may be determined, and in this case, the inverse transformation process may be omitted, so that the prediction signal may be used as a restoration signal. Here, skip may mean that motion information (motion vector, reference picture index, reference picture list, etc.) is not transmitted, or motion information is transmitted through only one syntax information, or additionally, the differential signal of the current block is not transmitted.

[0050] In some embodiments, if the current block is not a skip, a 1-bit flag may be signaled / parsed for the prediction technique of the current block to determine one of inter-screen prediction, intra-screen prediction, IBC technique, palette mode technique, or a technique that mixes intra-screen and inter-screen prediction methods.

[0051] Prediction mode decision unit

[0052] According to one embodiment, the determination of the prediction mode is performed according to the determined prediction technique.

[0053] In an embodiment, when the prediction technique of the current prediction unit block is intra-screen prediction, the prediction mode of the current prediction unit block may be a mode that generates a prediction signal of the current prediction unit block using at least one of a directional prediction mode, a Planar mode (Horizontal Planar or Vertical Planar or Regular Planar), a DC mode, a smooth mode, a recursive prediction mode or a prediction mode based on correlation between components (e.g., CfL (Chroma from luma), CCLM (Cross component linear model), MHCCP (Multi-hypothesis cross component prediction), CCCM (Convolutional cross component model), etc.).

[0054] In an embodiment, if the prediction technique of the current prediction unit block is within-screen prediction and is predicted in a directional mode, the current prediction unit can be predicted using reference pixel information in a corrected prediction direction using additional delta angle information. The delta angle information can be signaled / parsed from the encoder and used.

[0055] According to an embodiment, when the prediction technology of the current prediction unit block is intra-screen prediction, intra-screen prediction mode information of prediction blocks existing around the current prediction block may be used to predict the intra-screen prediction mode of the current prediction unit. When the intra-screen prediction mode of the current prediction block is predicted and used from the surrounding prediction blocks, the intra-screen prediction modes of the surrounding prediction blocks may be combined to form a candidate list of intra-screen prediction modes, and the intra-screen prediction mode of the current prediction block may be determined from the candidate list formed by signaling / parsing index information of the candidate list. According to an embodiment, delta angle information may be derived and determined together with the intra-screen prediction mode from the intra-screen prediction mode candidate list. According to an embodiment, the intra-screen prediction mode candidate list may have one or more sets, and index information of a candidate list set for determining the set may be additionally signaled / parsed.

[0056] According to an embodiment, when the prediction technology of the current prediction unit block is within-screen prediction and is predicted in a directional mode, a prediction block can be generated by weighting reference pixels existing in both directions based on the prediction direction of the current directional mode. At this time, the weights used in the weighted sum can be determined according to the size of the current prediction unit, the prediction mode of the current prediction unit, the position of the pixel to be predicted, etc. According to an embodiment, when the current prediction unit is predicted in a DC mode, a prediction block can be generated by weighting reference pixels existing on the top and left sides of the prediction block generated in the DC mode. At this time, the weights used in the weighted sum can be determined according to the size of the current prediction unit, the position of the pixel to be predicted, etc.

[0057] According to an embodiment, when the prediction technique of the current prediction unit block is intra-screen prediction, the current prediction unit block may be divided into one or more sub-regions through geometric division, and a prediction signal may be generated for each region using a prediction mode within the intra-screen prediction technique including a different directional prediction mode or a Planar mode or a DC mode, and the prediction signal of the current prediction unit block may be generated through a weighted sum of the respective prediction signals, which may be an intra-screen geometric division-based prediction mode.

[0058] In an embodiment, if the prediction technology of the current prediction unit block is intra-screen prediction, it may be a matrix-based intra-screen prediction mode that performs prediction by signaling / parsing the index of the matrix using a matrix predefined by an agreement between the encoder / decoder, or by signaling / parsing the matrix.

[0059] According to an embodiment, if the prediction technology of the current prediction unit block is an intra-screen prediction, it may be an intra-screen template matching prediction mode that defines a restored area around the current prediction unit block as a template and performs template matching in the restored surrounding area of ​​the current prediction unit block to generate a prediction signal.

[0060] In an embodiment, if the prediction technique of the current prediction unit block is intra-screen prediction, the surrounding restoration area of ​​the current prediction unit block may be defined as a template, and the surrounding restoration area of ​​the current prediction unit block may be defined as a template, and the intra-screen prediction mode may be derived using the template. Thereafter, the derived mode may be used to generate the final prediction signal of the current prediction unit block. At this time, the template may also include an area that is not adjacent to the current prediction unit block.

[0061] In an embodiment, if the current prediction unit block is a chrominance block and the prediction technology of the current prediction unit block is within-screen prediction, prediction may be performed through DM (Direct mode). As an example, DM mode may be a method of performing prediction of the current chrominance block using the same prediction method as the prediction method of the luminance block at the corresponding position of the current chrominance block.

[0062] According to one embodiment, when the current prediction unit block is a chrominance block and the prediction technique of the current prediction unit block is within-screen prediction, the relationship between the surrounding restored chrominance samples of the current chrominance component and the surrounding restored luminance samples of the luminance block at the corresponding position of the current chrominance prediction unit block can be modeled as a linear or / and nonlinear model to generate a prediction signal of the current chrominance prediction unit block.

[0063] If the prediction technology of the current prediction unit is inter-screen prediction, the prediction mode of the current prediction unit block may perform motion compensation using motion information or motion information about the prediction unit and determine a signal or signals predicted through motion compensation. In addition, it may be a mode for generating a prediction signal of the current prediction unit block as a weighted sum of a plurality of prediction signals including prediction signals generated through one or more motion compensations. In this case, one or more of the prediction signals may be signals of a restored area of ​​the same frame as the current block.

[0064] According to an embodiment, if the prediction technology of the current prediction unit block is inter-screen prediction, the current prediction block may configure an inter-screen prediction information list, signal / parse an index to the list, and perform inter-screen prediction using the inter-screen prediction information corresponding to the index in the list as a predictor. The inter-screen prediction information list may be composed of inter-screen prediction information of an area spatially adjacent to the current block in the current picture, or inter-screen prediction information of an area adjacent to the position of the current block in the reference picture, or / and inter-screen prediction information of a decoded area before the current block. The inter-screen prediction information may include information such as a motion vector and a reference picture. Inter-screen prediction may be performed by signaling / parsing a motion vector predictor and a motion vector difference value derived through an index to the inter-screen prediction information list of the current prediction block. The above motion vector difference value can be signaled / parsed as an index to a motion vector difference table composed of a motion vector difference value calculated by the encoder, a motion vector difference value approximated by one of the values ​​calculated by the encoder and agreed upon between the sub- and decoders, or values ​​agreed upon between the sub- and decoders. In this case, when the current prediction unit block performs intra-screen prediction using two or more pieces of motion information, only one motion vector difference value is signaled / parsed, and the remaining motion vector difference values ​​can be derived by scaling from the parsed motion vector difference value.

[0065] In an embodiment, if the current block is not a skip and the prediction technique is determined as an inter-screen prediction or IBC technique, a 1-bit flag may be signaled / parsed to determine whether to perform prediction on the current block in merge mode or in Advanced Motion Vector Prediction (AMVP) mode. The AMVP mode may refer to a mode in which each element in the motion information is transmitted separately, and in addition to the motion information, a motion vector difference value is signaled / parsed to perform motion compensation using the corrected motion information in a subsequent prediction process.

[0066] In an embodiment, if the prediction technology of the current prediction unit block is inter-screen prediction, the current prediction unit may have a shape and size determined through geometric partitioning. In this case, the prediction signal may be a geometric partitioning-based prediction mode that generates the prediction signal of the current prediction unit block through a weighted sum of the prediction signals for each geometric partitioning unit.

[0067] In some embodiments, when the prediction technology of the current prediction unit uses a technology that combines intra-screen and inter-screen prediction methods, the surrounding restoration area of ​​the current prediction unit block may be defined as a template, and the intra-screen prediction mode may be derived using information on some or all pixels of the template. In this case, the template may include both adjacent and non-adjacent areas of the current prediction unit block, and the non-adjacent area may be an area within a certain distance of a pixel line from the current prediction unit block. When the non-adjacent area is used as a template, the distance may be transmitted from the encoder to the decoder. In some embodiments, information about the distance may be defined by an agreement between the encoder and the decoder, and the transmission of the distance information may be omitted. When defined by an agreement, the value may be fixed to a specific constant value, or may be variably determined by the horizontal and vertical pixel lengths of the prediction unit block, the block width, etc. Thereafter, the derived intra-screen prediction mode may be used to generate the final prediction signal of the current block.

[0068] Prediction Performance Department

[0069] Figure 2 illustrates a process of generating a prediction block using an optical flow-based inter-screen prediction correction method.

[0070] FIG. 2 illustrates an embodiment to which the present invention is applied, in which, when a current block is in a mode in which inter-screen prediction is performed using two reference pictures, a process of generating a prediction block by correcting inter-screen prediction information for both reference pictures through an optical flow-based inter-screen prediction correction method is illustrated.

[0071] Referring to FIG. 2, a reference sample can be configured according to inter-screen prediction mode information of the current prediction block (S200).

[0072] The above inter-screen prediction mode information may include an inter-screen prediction mode, a prediction direction flag, a reference picture index, a motion vector, etc.

[0073] The above reference samples may include a first reference sample group and a second reference sample group, and the reference sample group may be composed of reference blocks determined through motion vector information for each reference picture.

[0074] Figures 3 and 4 illustrate the distance values ​​between the current picture and the reference picture.

[0075] As an example, as shown in FIGS. 3 and 4, if the inter-screen prediction mode of the current prediction block is a mode that performs inter-screen prediction using two reference pictures (a first reference picture and a second reference picture), the reference samples may include a first reference sample group and a second reference sample group, and the first reference sample group may be composed of a reference block (a first reference block) determined through motion vector information for the first reference picture, and the second reference sample group may be composed of a reference block (a second reference block) determined through motion vector information for the second reference picture.

[0076] The above reference block may include pixel information below integer pixel units generated through an interpolation filter such as a DCT-based interpolation filter, and the motion vector may be expressed with a precision of decimal pixel units below integer pixel units in order to refer to the interpolated pixels of the reference block. At this time, when an interpolation filter is applied to generate decimal pixel unit pixel information for the reference block, a reference area for interpolation filtering may be configured based on the motion vector in the reference picture. The reference area may be configured with integer pixel information, and the size of the reference area may be determined based on the size of the current block and the number of taps of the interpolation filter. As an example, when the current block has a size of 8 × 8 and the number of taps of the interpolation filter is 15, the reference area for interpolation filtering may have a size of 23 × 23, which is the sum of the size of the current block and the number of taps of the interpolation filter. At this time, since the number of taps is an odd number, the number of reference sample lines of the upper and lower reference areas and the number of reference sample lines of the left and right may be different. For example, in an example where the current block has a size of 8 × 8 and the number of taps of the interpolation filter is 15, the number of reference sample lines in the left reference area may be 7, and the number of reference sample lines in the right reference area may be 8.

[0077] Depending on the embodiment, the size of the reference area for interpolation filtering may be limited to a specific size or less, and when the size of the reference area is limited to a specific size or less, pixel information located at the edge of the reference area may be padded and used in the interpolation filtering process.

[0078] In the example where the current block above has a size of 8 × 8 and the number of taps of the interpolation filter is 15, if the size of the reference area is limited, the number of reference sample lines in the left reference area can be limited to 3, and the number of reference sample lines in the right reference area can be limited to 4. The number of reference sample lines at the top and bottom can also be limited in the same way as above.

[0079]

[0080] Referring to FIG. 2, the distance value between the current picture and the reference picture and the pixel gradient value of the configured reference sample can be calculated (S210).

[0081] The distance value between the current picture and the reference picture can be calculated as the difference value between the display order of the current picture and the display order of the reference picture.

[0082] As an example, as in FIG. 3, when the inter-screen prediction mode of the current prediction block refers to both a past frame (the first reference picture) and a future frame (the second reference picture) based on the current frame, the distance value between the current picture and the first reference picture can be calculated as C-(C-δ1)=δ1, and the distance value between the current picture and the second reference picture can be calculated as C-(C+δ2)=-δ2, and the two distance values ​​can have different signs.

[0083] As an example, as in FIG. 4(a), when the inter-screen prediction mode of the current prediction block refers only to past frames based on the current frame, the distance value between the current picture and the first reference picture can be calculated as C-(C-δ1)=δ1, and the distance value between the current picture and the second reference picture can be calculated as C-(C-δ2)=δ2, and the two distance values ​​can have the same sign.

[0084] As an example, as in FIG. 4(b), when the inter-screen prediction mode of the current prediction block refers only to future frames based on the current frame, the distance value between the current picture and the first reference picture can be calculated as C-(C+δ1)=-δ1, and the distance value between the current picture and the second reference picture can be calculated as C-(C+δ2)=-δ2, and the two distance values ​​can have the same sign.

[0085] The pixel gradient values ​​of the above reference sample may include horizontal gradient values ​​and vertical gradient values.

[0086] Figure 5 shows the filter coefficients of the horizontal filter and the vertical filter.

[0087] The horizontal and vertical gradient values ​​of the above reference sample can be calculated for each pixel of the reference sample using an edge detection filter, such as the example of Fig. 5. At this time, the filter size and filter coefficients of the edge detection filter for calculating the gradient value can be defined according to an agreement between the image encoding device and the decoding device, or can be determined based on the size of the current prediction block and inter-screen prediction mode information.

[0088] For pixels located at the edge of the reference sample, a padding operation may be performed to fill in pixel values ​​for locations outside the reference sample area, and the gradient may be calculated using the filled pixel values. Depending on the embodiment, the padding operation may be performed by filling in pixel values ​​with 0, filling in values ​​with the nearest pixel value, etc. Here, the reference sample adjacent to the boundary of the reference area may be a reference sample adjacent to the boundary of the limited reference area when the size of the reference area is limited.

[0089] The above reference sample may include pixel information below an integer pixel unit generated through an interpolation filter such as a DCT-based interpolation filter, a bilinear interpolation filter, or a bicubic interpolation filter, and the gradient value of the reference sample may be calculated using the pixel information below an integer pixel unit. According to an embodiment, a filter combining an interpolation filter and an edge detection filter may be used to calculate the gradient value for pixel information below an integer pixel unit from the reference sample.

[0090]

[0091] Figure 6 illustrates one embodiment of subsampling.

[0092] The pixel gradient value of the above reference sample can be calculated only for the subsampled pixels after subsampling the pixels of the reference sample at a preset sampling interval, as in the example of Fig. 6. At this time, the sampling interval and the subsampling position can be defined according to an agreement between the video encoding device and the decoding device, or can be determined based on the size of the current prediction block and the inter-screen prediction mode information.

[0093] Referring to Fig. 2, a motion vector correction value can be calculated using a formula based on bidirectional optical flow (S220).

[0094] The above formula can be derived by applying the optical flow formula between reference sample groups included in the reference sample.

[0095] As an example, by applying the optical flow equations (Equations 1 and 2) between the first reference sample group and the second reference sample group, an equation such as Equation 3 can be derived.

[0096] Formula 1

[0097]

[0098] Formula 2

[0099]

[0100] Formula 3

[0101]

[0102] An equation for applying optical flow from the first reference sample group to the second reference sample group can be expressed as Equation 1. Equation 1 represents an optical flow equation for each pixel unit included in the reference sample group, and in Equation 1, gx1 and gy1 represent horizontal and vertical gradients of the first reference sample group, d1 represents a distance value between the current picture and the reference picture including the first reference sample group, p1 and p2 represent pixel values ​​of the first reference sample group and the second reference sample group, and v x , v y may mean a motion vector compensation value.

[0103] The equation for applying optical flow in the direction of the first reference sample group from the second reference sample group can be expressed as Equation 2. Equation 2 represents the optical flow equation for each pixel unit included in the reference sample group, and in Equation 2, gx2 and gy2 represent the horizontal and vertical gradients of the second reference sample group, d2 represents the distance value between the current picture and the reference picture including the second reference sample group, p1 and p2 represent the pixel values ​​of the first reference sample group and the second reference sample group, and v x , v y may mean a motion vector compensation value.

[0104] In the above formula 3, p1 and p2 represent pixel values ​​of the first reference sample group and the second reference sample group, d1 and d2 represent distance values ​​between the current picture and the reference picture including the first reference sample group and distance values ​​between the current picture and the reference picture including the second reference sample group, gx1, gy1, gx2, and gy2 represent horizontal and vertical gradients of the first reference sample group and the second reference sample group, and P diff refers to the difference in pixel values ​​between the first reference sample group and the second reference sample group, and v x , v y may mean a motion vector compensation value.

[0105] The above motion vector correction value is P in Equation 3 in units of a specific area of ​​the reference sample group. diff v that minimizes x , v y can be calculated and produced. According to an embodiment, one motion vector correction value can be produced by applying Equation 3 to the entire region of the reference sample group. According to an embodiment, each reference sample group can be divided into M sub-regions, and M motion vector correction values ​​for each sub-region can be produced by applying Equation 3 to each sub-region. At this time, if the SAD (Sum of Absolute Difference) between two reference sample groups of each sub-region is less than a specific threshold, the motion vector correction value production process can be omitted for the corresponding sub-region.

[0106] P in the above formula 3 diff v that minimizes x , v y can be calculated through an optimization algorithm.

[0107] As an example, P of Equation 3 above can be obtained through the Least Square Method as in Equation 4. diff v that minimizes x , v ycan be calculated.

[0108] Formula 4

[0109]

[0110] In the above formula 4, v x , v y is the motion vector correction value, P diff represents the difference in pixel values ​​between the first reference sample group and the second reference sample group, Ω represents a specific area of ​​the reference sample group from which a motion vector correction value is to be calculated, and i, j represent indices indicating pixel positions within a specific area of ​​the reference sample group from which a motion vector correction value is to be calculated.

[0111] The motion vector correction value may be calculated with a precision in decimal pixel units below integer pixel units. Depending on the embodiment, the precision of the motion vector correction value may be the same as or different from the precision of the motion vector included in the inter-screen prediction mode information of the current prediction block. If the precision of the motion vector correction value is calculated differently from the precision of the motion vector, the precision of the motion vector correction value and / or the precision of the motion vector may be scaled and used so that the precision of the motion vector correction value and the precision of the motion vector become the same.

[0112] The above motion vector correction value may have a limited range depending on the precision of the motion vector.

[0113] For example, if the motion vector included in the inter-screen prediction mode information of the current prediction block has a precision of 1 / 8 and the motion vector correction value is calculated to have a precision of 1 / 16, the precision of the motion vector can be scaled to 1 / 16 and used. In this case, the correction value of the motion vector can be limited to a range such as -1 to +1.

[0114] The process of calculating the motion vector compensation value through the above optical flow-based formula can be performed using pixels at the same location as the subsampled pixels for calculating the pixel gradient value of S210.

[0115] Referring to FIG. 2, an inter-screen prediction block can be generated using the inter-screen prediction mode information of the current block and the correction value of the calculated motion vector (S230).

[0116] According to an embodiment, a corrected motion vector can be derived by weighting the motion vector included in the inter-screen prediction mode information of the current block and the correction value of the calculated motion vector, and an inter-screen prediction block of the current block can be generated through a reference block referenced by the corrected motion vector.

[0117] As an example, when a motion vector correction value is derived through a formula based on optical flow between a first reference sample group and a second reference sample group, the corrected motion vectors for the first reference picture and the second reference picture can be expressed as in Formula 5, and the inter-screen prediction block of the current block generated through the corrected motion vector can be expressed as in Formula 6.

[0118] Formula 5

[0119] refine_MV1 = MV1 + (d 1· v x , d 1· v y )

[0120] refine_MV2 = MV2 + (d 2· v x , d 2· v y )

[0121] Formula 6

[0122] PredBlock inter = w 1· RefBlockrefine_MV1+ w 2· RefBlockrefine_MV2

[0123] In the above formula 5, refine_MV1 and refine MV2 represent corrected motion vectors for the first and second reference pictures, MV1 and MV2 represent motion vectors for the first and second reference pictures included in the inter-screen prediction mode information of the current block, and d1 and d2 represent distance values ​​between the current picture and the first reference picture and distance values ​​between the current picture and the second reference picture, respectively. x , v y may mean the calculated motion vector correction value. According to an embodiment, if one motion vector correction value is calculated for the entire region of the reference sample group, one corrected motion vector may be calculated for the entire region of the current block. According to an embodiment, if M motion vector correction values ​​are calculated for M sub-regions divided from each reference sample group, a corrected motion vector may be calculated for each of the M sub-blocks divided from the current block.

[0124] PredBlock in the above formula 6 inter refers to the inter-screen prediction block of the current block generated through the corrected motion vector, RefBlockrefine_MV1 and RefBlockrefine_MV2 refer to the reference blocks referenced by the corrected motion vector, and w1 and w2 refer to the weights for each reference block. At this time, a prediction block can be generated through the corrected motion vector of each sub-block in units of sub-blocks in which the corrected motion vector is calculated.

[0125] The weight for the above reference block may be determined to be the same as the weight for the reference block included in the inter-screen prediction mode information of the current block, or may be determined based on the distance between the current picture and the reference picture.

[0126] According to an embodiment, a prediction block generated through an inter-screen prediction mode of a current block can be corrected with an offset value derived through an optical flow equation and an output motion vector correction value to generate an inter-screen prediction block of the current block.

[0127] As an example, when a motion vector compensation value is derived through a formula based on optical flow between the first reference sample group and the second reference sample group, an offset value in pixel units can be derived as in formula 7, and an inter-screen prediction block of the current block can be generated as in formula 8.

[0128] Formula 7

[0129]

[0130] Formula 8

[0131] PredBlock inter (x,y) = w 1· RefBlock MV1 (x,y) + w 2· RefBlock MV2 (x,y) + w 3· offset(x,y)

[0132] In the above formula 7, offset(x,y) means the offset value of each pixel unit of the current prediction block, d1 and d2 mean the distance value between the current picture and the first reference picture and the distance value between the current picture and the second reference picture, gx1(x,y), gy1(x,y), gx2(x,y), gy2(x,y) mean the horizontal and vertical gradients of the first reference sample group and the second reference sample group for each pixel unit, and v x , v ymay mean a motion vector correction value. In some embodiments, when one motion vector correction value is calculated for the entire region of the reference sample group, the offset value for each pixel unit may be calculated using one motion vector correction value for the entire region of the current block. In some embodiments, when M motion vector correction values ​​are calculated for M sub-regions divided from each reference sample group, the offset value for each pixel unit may be calculated using each motion vector correction value in the M sub-blocks divided from the current block.

[0133] PredBlock in the above formula 8 inter (x,y) means each pixel value of the inter-screen prediction block of the current block corrected by the offset value in pixel units, and RefBlock MV1 (x,y), RefBlock MV2 (x,y) refers to a reference block referenced by a motion vector for the first reference picture and the second reference picture included in the inter-screen prediction mode information of the current block, offset(x,y) refers to an offset value for each pixel unit of the current prediction block, and w1, w2, and w3 may refer to weights for each reference block and offset value.

[0134] FIG. 7 illustrates a process of generating a prediction block by correcting inter-screen prediction information for one of two reference pictures using an optical flow-based inter-screen prediction correction method.

[0135] FIG. 7 illustrates an embodiment to which the present invention is applied, in which, when a current block is in a mode in which inter-screen prediction is performed using two reference pictures, a process of generating a prediction block by correcting inter-screen prediction information for one of two reference pictures through an optical flow-based inter-screen prediction correction method is illustrated.

[0136] Referring to FIG. 7, a reference sample can be configured according to the inter-screen prediction mode information of the current prediction block (S700).

[0137] The above inter-screen prediction mode information may include an inter-screen prediction mode, a prediction direction flag, a reference picture index, a motion vector, etc.

[0138] The above reference samples may include a first reference sample group and a second reference sample group, and the reference sample group may be composed of reference blocks determined through motion vector information for each reference picture.

[0139] As an example, as shown in FIGS. 3 and 4, if the inter-screen prediction mode of the current prediction block is a mode that performs inter-screen prediction using two reference pictures (a first reference picture and a second reference picture), the reference samples may include a first reference sample group and a second reference sample group, and the first reference sample group may be composed of a reference block (a first reference block) determined through motion vector information for the first reference picture, and the second reference sample group may be composed of a reference block (a second reference block) determined through motion vector information for the second reference picture.

[0140] The above reference block may include pixel information below integer pixel units generated through an interpolation filter such as a DCT-based interpolation filter, and the motion vector may be expressed with a precision of decimal pixel units below integer pixel units in order to refer to interpolated pixels of the reference block. At this time, when an interpolation filter is applied to generate decimal pixel unit pixel information for the reference block, a reference area for interpolation filtering may be configured based on the motion vector in the reference picture. The reference area may be configured with integer pixel information, and the size of the reference area may be determined based on the size of the current block and the number of taps of the interpolation filter. As an example, when the current block has a size of 8 × 8 and the number of taps of the interpolation filter is 15, the reference area for interpolation filtering may have a size of 23 × 23, which is the sum of the size of the current block and the number of taps of the interpolation filter. In some embodiments, the size of the reference area for interpolation filtering may be limited to a specific size or less, and when limited to a specific size or less of the reference area, pixel information located at the edge of the reference area may be padded and used in the interpolation filtering process. In this case, since the number of taps is an odd number, the number of reference sample lines of the upper and lower reference areas and the number of reference sample lines of the left and right areas may be different. For example, in an example where the current block has a size of 8 × 8 and the number of taps of the interpolation filter is 15, the number of reference sample lines of the left reference area may be 7, and the number of reference sample lines of the right reference area may be 8.

[0141]

[0142] Referring to Fig. 7, the pixel gradient value of the configured reference sample can be calculated. (S710)

[0143] The process of calculating the gradient value of the above reference sample may be performed only for one of the first reference sample group and the second reference sample group. According to an embodiment, when correcting the inter-screen prediction information of the first reference picture among the two reference pictures, the process of calculating the gradient value of the reference sample may be performed only for the second reference sample group. According to an embodiment, when correcting the inter-screen prediction information of the second reference picture among the two reference pictures, the process of calculating the gradient value of the reference sample may be performed only for the first reference sample group.

[0144] Among the first reference sample group and the second reference sample group, the reference sample group to be subject to inter-screen prediction information correction may be determined by signaling / parsing a 1-bit flag, or may be determined based on the inter-screen prediction information of the current block. According to an embodiment, a reference sample group for a reference picture that does not include a motion vector differential value in the inter-screen prediction information of the current block may be determined as a target of inter-screen prediction information correction. For example, if the inter-screen prediction information of the current block includes both a motion vector predictor and a motion vector differential value for the first reference picture, and only a motion vector predictor for the second reference picture, the second reference sample group may be determined as a reference sample group to be subject to inter-screen prediction information correction. According to an embodiment, if the inter-screen prediction information of the current block signals / parses only one motion vector differential value, and the remaining motion vector differential values ​​are derived by scaling from the parsed motion vector differential values, a reference sample group for a reference picture that uses the signaled / parsed motion vector differential value may be determined as a target of inter-screen prediction information correction. For example, if the inter-screen prediction information of the current block is signaled / parsed by receiving a motion vector difference value for a first reference picture, and the motion vector difference value for a second reference picture is derived by scaling the motion vector difference value for the first reference picture, the second reference sample group can be determined as the reference sample group that is the target of inter-screen prediction information correction.

[0145] The pixel gradient values ​​of the above reference sample may include horizontal gradient values ​​and vertical gradient values.

[0146] The horizontal and vertical gradient values ​​of the above reference sample can be calculated for each pixel of the reference sample using an edge detection filter, such as the example of Fig. 5. At this time, the filter size and filter coefficients of the edge detection filter for calculating the gradient value can be defined according to an agreement between the image encoding device and the decoding device, or can be determined based on the size of the current prediction block and inter-screen prediction mode information.

[0147] For pixels located at the edge of the reference sample, a padding operation may be performed to fill in pixel values ​​for locations outside the reference sample area, and the gradient may be calculated using the filled pixel values. Depending on the embodiment, the padding operation may be performed by filling in pixel values ​​with 0, filling in values ​​with the nearest pixel value, etc. Here, the reference sample adjacent to the boundary of the reference area may be a reference sample adjacent to the boundary of the limited reference area when the size of the reference area is limited.

[0148] The above reference sample may include pixel information below an integer pixel unit generated through an interpolation filter such as a DCT-based interpolation filter, a bilinear interpolation filter, or a bicubic interpolation filter, and the gradient value of the reference sample may be calculated using the pixel information below an integer pixel unit. According to an embodiment, a filter combining an interpolation filter and an edge detection filter may be used to calculate the gradient value for pixel information below an integer pixel unit from the reference sample.

[0149] The pixel gradient value of the above reference sample can be calculated for the subsampled pixels after subsampling the pixels of the reference sample at a preset sampling interval, as in the example of Fig. 6. At this time, the sampling interval and the subsampling position can be defined according to an agreement between the image encoding device and the decoding device, or can be determined based on the size of the current prediction block and the inter-screen prediction mode information.

[0150] Referring to Fig. 7, a motion vector correction value can be calculated using a unidirectional optical flow equation. (S720)

[0151] When a pixel gradient value for the second reference sample group is calculated in S710, a motion vector correction value for the first reference picture can be calculated through the motion vector correction value calculation process, and when a pixel gradient value for the first reference sample group is calculated in S710, a motion vector correction value for the second reference picture can be calculated through the motion vector correction value calculation process.

[0152] The above optical flow equation can be derived by applying optical flow between groups of reference samples included in the reference sample.

[0153] As an example, when correcting inter-screen prediction information of a second reference picture among two reference pictures, an equation for applying optical flow from the first reference sample group toward the second reference sample group can be expressed as Equation 9.

[0154] Formula 9

[0155]

[0156] The above equation 9 represents the optical flow equation for each pixel unit included in the reference sample group, and in the equation 9, gx and gy represent the horizontal and vertical gradients of the first reference sample group, p1 and p2 represent the pixel values ​​of the first reference sample group and the second reference sample group, and v x, v y may mean a motion vector compensation value.

[0157] The above motion vector compensation value is calculated by using the optical flow equation of Equation 9 for each specific area of ​​the reference sample group v x , v y can be calculated. According to an embodiment, one motion vector correction value can be calculated by applying Equation 9 to the entire region of the reference sample group. According to an embodiment, each reference sample group can be divided into M sub-regions, and M motion vector correction values ​​for each sub-region can be calculated by applying Equation 9 to each sub-region. At this time, if the SAD (Sum of Absolute Difference) between two reference sample groups of each sub-region is less than a specific threshold, the motion vector correction value calculation process can be omitted for the corresponding sub-region.

[0158] The optical flow equation in Equation 9 above can be calculated through an optimization algorithm.

[0159] As an example, the optical flow equation of Equation 9 can be calculated using the Least Square Method, as in Equation 10.

[0160] Formula 10

[0161]

[0162]

[0163] In the above formula 10, v x , v y represents a motion vector compensation value, f represents an optical flow equation, Ω represents a specific area of ​​a reference sample group from which a motion vector compensation value is to be calculated, and i and j represent indices indicating pixel locations within a specific area of ​​a reference sample group from which a motion vector compensation value is to be calculated.

[0164] The motion vector correction value may be calculated with a precision in decimal pixel units below integer pixel units. Depending on the embodiment, the precision of the motion vector correction value may be the same as or different from the precision of the motion vector included in the inter-screen prediction mode information of the current prediction block. If the precision of the motion vector correction value is calculated differently from the precision of the motion vector, the precision of the motion vector correction value and / or the precision of the motion vector may be scaled and used so that the precision of the motion vector correction value and the precision of the motion vector become the same.

[0165] The above motion vector correction value may have a limited range depending on the precision of the motion vector.

[0166] For example, if the motion vector included in the inter-screen prediction mode information of the current prediction block has a precision of 1 / 8 and the motion vector correction value is calculated to have a precision of 1 / 16, the precision of the motion vector can be scaled to 1 / 16 and used. In this case, the correction value of the motion vector can be limited to a range such as -1 to +1.

[0167] The process of calculating the motion vector compensation value through the above optical flow equation can be performed using pixels at the same location as the subsampled pixels for calculating the pixel gradient value of S710.

[0168] Referring to Fig. 7, an inter-screen prediction block can be generated using the inter-screen prediction mode information of the current block and the correction value of the calculated motion vector. (S730)

[0169] According to an embodiment, a corrected motion vector can be derived by weighting the motion vector included in the inter-screen prediction mode information of the current block and the correction value of the calculated motion vector, and an inter-screen prediction block of the current block can be generated through a reference block referenced by the corrected motion vector.

[0170] As an example, if a motion vector correction value for the second reference picture among two reference pictures is calculated, the corrected motion vector for the second reference picture can be expressed as in Equation 11, and the inter-screen prediction block of the current block generated through the corrected motion vector can be expressed as in Equation 12.

[0171] Equation 11

[0172] refine_MV2 = MV2 + (v x ,v y )

[0173] Equation 12

[0174] PredBlock inter = w 1· RefBlock MV1 + w 2· RefBlockrefine_MV2

[0175] In the above equation 11, refine_MV2 means the corrected motion vector for the second reference picture, MV2 means the motion vector for the second reference picture included in the inter-screen prediction mode information of the current block, and v x ,v y may mean a calculated motion vector correction value. In some embodiments, when one motion vector correction value is calculated for the entire region of the reference sample group, one corrected motion vector may be calculated for the entire region of the current block. In some embodiments, when M motion vector correction values ​​are calculated for M sub-regions divided from each reference sample group, a corrected motion vector may be calculated for each of the M sub-blocks divided from the current block.

[0176] PredBlock in the above formula 12 inter means the inter-screen prediction block of the current block generated through the compensated motion vector, and RefBlock MV1refers to a reference block referenced by a motion vector for a first reference picture included in the inter-screen prediction mode information of the current block, RefBlockrefine_MV2 refers to a reference block referenced by a corrected motion vector for a second reference picture, and w1 and w2 refer to weights for each reference block. At this time, a prediction block can be generated through the corrected motion vector of each sub-block in units of sub-blocks in which the corrected motion vector is calculated.

[0177]

[0178] Figure 8 illustrates a process of correcting a motion vector of inter-screen prediction information using an optical flow-based inter-screen prediction correction method.

[0179] FIG. 8 illustrates an example of an application of the present invention, which illustrates a process of correcting a motion vector of inter-screen prediction information using an optical flow-based inter-screen prediction correction method.

[0180] Referring to FIG. 8, a reference sample can be configured according to the inter-screen prediction mode information of the current prediction block (S800).

[0181] The above inter-screen prediction mode information may include an inter-screen prediction mode, a prediction direction flag, a reference picture index, a motion vector, etc.

[0182] The above reference samples may include a first reference sample group and a second reference sample group, and the reference sample group may be composed of a surrounding area adjacent to the current block and a surrounding area adjacent to the reference block determined through motion vector information for the reference picture.

[0183] Figure 9 illustrates one embodiment of a reference sample group configuration.

[0184] As an example, as shown in FIG. 9, the first reference sample group may be composed of a peripheral area adjacent to the current block, and the second reference sample group may be composed of a peripheral area adjacent to the first reference block.

[0185] Figure 10 shows the surrounding lines of the current block.

[0186] The above peripheral area may be determined based on the size and aspect ratio of the current block. For example, as shown in FIG. 10, the peripheral area may be composed of X vertical lines adjacent to the left side of the current block, or Y horizontal lines adjacent to the top of the current block, or / and an area of ​​size X×Y adjacent to the top left side of the current block.

[0187] Referring to FIG. 8, the pixel gradient value of the configured reference sample can be calculated (S810).

[0188] The process of calculating the gradient value of the above reference sample can be performed only for the reference sample group consisting of the surrounding area of ​​the current block.

[0189] The pixel gradient values ​​of the above reference sample may include horizontal gradient values ​​and vertical gradient values.

[0190] The horizontal and vertical gradient values ​​of the above reference sample can be calculated for each pixel of the reference sample using an edge detection filter, such as the example of Fig. 5. At this time, the filter size and filter coefficients of the edge detection filter for calculating the gradient value can be defined according to an agreement between the image encoding device and the decoding device, or can be determined based on the size of the current prediction block and inter-screen prediction mode information.

[0191] For pixels located at the edge of the reference sample, a padding operation may be performed to fill in pixel values ​​for locations outside the reference sample area, and the gradient may be calculated using the filled pixel values. Depending on the embodiment, the padding operation may be performed by a method of filling pixel values ​​with 0, a method of filling values ​​with the nearest pixel value, etc. Here, the reference sample adjacent to the boundary of the reference area may be a reference sample adjacent to the boundary of the limited reference area when the size of the reference area is limited.

[0192] The above reference sample may include pixel information below an integer pixel unit generated through an interpolation filter such as a DCT-based interpolation filter, a bilinear interpolation filter, or a bicubic interpolation filter, and the gradient value of the reference sample may be calculated using the pixel information below an integer pixel unit. According to an embodiment, a filter combining an interpolation filter and an edge detection filter may be used to calculate the gradient value for pixel information below an integer pixel unit from the reference sample.

[0193] The pixel gradient value of the above reference sample can be calculated for the subsampled pixels after subsampling the pixels of the reference sample at a preset sampling interval, as in the example of Fig. 6. At this time, the sampling interval and the subsampling position can be defined according to an agreement between the image encoding device and the decoding device, or can be determined based on the size of the current prediction block and the inter-screen prediction mode information.

[0194] Referring to Fig. 8, a motion vector correction value can be calculated using a unidirectional optical flow equation (S820).

[0195] The above optical flow equation can be derived by applying optical flow between a reference sample group consisting of a surrounding area of ​​a current block and a reference sample group consisting of a surrounding area of ​​a reference block.

[0196] As an example, an equation for applying optical flow from a first reference sample group consisting of a surrounding area of ​​a current block toward a second reference sample group consisting of a surrounding area of ​​a reference block can be expressed as Equation 13 above.

[0197] Equation 13

[0198]

[0199] The above equation 13 represents the optical flow equation for each pixel unit included in the reference sample group, and in the equation 13, gx and gy represent the horizontal and vertical gradients of the first reference sample group consisting of the surrounding area of ​​the current block, p1 and p2 represent the pixel values ​​of the second reference sample group consisting of the surrounding area of ​​the first reference sample group and the reference block, and v x , v y may mean a motion vector compensation value.

[0200] The above motion vector compensation value is calculated by calculating the optical flow equation of Equation 13 for each specific area of ​​the reference sample group v x , v y can be calculated. According to an embodiment, one motion vector correction value can be calculated by applying Equation 13 to the entire region of the reference sample group. According to an embodiment, each reference sample group can be divided into M sub-regions, and M motion vector correction values ​​for each sub-region can be calculated by applying Equation 13 to each sub-region. At this time, if the SAD (Sum of Absolute Difference) between two reference sample groups of each sub-region is less than a specific threshold, the motion vector correction value calculation process can be omitted for the corresponding sub-region.

[0201] The optical flow equation in Equation 13 above can be calculated through an optimization algorithm.

[0202] As an example, the optical flow equation of Equation 13 can be calculated using the Least Square Method, as in Equation 14.

[0203] Equation 14

[0204]

[0205]

[0206] In the above formula 14, v x , v y represents a motion vector compensation value, f represents an optical flow equation, Ω represents a specific area of ​​a reference sample group from which a motion vector compensation value is to be calculated, and i and j represent indices indicating pixel locations within a specific area of ​​a reference sample group from which a motion vector compensation value is to be calculated.

[0207] The above motion vector correction value can be calculated with a precision in decimal pixel units below integer pixel units. Depending on the embodiment, the precision of the motion vector correction value can be the same as or different from the precision of the motion vector included in the inter-screen prediction mode information of the current block. If the precision of the motion vector correction value is calculated differently from the precision of the motion vector, the precision of the motion vector correction value and / or the precision of the motion vector can be scaled and used so that the precision of the motion vector correction value and the precision of the motion vector become the same.

[0208] The above motion vector correction value may have a limited range depending on the precision of the motion vector.

[0209] For example, if the motion vector included in the inter-screen prediction mode information of the current prediction block has a precision of 1 / 8 and the motion vector correction value is calculated to have a precision of 1 / 16, the precision of the motion vector can be scaled to 1 / 16 and used. In this case, the correction value of the motion vector can be limited to a range such as -1 to +1.

[0210] The process of calculating the motion vector compensation value through the above optical flow equation can be performed using pixels at the same location as the subsampled pixels for calculating the pixel gradient value of S810.

[0211] Referring to FIG. 8, the motion vector of the boundary region of the current block can be corrected using the calculated motion vector correction value (S830). For example, the motion vector of the boundary region of the current block can represent the motion vector of a region (or sub-region) adjacent to the boundary of the current block and located within the current block.

[0212] Figure 11 shows the boundary area of ​​the current block.

[0213] The above boundary area may be an area within the current block adjacent to the reference sample group for which the motion vector compensation value is calculated, as shown in FIG. 11. For example, as shown in FIG. 11, the boundary area may be composed of I vertical lines at the left edge within the current block and / or J horizontal lines at the top edge within the current block. In this case, the size of the boundary area may be determined according to the size and aspect ratio of the current block.

[0214] Figure 12 shows an example of correcting the motion vector of the boundary area of ​​the current block.

[0215] The motion vector of the above boundary region can be corrected using the motion vector correction values ​​calculated from the reference sample group regions adjacent to the left, top, and upper left. In this case, if the reference sample group is divided into M sub-regions and M motion vector correction values ​​are calculated for each sub-region, as in the example of Fig. 12, the boundary region can be divided into K sub-regions, and the motion vector for each of the divided K sub-regions can be corrected using the motion vector correction values ​​of the sub-regions of the adjacent reference sample group.

[0216] A prediction block can be generated using an existing motion vector included in the inter-screen prediction mode information of the current prediction block and a motion vector of a boundary area corrected by a motion vector correction value for each sub-area. At this time, a filtering operation can be performed on the sub-block boundary in units of sub-blocks in the generated prediction block. According to an embodiment, a gradient can be obtained for each sub-block boundary, and if the gradient is greater than a specific threshold, a filtering operation can be performed using a smoothing filter. According to an embodiment, different filtering operations can be performed on the boundary between prediction sub-blocks generated using the motion vector of the corrected boundary area and the boundary between the prediction sub-block generated using the existing motion vector and the prediction sub-block generated using the motion vector of the corrected boundary area. As an example, a stronger smoothing operation can be performed on the boundary between the prediction sub-block generated using the existing motion vector and the prediction sub-block generated using the motion vector of the corrected boundary area.

[0217] If the current block is in a mode that performs inter-screen prediction using two or more reference blocks, the motion vector of the boundary area of ​​the current block can be corrected by applying the process of Fig. 8 between the current block and each reference block.

[0218] As an example, if the current block is in a mode of performing inter-screen prediction using three reference blocks (a first reference block, a second reference block, and a third reference block), the motion vector can be corrected using a weighted sum of a motion vector correction value generated through the process of FIG. 8 (the first process) between the current block and the first reference block, a motion vector correction value generated through the process of FIG. 8 (the second process) between the current block and the second reference block, and a motion vector correction value generated through the process of FIG. 8 (the third process) between the current block and the third reference block.

[0219] FIG. 13 illustrates a method for generating an inter-screen prediction block by applying an optical flow-based inter-screen prediction correction method when the current block is in a mode of performing inter-screen prediction using a temporal interpolation block and one reference block.

[0220] FIG. 13 illustrates an embodiment to which the present invention is applied, in which a method for generating an inter-screen prediction block is applied by applying the optical flow-based inter-screen prediction correction method of FIG. 2 and FIG. 7 when the current block is in a mode for performing inter-screen prediction using a temporal interpolation block and one reference block.

[0221] The above temporal interpolation block generates a temporal interpolation motion vector by interpolating or extrapolating a motion vector of a first temporal interpolation picture block that refers to a second temporal interpolation picture block among reference pictures included in the inter-screen prediction information of the current frame based on a distance between pictures, and can generate inter-screen prediction by referencing the first temporal interpolation reference block and the second temporal interpolation reference block through the temporal interpolation motion vector.

[0222] The above picture-to-picture distance can be calculated as the difference between the display order of the current picture and the display order of the temporal interpolation reference picture.

[0223] The inter-screen prediction process of FIG. 13 can be composed of a process (first process) of generating a temporal interpolation block through the optical flow-based inter-screen prediction correction method of FIG. 2 and a process (second process) of generating an inter-screen prediction block of the current block through the optical flow-based inter-screen prediction correction method of FIG. 7.

[0224] The above first process can generate a temporal interpolation block using the optical flow-based inter-screen prediction correction method of FIG. 2 between the first temporal interpolation reference block and the second temporal interpolation reference block. At this time, the distance value between the current picture and the reference picture of S210 can be calculated as the display order difference value (C-T1) between the current picture and the first temporal interpolation reference picture and the display order difference value (C-T2) between the current picture and the second temporal interpolation reference picture.

[0225] The second process may generate an inter-screen prediction block of the current block using the optical flow-based inter-screen prediction correction method of FIG. 7 between the temporal interpolation block generated in the first process and the first reference block. At this time, the pixel gradient calculation process of S710 may be performed only for the temporal interpolation block, and accordingly, only the motion vector correction value for the first reference block may be calculated in S720.

[0226] FIG. 14 illustrates a method for generating an inter-screen prediction block by applying an optical flow-based inter-screen prediction correction method when the current block is in a mode that performs inter-screen prediction using a geometric segmentation mode.

[0227] The above geometric segmentation mode can divide the current block into five sub-regions along geometric segmentation lines, and generate an inter-screen prediction block of the current block using an inter-screen prediction block generated using N reference blocks as reference blocks for each sub-region.

[0228] The above geometric dividing line can be defined according to the angle value formed by the dividing line and the horizontal or vertical line and the distance value between the current block center and the dividing line.

[0229] Figure 14 is an example in which the current block is divided into two sub-regions according to a geometric dividing line, and inter-screen prediction blocks (first prediction block and second prediction block) generated using each of two reference blocks are used as reference pictures for each sub-region.

[0230] The reference pictures (first reference picture, second reference picture, third reference picture, fourth reference picture) including the reference blocks (first reference block, second reference block, third reference block, fourth reference block) of the first prediction block and the second prediction block may be the same or different pictures.

[0231] The inter-screen prediction process of FIG. 14 may be composed of a process (first process) of generating a first prediction block through the optical flow-based inter-screen prediction correction method of FIG. 2 or FIG. 7, a process (second process) of generating a second prediction block through the optical flow-based inter-screen prediction correction method of FIG. 2 or FIG. 7, and a process (third process) of generating an inter-screen prediction block of a geometric division mode using the first prediction block and the second prediction block.

[0232] The above first process can generate a first prediction block using the optical flow-based inter-screen prediction correction method of FIG. 2 or FIG. 7 between a first reference block and a second reference block. When the method of FIG. 2 is used, the distance value between the current picture and the reference picture of S210 can be calculated as the display order difference value (C-R1) between the current picture and the first reference picture and the display order difference value (C-R2) between the current picture and the second reference picture. When the method of FIG. 7 is used, the pixel gradient calculation process of S710 can be performed only for the first reference block or the second reference block, and accordingly, only the motion vector correction value for the second reference block or the first reference block can be calculated in S720.

[0233] The second process may generate a second prediction block using the optical flow-based inter-screen prediction correction method of FIG. 2 or FIG. 7 between the third reference block and the fourth reference block. When the method of FIG. 2 is used, the distance value between the current picture and the reference picture of S210 may be calculated as the display order difference value (C-R3) between the current picture and the third reference picture and the display order difference value (C-R4) between the current picture and the fourth reference picture. When the method of FIG. 7 is used, the pixel gradient calculation process of S710 may be performed only for the third reference block or the fourth reference block, and accordingly, only the motion vector correction value for the fourth reference block or the third reference block may be calculated in S720.

[0234] In the third process, the first prediction block generated in the first process and the second prediction block generated in the second process are used as reference pictures for each sub-region divided along a geometric dividing line to perform inter-screen prediction, thereby generating an inter-screen prediction block of the current block.

[0235] Inverse transformation unit decision unit

[0236] According to an embodiment, the inverse transform unit determination unit may be a single transform unit (TU) or a sub-block each of which is a single TU divided into multiple units in the decoder / decoder.

[0237] Inverse transform kernel decision unit

[0238] According to an embodiment, the inverse transform kernel determining unit may determine separable vertical and horizontal first-order inverse transform kernels and / or non-separable second-order inverse transform kernels, or may determine a non-separable first-order inverse transform kernel.

[0239] Inverse transformation execution unit

[0240] According to an embodiment, the inverse transform performing unit can perform inverse transform on the inverse quantized transform coefficients using the inverse transform unit determined by the inverse transform unit determining unit and the inverse transform kernel determined by the inverse transform kernel determining unit.

[0241] Depending on the embodiment, whether to perform a non-separable first-order inverse transform and whether to perform a non-separable second-order inverse transform can be determined through signaling / parsing, or can be determined implicitly based on the size of the current TU, etc.

[0242] Transform coefficient entropy decoding unit

[0243] According to an embodiment, the transform coefficient entropy decoding unit can restore the quantized second transform coefficient when the second transform is applied through an entropy decoding method (e.g., Context-Adaptive Binary Arithmetic Coding (CABAC), Context-Adaptive Variable-Length Coding (CAVLC)), and can perform restoration of the quantized first transform coefficient when the second transform is not applied.

[0244] Transformation coefficient inverse quantization part

[0245] According to an embodiment, the transform coefficient dequantization unit can parse information such as a quantization method and quantization parameters, and perform dequantization on the restored transform coefficients to obtain dequantized transform coefficients.

[0246] Restoration Department

[0247] According to an embodiment, the inverse transform performing unit can generate a final restored signal by combining the final predicted signal generated through the prediction performing unit and the restored residual signal.

[0248]

[0249] While the exemplary methods of this disclosure are presented as a series of operations for clarity of description, this is not intended to limit the order in which the steps are performed, and individual steps may be performed simultaneously or in different orders, if desired. To implement a method according to this disclosure, additional steps may be included in addition to the steps illustrated, some steps may be excluded and the remaining steps included, or some steps may be excluded and additional steps included.

[0250] The various embodiments of the present disclosure are not intended to list all possible combinations but rather to illustrate representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combinations of two or more.

[0251] Additionally, various embodiments of the present disclosure may be implemented by hardware, firmware, software, or a combination thereof. In the case of hardware implementation, the embodiments may be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.

[0252] The scope of the present disclosure may include software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium having such software or instructions stored thereon and executable on the device or computer.

[0253] The present disclosure may be applicable to industrial fields utilizing an optical flow-based inter-screen prediction correction method, device, and recording medium.

Claims

1. A step of constructing a reference sample based on the inter-screen prediction mode information of the current block; A step of calculating a pixel gradient value of the reference sample based on a distance value between the current picture including the current block and the reference picture: A step of calculating a motion vector correction value of the current block based on the pixel gradient value; and A video decoding method, characterized by comprising a step of generating a prediction block of the current block based on the motion vector correction value and the inter-screen prediction mode information.

2. In paragraph 1, The above reference sample includes a first reference block determined by a motion vector for a first reference picture of the current block and a second reference block determined by a motion vector for a second reference picture of the current block, An image decoding method, characterized in that the directions of the first reference picture and the second reference picture are different from each other.

3. In paragraph 1, The above reference sample includes a first reference block determined by a motion vector for a first reference picture of the current block and a second reference block determined by a motion vector for a second reference picture of the current block, A video decoding method, characterized in that the directions of the first reference picture and the second reference picture are the same.

4. In paragraph 1, The above reference sample consists of a reference block, A video decoding method, characterized in that the motion vector that determines the reference block is determined based on pixel-by-pixel information for the reference block.

5. In paragraph 4, An image decoding method, characterized in that the pixel unit pixel information for the above reference block is obtained by applying interpolation filtering to the above reference block.

6. In paragraph 5, An image decoding method, characterized in that the size of the reference area used for the interpolation filtering is determined based on the size of the current block and the number of taps of the interpolation filter.

7. In paragraph 6, An image decoding method, characterized in that the size of the above-determined reference area is limited to a specific size or less.

8. In paragraph 1, A video decoding method, characterized in that the distance value between the current block and the reference picture is calculated as a difference value between the display order of the current picture and the display order of the reference picture.

9. In paragraph 1, An image decoding method, characterized in that the pixel gradient value includes a horizontal gradient value and a vertical gradient value.

10. In paragraph 9, An image decoding method, characterized in that the above pixel gradient value is calculated using a boundary detection filter.

11. In paragraph 10, An image decoding method, characterized in that the filter size and filter coefficient of the boundary detection filter are determined based on the size of the current block and the inter-screen prediction mode information.

12. In paragraph 11, A method for decoding an image, characterized in that, in response to the edge detection filter being applied to pixels located at the edge of the reference sample, pixels located outside the reference sample area of ​​the reference sample are filled through a padding operation.

13. A step of constructing a reference sample based on the inter-screen prediction mode information of the current block; A step of calculating a pixel gradient value of the reference sample based on a distance value between the current picture including the current block and the reference picture: A step of calculating a motion vector correction value of the current block based on the pixel gradient value; and A video encoding method, characterized by comprising a step of generating a prediction block of the current block based on the motion vector correction value and the inter-screen prediction mode information.

14. In a non-transitory computer-readable recording medium storing a bitstream generated by a video encoding method, The above image encoding method is, A step of constructing a reference sample based on inter-screen prediction mode information of the current block; A step of calculating a pixel gradient value of the reference sample based on a distance value between the current picture including the current block and the reference picture: A step of calculating a motion vector correction value of the current block based on the pixel gradient value; and A non-transitory computer-readable recording medium, characterized in that it comprises a step of generating a prediction block of the current block based on the motion vector correction value and the inter-screen prediction mode information.

Citation Information

Patent Citations

  • Inter prediction mode-based image processing method and apparatus therefor

    KR1020180043787A

  • Active mixer and method for improving noise and gain

    KR1020220034536A

  • Mission level autonomous system and method thereof for unmanned aerial vehicle

    KR1020230001668A

  • Energy saving system and method for Cargo Hold within ventilation system

    KR1020240045630A

  • KR20240025058A