Video decoding method and apparatus as well as video encoding method and apparatus
By using pixel and gradient values within reference blocks for bi-directional motion prediction, the method addresses inefficiencies in existing codecs, enhancing encoding/decoding efficiency and reducing memory access and complex operations.
Patent Information
- Application Number
- JP2025064119
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-01-04
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2038-01-04
AI Technical Summary
Existing video codecs face inefficiencies in encoding and decoding high-resolution video due to the need to access pixel and gradient values outside reference blocks during bi-directional motion prediction, leading to increased memory access and complex multiplication operations.
The method utilizes pixel and gradient values only from within the reference blocks to determine displacement vectors, applying interpolation and gradient filters to fractional pixel positions, minimizing memory access and complex operations.
This approach enhances encoding/decoding efficiency by predicting pixel values similar to the original block, reducing memory access and complex operations, thus improving video processing speed and efficiency.
Smart Images

Figure 2025103014000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video decoding method and video encoding, and more particularly, to video decoding and video encoding that perform inter prediction in a bi-directional motion prediction mode.
Background Art
[0002] With the development and spread of hardware that can play and store high-resolution or high-quality video content, the need for a video codec that can effectively encode and decode high-resolution or high-quality video content is increasing. According to existing video codecs, video is encoded by a limited encoding method based on a tree-structured encoding unit.
[0003] Using frequency conversion, video data in the spatial domain is converted into coefficients in the frequency domain. For rapid calculation of frequency conversion, a video codec divides a video into blocks of a predetermined size, performs DCT conversion for each block, and encodes frequency coefficients in units of blocks. Compared with video data in the spatial domain, the coefficients in the frequency domain have a form that is easily compressed. In particular, through inter prediction or intra prediction of a video codec, video pixel values in the spatial domain are expressed as prediction errors, so if frequency conversion is performed on the prediction errors, much data is also converted to 0. The video codec reduces the amount of data by replacing continuously repeating data with small-size data.
Summary of the Invention
Problems to be Solved by the Invention
[0004] According to various embodiments, in the bidirectional motion prediction mode, not only the pixel values of the first reference block of the first reference picture and the pixel values of the second reference block of the second reference picture are used, but also the first gradient value of the first reference block and the second gradient value of the second reference block are both used to generate the predicted pixel value of the current block. Therefore, since a prediction block similar to the original block can also be generated, the encoding / decoding efficiency can be increased.
[0005] The pixel value of the first reference block, the pixel value of the second reference block, the first gradient value of the first reference block, and the second gradient value of the second reference block are used to determine the displacement vector in the horizontal or vertical direction of the current block when performing motion compensation in pixel group units. In particular, in order to determine the displacement vector in the horizontal or vertical direction of the current pixel in the current block, not only the pixel value and gradient value of the first reference pixel in the first reference block corresponding to the current pixel and the pixel value and gradient value of the second reference pixel in the second reference block are used, but also the pixel values and gradient values of the peripheral pixels included in a window of a predetermined size centered on the first reference pixel and the second reference pixel are used. Therefore, when the current pixel is located at the boundary, since the peripheral pixels of the reference pixel corresponding to the current pixel are located outside the reference block, there is a problem that additional memory access must be made because the pixel values and gradient values of the pixels located outside the reference block must be referred to.
[0006] According to various embodiments, by referring only to the pixel values and gradient values of the pixels located inside the reference block without referring to the pixel values and gradient values stored for the pixels located outside the reference block and determining the displacement vector in the horizontal or vertical direction of the current block, the number of memory accesses can be minimized.
[0007] According to various embodiments, instead of using the pixel values of integer pixels as inputs to determine the gradient values in the horizontal or vertical direction of a reference pixel and using horizontal and vertical gradient filters and interpolation filters, an interpolation filter is applied to the pixel values of integer pixels to determine the pixel values of pixels at positions in units of fractional pixels, and a horizontal or vertical gradient filter with a relatively short filter length is applied to the pixel values of pixels at positions in units of fractional pixels to determine the gradient values in the horizontal or vertical direction of the reference pixel, thereby minimizing the performance of complex multiplication operations.
[0008] According to various embodiments, by performing motion compensation in units of pixel groups, it is possible to minimize the performance of even more complex multiplication operations when performing motion compensation in units of pixels.
[0009] A computer-readable recording medium on which a program for implementing the method according to various embodiments is recorded may be included.
[0010] Here, the technical problems of various embodiments are not limited to the features mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art from the following description.
Means for Solving the Problems
[0011] The technical problems of the present invention are not limited to the features mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art from the following description.
[0012] A video decoding method according to various embodiments includes a step of obtaining motion prediction mode information related to a current block in a current picture from a bitstream, and When the obtained motion prediction mode information indicates a bi-direction motion prediction mode, in the first reference picture, information about a first motion vector indicating a first reference block of the current block, and in the second reference picture, information about a second motion vector indicating a second reference block of the current block are obtained from the bitstream; Using values of reference pixels included in the first reference block and the second reference block without using values stored for pixels located outside the boundaries of the first reference block and the second reference block, determining a displacement vector per unit time in the horizontal or vertical direction of pixels of the current block including pixels adjacent to the inside of the boundary of the current block; Based on a gradient value in the horizontal or vertical direction of a first corresponding reference pixel in the first reference block corresponding to a current pixel included in a current pixel group in the current block, a gradient value in the horizontal or vertical direction of a second corresponding reference pixel in the second reference block corresponding to the current pixel, a pixel value of the first corresponding reference pixel, a pixel value of the second corresponding reference pixel, and a displacement vector per unit time in the horizontal or vertical direction of the current pixel, performing motion compensation in block units and motion compensation in pixel group units of the current block to obtain a predicted block of the current block; Obtaining information about a residual block of the current block from the bitstream; Restoring the current block based on the predicted block and the residual block, including: The pixel group may include at least one pixel.
[0013] In a video decoding method according to various embodiments, the step of obtaining a predicted block of the current block includes: Applying a gradient filter in the horizontal or vertical direction to the pixel values of the fractional - position pixels within the first reference block or the second reference block, and calculating the horizontal or vertical gradient values of the first corresponding reference pixels or the second corresponding reference pixels, further comprising the step of: The gradient filter is a 5 - tap filter. The fractional - position pixel is also a pixel in which at least one of the horizontal component or the vertical component of the coordinates indicating the position of the pixel has a fractional value.
[0014] In a video decoding method according to various embodiments, the pixel values of the fractional - position pixels within the first reference block or the second reference block may also be calculated by applying an interpolation filter in the horizontal or vertical direction to the pixel values of the integer - position pixels.
[0015] In a video decoding method according to various embodiments, the size of the pixel group is also determined based on the minimum value of the height and width of the current block.
[0016] In a video decoding method according to various embodiments, the displacement vector per unit time in the horizontal or vertical direction for the current pixel group is also a displacement vector per unit time determined by using values based on the pixel values and gradient values of the first corresponding reference pixels and their surrounding pixels included in the first corresponding reference pixel group within the first reference picture corresponding to the current pixel group, the second corresponding reference pixels and their surrounding pixels included in the second corresponding reference pixel group within the second reference picture, the first POC (picture order count) difference between the first reference picture and the current picture, and the second POC difference between the second reference picture and the current picture.
[0017] In a video decoding method according to various embodiments, the step of determining the displacement vector per unit time in the horizontal or vertical direction of the pixels of the current block is: When the first corresponding reference pixel or the second corresponding reference pixel is a boundary peripheral pixel adjacent to the inside of the boundary of the first reference block or the second reference block, using the pixel value of the boundary peripheral pixel to derive the pixel value and gradient value of a pixel located outside the boundary of the first reference block or the second reference block; determining a displacement vector per unit time in the horizontal or vertical direction of the current pixel based on the pixel value and gradient value of the boundary peripheral pixel, and the pixel value and gradient value of a pixel located outside the boundary of the current block derived using the pixel value of the boundary peripheral pixel. This may include these steps.
[0018] In a video decoding method according to various embodiments, the step of determining a displacement vector per unit time in the horizontal or vertical direction of a pixel of the current block using the pixel value of the first corresponding reference pixel included in the first reference block, the pixel value of the second corresponding reference pixel included in the second reference block, the gradient value of the first corresponding reference pixel, the gradient value of the second corresponding reference pixel, the first POC difference between the first reference picture and the current picture, and the second POC difference between the second reference picture and the current picture to calculate a value related to the current pixel; using the pixel value of the first corresponding peripheral pixel of the first corresponding reference pixel, the pixel value of the second corresponding peripheral pixel of the second corresponding reference pixel, the gradient value of the first corresponding peripheral pixel, the gradient value of the second corresponding peripheral pixel, the first POC difference between the first reference picture and the current picture, and the second POC difference between the second reference picture and the current picture to calculate a value related to a peripheral pixel calculated thereby; using the value related to the current pixel, the value related to the peripheral pixel, and a weighting value to calculate a weighted average value for the current pixel necessary to calculate a displacement vector per unit time in the horizontal or vertical direction; Using the calculated weighted average value for the current pixel, determining a displacement vector per unit time in the horizontal or vertical direction of the current pixel, may be included.
[0019] In a video decoding method according to various embodiments, the weighted average value for the current pixel is also a value calculated by applying an exponential smoothing technique in the up, down, left, and right directions to values related to pixels included in the first reference block and the second reference block.
[0020] A video decoding apparatus according to various embodiments includes obtaining motion prediction mode information related to a current block in a current picture from a bitstream. When the obtained motion prediction mode information indicates a bidirectional motion prediction mode, in a first reference picture, information about a first motion vector indicating a first reference block of the current block, and in a second reference picture, information about a second motion vector indicating a second reference block of the current block are obtained from the bitstream, and an obtaining unit that obtains information about a residual block of the current block from the bitstream. Without using the values stored for pixels located outside the boundaries of the first reference block and the second reference block, use the values related to the reference pixels included in the first reference block and the second reference block, determine the displacement vector per unit time in the horizontal or vertical direction of the pixels of the current block including the pixels adjacent to the inside of the boundary of the current block, and based on the horizontal or vertical gradient value of the first corresponding reference pixel in the first reference block corresponding to the current pixel included in the current pixel group in the current block, the horizontal or vertical gradient value of the second corresponding reference pixel in the second reference block corresponding to the current pixel, the pixel value of the first corresponding reference pixel, the pixel value of the second corresponding reference pixel, and the displacement vector per unit time in the horizontal or vertical direction of the current pixel, perform block unit motion compensation and pixel group unit motion compensation related to the current block, and an inter prediction unit that obtains a predicted block of the current block, a decoding unit that restores the current block based on the predicted block and the residual block, and the pixel group may include at least one pixel.
[0021] In a video decoding apparatus according to various embodiments, the inter prediction unit applies a horizontal or vertical gradient filter to the pixel value of the fractional position pixel in the first reference block or the second reference block to calculate the horizontal or vertical gradient value of the first corresponding reference pixel or the second corresponding reference pixel, the gradient filter is a 5-tap filter, the fractional position pixel is also a pixel in which at least one of the horizontal component or the vertical component of the coordinate indicating the position of the pixel has a fractional value.
[0022] In a video decoding apparatus according to various embodiments, the inter prediction unit The displacement vector per unit time in the horizontal or vertical direction for the current pixel group is also a displacement vector per unit time determined by using values determined based on pixel values and gradient values of first corresponding reference pixels and their peripheral pixels included in a first corresponding reference pixel group in a first reference picture corresponding to the current pixel group, second corresponding reference pixels and their peripheral pixels included in a second corresponding reference pixel group in a second reference picture, a first POC difference between the first reference picture and the current picture, and a second POC difference between the second reference picture and the current picture.
[0023] In a video decoding apparatus according to various embodiments, the inter prediction unit uses the pixel value of the first corresponding reference pixel included in the first reference block, the pixel value of the second corresponding reference pixel included in the second reference block, the gradient value of the first corresponding reference pixel, the gradient value of the second corresponding reference pixel, the first POC difference between the first reference picture and the current picture, and the second POC difference between the second reference picture and the current picture to calculate a value related to the current pixel. uses the pixel value of the first corresponding peripheral pixel of the first corresponding reference pixel, the pixel value of the second corresponding peripheral pixel of the second corresponding reference pixel, the gradient value of the first corresponding peripheral pixel, the gradient value of the second corresponding peripheral pixel, the first POC difference between the first reference picture and the current picture, and the second POC difference between the second reference picture and the current picture to calculate a value related to the peripheral pixels calculated thereby. uses the value related to the current pixel, the value related to the peripheral pixels, and a weighting value to calculate a weighted average value for the current pixel necessary for calculating a displacement vector per unit time in the horizontal or vertical direction. Using the calculated weighted average value for the current pixel, the displacement vector per unit time in the horizontal or vertical direction of the current pixel can be determined.
[0024] Video encoding methods according to various embodiments perform motion compensation at the block unit and motion compensation at the pixel group unit for the current block, and obtain a predicted block, a first motion vector, and a second motion vector of the current block; generating a bitstream including information about the first motion vector and the second motion vector, and motion prediction mode information indicating whether the motion prediction mode related to the current block is a bidirectional motion prediction mode; The pixel group includes at least one pixel; The first motion vector is a motion vector indicating a first reference block of a first reference picture corresponding to the current block in the current picture from the current block; The second motion vector is a motion vector indicating a second reference block of a second reference picture corresponding to the current block in the current picture from the current block; Based on the horizontal or vertical gradient value of the first corresponding reference pixel in the first reference block corresponding to the current pixel included in the current pixel group in the current block, the horizontal or vertical gradient value of the second corresponding reference pixel in the second reference block corresponding to the current pixel, the pixel value of the first corresponding reference pixel, the pixel value of the second corresponding reference pixel, and the displacement vector per unit time in the horizontal or vertical direction of the current pixel, motion compensation at the block unit and motion compensation at the pixel group unit related to the current block are performed, and a predicted block of the current block is obtained. The displacement vector per unit time in the horizontal or vertical direction of the pixels of the current block including pixels adjacent to the inside of the boundary of the current block is determined using the values related to the reference pixels included in the first reference block and the second reference block without using the values stored for the pixels located outside the boundaries of the first reference block and the second reference block.
[0025] The video encoding device according to various embodiments performs motion compensation in units of blocks and motion compensation in units of pixel groups on a current block, and an inter prediction unit that obtains a prediction block, a first motion vector, and a second motion vector of the current block; a bitstream generation unit that generates a bitstream including information about the first motion vector and the second motion vector, and motion prediction mode information indicating whether or not the motion prediction mode related to the current block is a bidirectional motion prediction mode; the pixel group includes at least one pixel; the first motion vector is a motion vector indicating a first reference block of a first reference picture corresponding to the current block in the current picture from the current block, and the second motion vector is a motion vector indicating a second reference block of a second reference picture corresponding to the current block in the current picture from the current block; Based on the gradient value in the horizontal or vertical direction of the first corresponding reference pixel in the first reference block corresponding to the current pixel included in the current pixel group in the current block, the gradient value in the horizontal or vertical direction of the second corresponding reference pixel in the second reference block corresponding to the current pixel, the pixel value of the first corresponding reference pixel, the pixel value of the second corresponding reference pixel, and the displacement vector per unit time in the horizontal or vertical direction of the current pixel, motion compensation in units of blocks and motion compensation in units of pixel groups related to the current block are performed, and a prediction block of the current block is obtained; The displacement vector per unit time in the horizontal or vertical direction of the pixels of the current block including the pixels adjacent to the inside of the boundary of the current block is determined using the values related to the reference pixels included in the first reference block and the second reference block without using the values stored for the pixels located outside the boundaries of the first reference block and the second reference block. A computer-readable recording medium on which a program for implementing the method according to various embodiments is recorded may be included.
Advantages of the Invention
[0026] According to various embodiments, in the bidirectional motion prediction mode, by using the gradient value of the reference block of the reference picture to perform inter prediction related to the current block and predicting a value similar to the value of the original block of the current block, the encoding / decoding efficiency can be improved.
Brief Description of the Drawings
[0027]
Figure 1A
Figure 1B
Figure 1C
Figure 1D
Figure 1E
Figure 1F
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 7A
Figure 7B
Figure 7C
Figure 7D
Figure 7E
Figure 8A
Figure 8B
Figure 8C
Figure 8D
Figure 9A
Figure 9B
Figure 9C
Figure 9D
Figure 9E
Figure 9F
Figure 9G
Figure 9H
Figure 9I
Figure 9J
Figure 9K
Figure 9L
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Embodiments for Carrying Out the Invention
[0028] The video decoding method according to various embodiments includes the steps of obtaining motion prediction mode information related to a current block in a current picture from a bitstream, and when the obtained motion prediction mode information indicates a bi-directional motion prediction mode, obtaining, from the bitstream, information about a first motion vector indicating a first reference block of the current block in a first reference picture and information about a second motion vector indicating a second reference block of the current block in a second reference picture; determining a displacement vector per unit time in a horizontal or vertical direction of pixels of the current block including pixels adjacent to the inside of the boundary of the current block by using values related to reference pixels included in the first reference block and the second reference block without using values stored for pixels located outside the boundaries of the first reference block and the second reference block; performing block unit motion compensation and pixel group unit motion compensation of the current block based on a horizontal or vertical gradient value of a first corresponding reference pixel in the first reference block corresponding to a current pixel included in a current pixel group in the current block, a horizontal or vertical gradient value of a second corresponding reference pixel in the second reference block corresponding to the current pixel, a pixel value of the first corresponding reference pixel, a pixel value of the second corresponding reference pixel, and a displacement vector per unit time in a horizontal or vertical direction of the current pixel to obtain a predicted block of the current block; obtaining information about a residual block of the current block from the bitstream; restoring the current block based on the predicted block and the residual block, and the pixel group may include at least one pixel.
[0029] The video decoding apparatus according to various embodiments includes the steps of obtaining motion prediction mode information related to a current block in a current picture from a bitstream, and When the obtained motion prediction mode information indicates a bidirectional motion prediction mode, in the first reference picture, information about a first motion vector indicating a first reference block of a current block, and in the second reference picture, information about a second motion vector indicating a second reference block of the current block are obtained from the bitstream, and an acquisition unit that obtains information about a residual block of the current block from the bitstream; Without using the values stored for pixels located outside the boundaries of the first reference block and the second reference block, using the values related to the reference pixels included in the first reference block and the second reference block, determining a displacement vector per unit time in the horizontal or vertical direction of the pixels of the current block including the pixels adjacent to the inside of the boundary of the current block, and based on the horizontal or vertical gradient value of the first corresponding reference pixel in the first reference block corresponding to the current pixel included in the current pixel group in the current block, the horizontal or vertical gradient value of the second corresponding reference pixel in the second reference block corresponding to the current pixel, the pixel value of the first corresponding reference pixel, the pixel value of the second corresponding reference pixel, and the displacement vector per unit time in the horizontal or vertical direction of the current pixel, performing block unit motion compensation and pixel group unit motion compensation related to the current block, and obtaining a predicted block of the current block; an inter prediction unit; A decoding unit that restores the current block based on the predicted block and the residual block, and The pixel group may include at least one pixel.
[0030] A video encoding method according to various embodiments includes performing block unit motion compensation and pixel group unit motion compensation on a current block, and obtaining a predicted block, a first motion vector, and a second motion vector of the current block; generating a bitstream including information about the first motion vector and the second motion vector, and motion prediction mode information indicating whether the motion prediction mode related to the current block is a bidirectional motion prediction mode; The pixel group includes at least one pixel. The first motion vector is a motion vector indicating a first reference block of a first reference picture corresponding to the current block in the current picture from the current block. The second motion vector is a motion vector indicating a second reference block of a second reference picture corresponding to the current block in the current picture from the current block. Based on the horizontal or vertical gradient value of the first corresponding reference pixel in the first reference block corresponding to the current pixel included in the current pixel group in the current block, the horizontal or vertical gradient value of the second corresponding reference pixel in the second reference block corresponding to the current pixel, the pixel value of the first corresponding reference pixel, the pixel value of the second corresponding reference pixel, and the displacement vector per unit time in the horizontal or vertical direction of the current pixel, block-based motion compensation and pixel group-based motion compensation related to the current block are performed, and a predicted block of the current block is obtained. The displacement vector per unit time in the horizontal or vertical direction of the pixels of the current block including the pixels adjacent to the inside of the boundary of the current block is determined using the values related to the reference pixels included in the first reference block and the second reference block without using the values stored for the pixels located outside the boundaries of the first reference block and the second reference block.
[0031] An inter prediction unit according to various embodiments performs block-based motion compensation and pixel group-based motion compensation on a current block, and obtains a predicted block, a first motion vector, and a second motion vector of the current block. A bitstream generation unit that generates a bitstream including information about the first motion vector and the second motion vector, and motion prediction mode information indicating whether the motion prediction mode related to the current block is a bidirectional motion prediction mode. The pixel group includes at least one pixel. The first motion vector is a motion vector indicating a first reference block of a first reference picture corresponding to the current block in the current picture from the current block, and the second motion vector is a motion vector indicating a second reference block of a second reference picture corresponding to the current block in the current picture from the current block. Based on the gradient value in the horizontal or vertical direction of the first corresponding reference pixel in the first reference block corresponding to the current pixel included in the current pixel group in the current block, the gradient value in the horizontal or vertical direction of the second corresponding reference pixel in the second reference block corresponding to the current pixel, the pixel value of the first corresponding reference pixel, the pixel value of the second corresponding reference pixel, and the displacement vector per unit time in the horizontal or vertical direction of the current pixel, block-based motion compensation and pixel group-based motion compensation related to the current block are performed, and a predicted block of the current block is obtained. The displacement vector per unit time in the horizontal or vertical direction of the pixels of the current block including the pixels adjacent to the inside of the boundary of the current block is determined by using the values related to the reference pixels included in the first reference block and the second reference block without using the values stored for the pixels located outside the boundaries of the first reference block and the second reference block.
[0032] It may include a computer-readable recording medium on which a program for implementing a method according to various embodiments is recorded.
[0033] Hereinafter, "video" can indicate a still image or a moving image of a video, that is, the video itself.
[0034] Hereinafter, a "sample" is data assigned to a sampling position of a video, and means data to be processed. For example, in a video of a spatial region, a pixel is also a sample.
[0035] Hereinafter, a "current block" means a block of a video to be encoded or decoded.
[0036] FIG. 1A illustrates a block diagram of a video decoding apparatus according to various embodiments.
[0037] A video decoding apparatus 100 according to various embodiments includes an acquisition unit 105, an inter prediction unit 110, and a restoration unit 125.
[0038] The acquisition unit 105 receives a bitstream including information about a prediction mode of the current block, information indicating a motion prediction mode of the current block, and information about a motion vector.
[0039] The acquisition unit 105 can acquire information about a prediction mode of the current block, information indicating a motion prediction mode of the current block, and information about a motion vector from the received bitstream. Also, the acquisition unit 105 can acquire a reference picture index indicating a reference picture among previously decoded pictures from the received bitstream.
[0040] When the prediction mode of the current block is the inter prediction mode, the inter prediction unit 110 performs inter prediction related to the current block. That is, the inter prediction unit 110 can generate the predicted pixel value of the current block by using at least one of the pictures decoded before the current picture including the current block. For example, when the motion prediction mode of the current block is the bidirectional motion prediction mode, the inter prediction unit 110 can generate the predicted pixel value of the current block by using two pictures decoded before the current picture. That is, when the information about the motion prediction mode obtained from the bitstream indicates the bidirectional motion prediction mode, the inter prediction unit 110 can generate the predicted pixel value of the current block by using two pictures decoded before the current picture.
[0041] The inter prediction unit 110 may include a motion compensation unit 115 in block units and a motion compensation unit 120 in pixel group units.
[0042] The motion compensation unit 115 in block units can perform motion compensation in block units related to the current block.
[0043] The motion compensation unit 115 in block units can determine at least one reference picture of the pictures decoded before by using the reference picture index obtained from the bitstream. At this time, the reference picture index means the reference picture index related to each of the prediction directions including the L0 direction and the L1 direction. Here, the reference picture index related to the L0 direction means the index indicating the reference picture among the pictures included in the L0 reference picture list, and the reference picture index related to the L1 direction means the index indicating the reference picture among the pictures included in the L1 reference picture list.
[0044] The block-based motion compensation unit 115 can determine a reference block of a current block located in at least one reference picture by using information about motion vectors received from a bit stream. Here, the corresponding block in the reference picture corresponding to the current block in the current picture is also the reference block. That is, the block-based motion compensation unit 115 can determine the reference block of the current block by using a motion vector indicating the reference block from the current block. Here, the motion vector means a vector indicating the displacement between the reference coordinates of the current block in the current picture and the reference coordinates of the reference block in the reference picture. For example, when the upper left coordinates of the current block are (1, 1) and the upper left coordinates of the reference block in the reference picture are (3, 3), the motion vector is also (2, 2).
[0045] Here, the information about the motion vector may include a differential value of the motion vector, and the block-based motion compensation unit 115 uses a predictor of the motion vector and the differential value of the motion vector obtained from the bit stream to restore the motion vector, and uses the restored motion vector to determine a reference block of the current block located in at least one reference picture. At this time, the differential value of the motion vector means the differential value of the motion vector related to the reference picture related to each prediction direction including the L0 direction and the L1 direction. Here, the differential value of the motion vector related to the L0 direction means the differential value of the motion vector indicating the reference block in the reference picture included in the L0 reference picture list, and the differential value of the motion vector related to the L1 direction means the differential value of the motion vector indicating the reference block in the reference picture included in the L1 reference picture list.
[0046] The block-based motion compensation unit 115 can perform block-based motion compensation for the current block by using the pixel values of the reference blocks. The block-based motion compensation unit 115 can perform block-based motion compensation for the current block by using the reference pixel value in the reference block corresponding to the current pixel in the current block. Here, the reference pixel is a pixel included in the reference block, and the corresponding pixel corresponding to the current pixel in the current block is also the reference pixel.
[0047] The block-based motion compensation unit 115 can perform block-based motion compensation for the current block by using a plurality of reference blocks included in each of a plurality of reference pictures. For example, when the motion prediction mode of the current block is the bidirectional motion prediction mode, the block-based motion compensation unit 115 can determine two reference pictures among the previously encoded pictures and determine two reference blocks included in the two pictures.
[0048] The block-based motion compensation unit 115 can perform block-based motion compensation for the current block by using the pixel values of two reference pixels in the two reference blocks. The block-based motion compensation unit 115 can perform block-based motion compensation for the current block by using the average value or weighted sum related to the pixel values of the two reference pixels, and generate a block-based motion compensation value.
[0049] The reference position of the reference block is not only the position of an integer pixel but also the position of a fractional pixel. Here, an integer pixel is a pixel whose position component is an integer, meaning a pixel at an integer pixel position. A fractional pixel is a pixel whose position component is a fraction, meaning a pixel at a fractional pixel position.
[0050] For example, when the upper left coordinate of the current block is (1, 1) and the motion vector is (2.5, 2.5), the upper left coordinate of the reference block in the reference picture is also (3.5, 3.5). At this time, the position of the fractional pixel is also determined in units of 1 / 4 pel (pel: pixel element) or 1 / 16 pel. It is not limited thereto, and the position of the fractional pixel may also be determined by various fractional pel units.
[0051] When the reference position of the reference block is the position of a fractional pixel, the block-based motion compensation unit 115 applies an interpolation filter to a first peripheral region including a first pixel among the pixels of the first reference block indicated by the first motion vector and a second peripheral region including a second pixel among the pixels of the second reference block indicated by the second motion vector, and can generate a pixel value of the first pixel and a pixel value of the second pixel.
[0052] That is, the reference pixel value in the reference block is also determined by using the pixel values of the peripheral pixels whose predetermined direction components are integers. At this time, the predetermined direction is also the horizontal direction or the vertical direction.
[0053] For example, the block-based motion compensation unit 115 performs filtering on the pixel values of pixels whose predetermined direction components are integers by using an interpolation filter, determines the reference pixel value with the resulting value, and can determine the block-based motion compensation value related to the current block by using the reference pixel value. The motion compensation value in block units can also be determined by using the average value or weighted sum of the reference pixels. At this time, the interpolation filter can use an M-tap interpolation filter based on DCT (discrete cosine transformation). The coefficients of the M-tap interpolation filter based on DCT are also derived from DCT and IDCT (inverse discrete cosine transform). At this time, the coefficients of the interpolation filter are also filter coefficients scaled to integer coefficients in order to reduce real number operations during filtering. At this time, the interpolation filter is also a one-dimensional interpolation filter in the horizontal or vertical direction. For example, when expressing the position of a pixel with x and y orthogonal coordinate components, the horizontal direction means a direction parallel to the x-axis. The vertical direction means a direction parallel to the y-axis.
[0054] The block-based motion compensation unit 115 first uses a one-dimensional interpolation filter in the vertical direction to perform filtering on the pixel values at integer positions, and then uses a one-dimensional interpolation filter in the horizontal direction to perform filtering on the value generated by the filtering, so as to determine the reference pixel value at the fractional pixel position.
[0055] On the other hand, when using the scaled filter coefficients, the value generated by the filtering is larger than when using the unscaled filter coefficients. Therefore, the block-based motion compensation unit 115 can perform de-scaling on the value generated by the filtering.
[0056] The block-based motion compensation unit 115 can perform inverse scaling after filtering the pixel values at integer positions using a one-dimensional interpolation filter in the vertical direction. At this time, the inverse scaling may include bit-shifting to the right by the number of inverse scaling bits to the right. The number of inverse scaling bits is also determined based on the bit depth of the samples of the input video. For example, the number of inverse scaling bits is also the value obtained by subtracting 8 from the bit depth of the samples.
[0057] Also, the block-based motion compensation unit 115 can perform inverse scaling after filtering the pixel values at integer positions using a one-dimensional interpolation filter in the vertical direction, and then filtering the generated values using a one-dimensional interpolation filter in the horizontal direction. At this time, the inverse scaling may include bit-shifting to the right by the number of inverse scaling bits to the right. The number of inverse scaling bits is also determined based on the scaling bit number of the one-dimensional interpolation filter in the vertical direction, the scaling bit number of the one-dimensional interpolation filter in the horizontal direction, and the bit depth of the samples. For example, when the scaling bit number p of the one-dimensional interpolation filter in the vertical direction is 6, the scaling bit number q of the one-dimensional interpolation filter in the horizontal direction is 6, and the bit depth of the samples is b, the number of inverse scaling bits is also p + q + 8 - b, which is 20 - b.
[0058] If the block-based motion compensation unit 115 uses a one-dimensional interpolation filter to filter the pixels with a predetermined direction component being an integer and then only performs bit-shifting to the right by the number of inverse scaling bits, a rounding error may occur. Therefore, after using a one-dimensional interpolation filter to filter the pixels with a predetermined direction component being an integer, an offset can be added, and then inverse scaling can be performed. At this time, the offset is also 2^(number of inverse scaling bits - 1).
[0059] The motion compensation unit 120 at the pixel group level performs motion compensation at the pixel group level for the current block and can generate a motion compensation value at the pixel group level. When the motion prediction mode of the current block is the bidirectional motion prediction mode, the motion compensation unit 120 at the pixel group level can perform motion compensation at the pixel group level for the current block and generate a motion compensation value at the pixel group level.
[0060] Based on the optical flow of the pixel groups of the first reference picture and the second reference picture, the motion compensation unit 120 at the pixel group level performs motion compensation at the pixel group level for the pixel group related to the current block and can generate a motion compensation value at the pixel group level. The optical flow will be described later in the description related to FIG. 3A.
[0061] The motion compensation unit 120 at the pixel group level can perform motion compensation at the pixel group level for the pixel groups included in the reference block of the current block and generate a motion compensation value at the pixel group level. The pixel group may include at least one pixel. For example, the pixel group may be one pixel. Or, the pixel group may be a plurality of pixels including two or more pixels. The pixel group may also be a plurality of pixels included in a block of size KxK (K is an integer).
[0062] The motion compensation unit 120 at the pixel group level can determine a pixel group and perform motion compensation at the pixel group level for the pixel group related to the current block based on the determined pixel group.
[0063] The motion compensation unit 120 at the pixel group level can determine the size of the pixel group based on the size of the current block. For example, the motion compensation unit 120 at the pixel group level can determine the height and width of the pixel group as the maximum value between the value obtained by dividing the minimum value of the height and width of the current block by 8 and 2.
[0064] The motion compensation unit 120 at the pixel group level performs motion compensation in units of pixel groups each including a plurality of pixels, thereby reducing the encoding / decoding complexity as compared with performing motion compensation at the pixel level from a high video resolution. Also, the motion compensation unit 120 at the pixel group level performs motion compensation in units of pixel groups each including a plurality of pixels, thereby reducing the encoding / decoding complexity as compared with performing motion compensation at the pixel level from a high frame rate.
[0065] The acquisition unit 105 can acquire information about the size of a pixel group included in a bitstream. The information about the size of the pixel group is also information indicating the height or width K when the size of the pixel group is KxK. The information about the size of the pixel group is also included in a high level syntax carrier.
[0066] The motion compensation unit 120 at the pixel group level can determine at least one pixel group partition including pixels having similar pixel values among the plurality of pixels included in the pixel group, and perform motion compensation on the pixel group partition. At this time, since the pixel group partition including pixels having similar pixel values is highly likely to be the same object and highly likely to have similar motions, the motion compensation unit 120 at the pixel group level can perform more detailed motion compensation in units of pixel groups.
[0067] On the other hand, motion compensation at the pixel group level is performed when the motion prediction mode information indicates the bidirectional motion prediction mode, but even in that case, it is not always performed and is also selectively performed.
[0068] The motion compensation unit 120 at the pixel group level can determine the reference pixel group within the reference block corresponding to the current pixel group of the current block, and can determine the gradient value of the reference pixel group. For example, the motion compensation unit 120 at the pixel group level can utilize the gradient value of at least one pixel value included in the reference pixel group to determine the gradient value of the reference pixel group.
[0069] The motion compensation unit 120 at the pixel group level can utilize the gradient value of the reference pixel group to perform motion compensation at the pixel group level for the pixel group related to the current block, and can generate a motion compensation value at the pixel group level.
[0070] The motion compensation unit 120 at the pixel group level can apply a filter to the first peripheral region of the first pixel group including the first pixel group among the pixel groups of the first reference block indicated by the first motion vector, and the second peripheral region of the second pixel group including the second pixel group among the pixel groups of the second reference block indicated by the second motion vector, and can generate the gradient value of the first pixel group and the gradient value of the second pixel group.
[0071] The motion compensation unit 120 at the pixel group level can determine the pixel value and the gradient value of the pixels within the first window of a predetermined size including the first pixel group centered on the first pixel group within the first reference picture, and can determine the pixel value and the gradient value of the pixels within the second window of a predetermined size including the second reference pixel centered on the second reference pixel group within the second reference picture.
[0072] The motion compensation unit 120 at the pixel group level can determine the displacement vector per unit time related to the current pixel group by using the pixel values and gradient values of the pixels within the first window and the pixel values and gradient values of the pixels within the second window. At this time, the displacement vector per unit time related to the current pixel group may also have its value adjusted by a regularization parameter. The regularization parameter is a parameter introduced to prevent an error from occurring when a displacement vector per unit time related to an ill-posed current pixel group is determined in order to perform motion compensation for the pixel group. The motion compensation unit 120 at the pixel group level can perform motion compensation for the pixel group related to the current block based on the regularization parameter for the displacement vector per unit time in the horizontal or vertical direction. The regularization parameter will be described later in the explanation of FIG. 8A.
[0073] The motion compensation unit 120 at the pixel group level can perform motion compensation for the pixel group related to the current block by using the displacement vector per unit time related to the current pixel group and the gradient value of the reference pixel.
[0074] The reference position of the reference block is not only the position of an integer pixel but also not limited thereto and can be the position of a fractional pixel.
[0075] When the reference position of the reference block is the position of a fractional pixel, the gradient value of the reference pixel within the reference block is also determined by using the pixel values of the surrounding pixels where a predetermined direction component is an integer.
[0076] For example, the motion compensation unit 120 at the pixel group level can perform filtering on the pixel values of peripheral pixels whose predetermined direction components are integers using a gradient filter, and determine the gradient value of the reference pixel with the resulting value. At this time, the filter coefficients of the gradient filter can also be determined using coefficients determined in advance for the DCT-based interpolation filter. The coefficients of the gradient filter are also filter coefficients scaled to integer coefficients in order to reduce real-number operations during the execution of filtering.
[0077] At this time, the gradient filter is also a one-dimensional gradient filter in the horizontal or vertical direction.
[0078] The motion compensation unit 120 at the pixel group level can use a one-dimensional gradient filter in the horizontal or vertical direction to perform filtering on peripheral pixels whose corresponding direction components are integers in order to determine the gradient value in the horizontal or vertical direction related to the reference pixel.
[0079] For example, the motion compensation unit 120 at the pixel group level can use a one-dimensional gradient filter in the horizontal direction to perform filtering on the pixels located horizontally from the pixels whose horizontal direction components are integers among the pixels located around the reference pixel, and determine the gradient value in the horizontal direction related to the reference pixel.
[0080] If the position of the reference pixel is (x + α, y + β) (where x and y are integers and α and β are decimals), the motion compensation unit 120 at the pixel group level can perform filtering on the pixel at the (x, y) position and the pixels whose vertical components are integers among the pixels located vertically from the pixel at the (x, y) position using a one-dimensional interpolation filter in the vertical direction, and determine the pixel value of (x, y + β) with the resulting value.
[0081] The motion compensation unit 120 at the pixel group level performs filtering on the pixel value at the (x, y+β) position and the pixels among the pixels horizontally positioned from the (x, y+β) position whose horizontal components are integers, using a horizontal gradient filter, and can determine the horizontal gradient value at the (x+α, y+β) position with the resulting value.
[0082] The order of using the one-dimensional gradient filter and the one-dimensional interpolation filter is not restricted. Above, first, the vertical interpolation filter is used to perform filtering on the pixels at integer positions to generate the vertical interpolation filtering value, and then the one-dimensional horizontal gradient filter is used to perform filtering on the vertical interpolation filtering value. However, first, the one-dimensional horizontal gradient filter can be used to perform filtering on the pixels at integer positions to generate the horizontal interpolation filtering value, and then the one-dimensional vertical interpolation filter can be used to perform filtering on the horizontal interpolation filtering value.
[0083] Above, the content of the motion compensation unit 120 at the pixel group level determining the horizontal gradient value at the (x+α, y+β) position has been described in detail. Since the motion compensation unit 120 at the pixel group level also determines the vertical gradient value at the (x+α, y+β) position in a manner similar to where the horizontal gradient value is determined, the detailed description is omitted.
[0084] As described above in detail, the motion compensation unit (230) at the pixel group unit uses a one-dimensional gradient filter and a one-dimensional interpolation filter to determine the gradient value at the fractional pixel position. However, it is not limited to these, and a gradient filter and an interpolation filter can also be used to determine the gradient value at the integer pixel position. However, in the case of integer pixels, even if the interpolation filter is not used, the pixel value can be determined. For the sake of consistent processing with the processing at fractional pixels, for integer pixels and surrounding pixels where a predetermined direction component is an integer, filtering can be performed using an interpolation filter to determine the integer pixel value. For example, the interpolation filter coefficient at integer pixels is also {0, 0, 64, 0, 0}. Since the interpolation filter coefficient related to the surrounding integer pixels is 0, filtering is performed using only the pixel value of the current integer pixel. As a result, filtering is also performed on the current integer pixel and the surrounding integer pixels using an interpolation filter, and the pixel value at the current integer pixel can also be determined.
[0085] The motion compensation unit 120 at the pixel group unit can perform inverse scaling after filtering the pixels at integer positions using a one-dimensional interpolation filter in the vertical direction. At this time, the inverse scaling may include bit-shifting by the number of inverse scaling bits to the right. The number of inverse scaling bits is also determined based on the bit depth of the sample. Also, the number of inverse scaling bits is determined based on the specific input data within the block.
[0086] For example, the number of bit-shifting is also a value obtained by subtracting 8 from the bit depth of the sample.
[0087] The motion compensation unit 120 at the pixel group level can perform inverse scaling after filtering the value generated by performing inverse scaling using a horizontal gradient filter. Similarly, here, the inverse scaling may include bit-shifting to the right by the number of inverse scaling bits. The number of inverse scaling bits is also determined based on the scaled number of bits of the one-dimensional interpolation filter in the vertical direction, the scaled number of bits of the one-dimensional gradient filter in the horizontal direction, and the bit depth of the sample. For example, when the scaled number of bits p of the one-dimensional interpolation filter is 6, the scaled number of bits q of the one-dimensional gradient filter is 4, and the bit depth of the sample is b, the number of inverse scaling bits is p + q + 8 - b, which is also 18 - b.
[0088] If the motion compensation unit 120 at the pixel group level only performs bit-shifting to the right by the number of inverse scaling bits on the value generated after filtering, a rounding error may occur. Therefore, an offset can be added to the value generated after filtering, and then inverse scaling can be performed. At this time, the offset is also 2 ^ (the number of inverse scaling bits - 1).
[0089] The inter prediction unit 110 can generate a predicted pixel value of the current block by using the motion compensation value in block units and the motion compensation value in pixel group units related to the current block. For example, the inter prediction unit 110 can combine the motion compensation value in block units and the motion compensation value in pixel group units related to the current block to generate a predicted pixel value of the current block. Here, the motion compensation value in block units means a value generated by performing motion compensation in block units, and the motion compensation value in pixel group units is a value generated by performing motion compensation in pixel group units. The motion compensation value in block units is also the average value or weighted sum of reference pixels, and the motion compensation value in pixel group units is also a value determined based on the displacement vector per unit time for the current pixel and the gradient value of the reference pixels.
[0090] The motion compensation unit 120 in pixel group units can obtain a shift value for descaling after interpolation operation or gradient operation based on at least one of the bit depth of the sample, the input range of the filter used for the interpolation operation or gradient operation, and the coefficient of the filter. The motion compensation unit 120 in pixel group units can perform descaling after the interpolation operation or gradient operation related to the pixels included in the first reference block and the second reference block by using the shift value for descaling.
[0091] When performing motion compensation in block units, the inter prediction unit 110 can use a motion vector and save the motion vector. At this time, the motion vector unit is also a block of 4x4 size. On the other hand, when saving the motion vector after motion compensation in block units, the motion vector saving unit is also blocks of various sizes larger than 4x4 size (for example, a block of RxR size; R is an integer). At this time, the motion vector saving unit is also a block larger than 4x4 size. For example, it is also a block of 16x16 size.
[0092] On the one hand, when performing motion compensation in pixel group units, together with the size of the current block, based on the window size and the interpolation filter length, the size of the target block for performing motion compensation in pixel group units is also extended. The reason why the size of the target block is extended from the size of the current block based on the window size is that, in the case of pixels located at the edge of the current block, using the window, based on the pixels located at the current edge and the surrounding pixels, motion compensation in pixel group units related to the current block is performed.
[0093] Therefore, in the process of performing motion compensation in pixel group units using the window to reduce the memory access times and the execution of multiplication operations, the motion compensation unit 120 in pixel group units adjusts the positions of the pixels outside the current block among the pixels within the window to the positions of the pixels adjacent to the inside of the current block, and determines the pixel values and gradient values at the adjusted pixel positions, thereby also reducing the memory access times and the number of multiplication operations.
[0094] The motion compensation unit 120 in pixel group units does not use the pixel values of integer-position pixels to determine the gradient values of the reference pixels, which are the values necessary for motion compensation in pixel group units. That is, the motion compensation unit 120 in pixel group units applies a horizontal or vertical gradient filter to the pixel values of the fractional-position pixels to calculate the horizontal or vertical gradient values of the first corresponding reference pixels in the first reference block or the second corresponding reference pixels in the second reference block. At this time, the gradient filter length is also 5. At this time, the coefficients of the filter can have coefficients symmetric about the central coefficient of the filter. The fractional-position pixels are also pixels in which at least one of the horizontal and vertical components indicating the position of the pixel has a fractional value.
[0095] The pixel values of the fractional - position pixels within the first reference block or the second reference block are also calculated by applying an interpolation filter in the horizontal or vertical direction to the pixel values of the integer - position pixels.
[0096] The displacement vector per unit time in the horizontal or vertical direction for the current pixel group is also a displacement vector per unit time determined by using values determined based on the first corresponding reference pixels included in the first corresponding reference pixel group within the first reference picture corresponding to the current pixel group, the second corresponding reference pixels included in the second corresponding reference pixel group within the second reference picture, and the pixel values and gradient values of their surrounding pixels, the first POC (picture order count) difference between the first reference picture and the current picture, and the second POC difference between the second reference picture and the current picture.
[0097] When the first corresponding reference pixel or the second corresponding reference pixel is a boundary - peripheral pixel adjacent to the inside of the boundary of the first reference block or the second reference block, the pixel value compensation unit 120 for pixel groups can use the pixel value of the boundary - peripheral pixel to derive the pixel values of the pixels located outside the boundary of the first reference block or the second reference block.
[0098] The pixel value compensation unit 120 for pixel groups can determine the displacement vector per unit time in the horizontal or vertical direction of the current block based on the pixel value of the boundary - peripheral pixel and the pixel values of the pixels located outside the boundary of the current block derived by using the pixel value of the boundary - peripheral pixel. That is, there are pixels located outside the boundary within the pixels included in a window centered on the boundary - peripheral pixel. At this time, the pixel values and gradient values of the pixels located outside the boundary are also the pixel values and gradient values of the pixels derived from the boundary - peripheral pixels that are not the values stored in the memory.
[0099] The pixel group unit compensation unit 120 can calculate a value related to the current pixel by using the pixel value of the first corresponding reference pixel included in the first reference block, the pixel value of the second corresponding reference pixel included in the second reference block, the gradient value of the first corresponding reference pixel, the gradient value of the second corresponding reference pixel, the first POC difference between the first reference picture and the current picture, and the second POC difference between the second reference picture and the current picture. That is, the value related to the current pixel is also the result value of a function based on the pixel value and gradient value of the corresponding reference pixel of each reference picture, and the POC difference between each reference picture and the current picture.
[0100] The pixel group unit compensation unit 120 can calculate a value related to the peripheral pixel by using the pixel value of the first corresponding peripheral pixel of the first corresponding reference pixel, the gradient value of the first corresponding peripheral pixel, the pixel value of the second corresponding peripheral pixel of the second corresponding reference pixel, the gradient value of the second corresponding peripheral pixel, the first POC difference between the first reference picture and the current picture, and the second POC difference between the second reference picture and the current picture. That is, the value related to the peripheral pixel is also the result value of a function based on the pixel value and gradient value of the corresponding reference pixel of each reference picture, and the POC difference between each reference picture and the current picture. That is, the value related to the peripheral pixel is also the result value of a function based on the pixel value and gradient value of the corresponding peripheral pixel of each reference picture, and the POC difference between each reference picture and the current picture.
[0101] The pixel group unit compensation unit 120 can calculate a weighted average value for the current pixel required to calculate the displacement vector per unit time in the horizontal direction by using the value related to the current pixel, the value related to the peripheral pixel, and the weighting value. At this time, the weighting value is also determined based on the distance between the current pixel and the peripheral pixel, the distance between the pixel and the block boundary, the number of pixels located outside the boundary, or whether the pixel is located inside or outside the boundary.
[0102] The weighted average value for the current pixel is also a value calculated by applying the exponential smoothing technique in the vertical, horizontal, left, and right directions to the values related to the pixels included in the first reference block and the second reference block. By applying the exponential smoothing technique in the vertical, horizontal, left, and right directions to the values related to the pixels, the value calculated for the current pixel has the largest weighted value related to the value of the current pixel, and the weighted values related to the values of the surrounding pixels are exponentially decreased values depending on the distance from the current pixel.
[0103] The pixel group unit compensation unit 120 can use the weighted average value for the current pixel to determine the displacement vector per unit time in the horizontal or vertical direction of the current pixel.
[0104] The restoration unit 125 can obtain the residual block of the current block from the bitstream and use the residual block and the prediction block of the current block to restore the current block. For example, the restoration unit 125 can combine the pixel values of the residual block of the current block and the pixel values of the prediction block of the current block from the bitstream to generate the pixel values of the restored block.
[0105] The video decoding apparatus 100 may include a video decoding unit (not shown), and the video decoding unit (not shown) may include an acquisition unit 105, an inter prediction unit 110, and a restoration unit 125. The video decoding unit will be described with reference to FIG. 1E.
[0106] FIG. 1B illustrates a flowchart of a video decoding method according to various embodiments.
[0107] In step S105, the video decoding apparatus 100 can acquire motion prediction mode information related to the current block in the current picture from the bit stream. The video decoding apparatus 100 can receive a bit stream including the motion prediction mode information related to the current block in the current picture, and can acquire the motion prediction mode information related to the current block from the received bit stream. The video decoding apparatus 100 can acquire information about the prediction mode of the current block from the bit stream, and can determine the prediction mode of the current block based on the information about the prediction mode of the current block. At this time, when the prediction mode of the current block is the inter prediction mode, the video decoding apparatus 100 can acquire the motion prediction mode information related to the current block.
[0108] For example, the video decoding apparatus 100 can determine the prediction mode of the current block as the inter prediction mode based on the information about the prediction mode of the current block. When the prediction mode of the current block is the inter prediction mode, the video decoding apparatus 100 can acquire the motion prediction mode information related to the current block from the bit stream.
[0109] In step S110, when the motion prediction mode information indicates the bi - directional motion prediction mode, the video decoding apparatus 100 can acquire, from the bit stream, a first motion vector indicating the first reference block of the current block in the first reference picture and a second motion vector indicating the second reference block of the current block in the second reference picture.
[0110] That is, the video decoding apparatus 100 can receive a bit stream including information about the first motion vector and the second motion vector, and can acquire the first motion vector and the second motion vector from the received bit stream. The video decoding apparatus 100 can acquire a reference picture index from the bit stream, and can determine the first reference picture and the second reference picture among the previously decoded pictures based on the reference picture index.
[0111] In step S115, the video decoding apparatus 100 can determine a horizontal or vertical displacement vector of the pixels of the current block including the pixels adjacent to the inside of the boundary of the current block, by using the values related to the reference pixels included in the first reference block and the second reference block, without using the values stored for the pixels located outside the boundaries of the first reference block and the second reference block. At this time, the values stored for the pixels located outside the boundaries of the first reference block and the second reference block, as well as the values related to the reference pixels included in the first reference block and the second reference block, are also the pixel values of the related pixels, or the horizontal gradient values of the related pixels, or the vertical gradient values. Or, the values stored for the pixels located outside the boundaries of the first reference block and the second reference block, as well as the values related to the reference pixels included in the first reference block and the second reference block, are also the values determined by using the pixel values or gradient values of the related pixels.
[0112] In step S120, the video decoding apparatus 100 can perform block unit motion compensation and pixel group unit motion compensation of the current block, and obtain a predicted block of the current block, by using the horizontal or vertical gradient value of the first corresponding reference pixel in the first reference block corresponding to the current pixel included in the current pixel group in the current block, the horizontal or vertical gradient value of the second corresponding reference pixel in the second reference block, the pixel value of the first corresponding reference pixel, the pixel value of the second corresponding reference pixel, and the horizontal or vertical displacement vector of the current pixel.
[0113] That is, the video decoding apparatus 100 can perform block-based motion compensation and pixel group-based motion compensation on the current block based on the first motion vector and the second motion vector, and generate a predicted block of the current block. The video decoding apparatus 100 can perform block-based motion compensation related to the current block by using the pixel values of the first reference block indicated by the first motion vector and the pixel values of the second reference block indicated by the second motion vector. Further, the video decoding apparatus 100 uses the horizontal or vertical gradient value of at least one first corresponding reference pixel in the first reference block corresponding to at least one pixel included in the current pixel group, the horizontal or vertical gradient value of at least one second corresponding reference pixel in the second reference block, the pixel value of the first corresponding reference pixel, the pixel value of the second corresponding reference pixel, and the horizontal or vertical displacement vector of the current pixel, and can perform motion compensation on the current pixel group in units of pixel groups.
[0114] The video decoding apparatus 100 can obtain a predicted block of the current block by using the block-based motion compensation value generated by performing block-based motion compensation on the current block and the pixel group-based motion compensation value generated by performing pixel group-based motion compensation on the current pixel group.
[0115] In step S125, the video decoding apparatus 100 can obtain information about the residual block of the current block from the bitstream.
[0116] In step S130, the video decoding apparatus 100 can restore the current block based on the predicted block and the residual block. That is, the video decoding apparatus 100 can combine the pixel value of the residual block indicated by the residual block related to the current block and the predicted pixel value of the predicted block to generate the pixel value of the restored block of the current block.
[0117] Figure 1C illustrates a block diagram of a video encoding apparatus according to various embodiments.
[0118] A video encoding apparatus 150 according to various embodiments includes an inter prediction unit 155 and a bitstream generation unit 170.
[0119] The inter prediction unit 155 performs inter prediction by referring to various blocks for a current block based on rate and distortion costs. That is, the inter prediction unit 155 can generate predicted pixel values for the current block by using at least one of the pictures encoded before the current picture in which the current block is included.
[0120] The inter prediction unit 155 may include a motion compensation unit 160 in units of blocks and a motion compensation unit 165 in units of pixel groups.
[0121] The motion compensation unit 160 in units of blocks can perform motion compensation in units of blocks for the current block and generate motion compensation values in units of blocks.
[0122] The motion compensation unit 160 in units of blocks can determine at least one reference picture among the previously decoded pictures and determine a reference block of the current block located in at least one reference picture.
[0123] The motion compensation unit 160 in units of blocks can use the pixel values of the reference block to perform motion compensation in units of blocks related to the current block and generate motion compensation values in units of blocks. The motion compensation unit 160 in units of blocks can use the reference pixel values of the reference block corresponding to the current pixels of the current block to perform motion compensation in units of blocks related to the current block and generate motion compensation values in units of blocks.
[0124] The block-level motion compensation unit 160 can perform block-level motion compensation related to the current block by using a plurality of reference blocks included in each of a plurality of reference pictures, and generate a block-level motion compensation value. For example, when the motion prediction mode of the current block is the bidirectional prediction mode, the block-level motion compensation unit 160 can determine two reference pictures among the previously encoded pictures, and determine two reference blocks included in the two pictures. Here, the bidirectional prediction is not limited to meaning performing inter prediction by using a picture whose display order is before the current picture and a picture whose display order is after the current picture, but means performing inter prediction by using two pictures encoded before the current picture regardless of the display order.
[0125] The block-level motion compensation unit 160 can perform block-level motion compensation related to the current block by using two reference pixel values within the two reference blocks, and generate a block-level motion compensation value. The block-level motion compensation unit 160 can perform block-level motion compensation related to the current block by using the average pixel value or weighted sum of the two reference pixels, and generate a block-level motion compensation value.
[0126] The block-level motion compensation unit 160 can output a reference picture index indicating a reference picture for motion compensation of the current block among the previously encoded pictures.
[0127] The block-level motion compensation unit 160 can determine a motion vector starting from the current block and ending at the reference block of the current block, and output the motion vector. The motion vector means a vector indicating the displacement between the reference coordinates of the current block in the current picture and the reference coordinates of the reference block in the reference picture. For example, when the coordinates of the upper left corner of the current block are (1, 1) and the upper left coordinates of the reference block in the reference picture are (3, 3), the motion vector is also (2, 2).
[0128] The reference position of the reference block is not only at an integer pixel position but also at a fractional pixel position, not limited thereto. At this time, the position of the fractional pixel may be determined in units of 1 / 4 pel or 1 / 16 pel. However, it is not limited thereto, and the position of the fractional pixel may also be determined by various fractional pel units.
[0129] For example, when the reference position of the reference block is (1.5, 1.5) and the coordinates of the upper left corner of the current block are (1, 1), the motion vector is also (0.5, 0.5). In order to indicate the reference position of the reference block where the motion vector is at a fractional pixel position and is determined in units of 1 / 4 pel or 1 / 16 pel, the motion vector can be scaled to determine an integer motion vector, and the upscaled motion vector can be used to determine the reference position of the reference block. When the reference position of the reference block is at a fractional pixel position, the position of the reference pixel of the reference block is also at a fractional pixel position. Therefore, in the reference block, the pixel value at the fractional pixel position is also determined using the pixel values of the surrounding pixels whose predetermined direction components are integers.
[0130] For example, the block-based motion compensation unit 160 uses an interpolation filter to perform filtering on the pixel values of the surrounding pixels whose predetermined direction components are integers, and determines the reference pixel value at the fractional pixel position with the resulting value, and uses the pixel value of the reference pixel to determine the block-based motion compensation value related to the current block. At this time, the interpolation filter can use an M-tap interpolation filter based on DCT. The coefficients of the M-tap interpolation filter based on DCT can also be derived from DCT and IDCT. At this time, the coefficients of the interpolation filter are also filter coefficients scaled to integer coefficients in order to reduce real number operations during filtering.
[0131] At this time, the interpolation filter is also a one-dimensional interpolation filter in the horizontal or vertical direction.
[0132] The motion compensation unit 160 in block units first performs filtering on the surrounding integer pixels using a one-dimensional interpolation filter in the vertical direction, and then performs filtering on the filtered value using a one-dimensional interpolation filter in the horizontal direction, so as to determine the reference pixel value at the fractional pixel position. When using the scaled filter coefficients, the motion compensation unit 160 in block units can perform inverse scaling on the value obtained by filtering the pixels at integer positions after filtering the pixels at integer positions using a one-dimensional interpolation filter in the vertical direction. At this time, the inverse scaling may include bit-shifting to the right by the number of inverse scaling bits. The number of inverse scaling bits is also determined based on the bit depth of the sample. For example, the number of bit-shifting is also the value obtained by subtracting 8 from the bit depth of the sample.
[0133] Also, the motion compensation unit 160 in block units can perform filtering on the pixels whose horizontal components are integers using a one-dimensional interpolation filter in the horizontal direction, and then include bit-shifting to the right by the number of inverse scaling bits. The number of inverse scaling bits is also determined based on the number of bits scaled for the one-dimensional interpolation filter coefficient in the vertical direction, the number of bits scaled for the one-dimensional interpolation filter coefficient in the horizontal direction, and the bit depth of the sample.
[0134] If the motion compensation unit 160 in block units only performs bit-shifting to the right by the number of inverse scaling bits, a rounding error may occur. Therefore, a one-dimensional interpolation filter in a predetermined direction is used to perform filtering on the pixels whose components in the predetermined direction are integers, an offset is added to the filtered value, and inverse scaling can be performed on the value to which the offset is added. At this time, the offset is also 2^(inverse scaling bit number - 1).
[0135] Previously, it was described that after filtering using a one-dimensional interpolation filter in the vertical direction, the inverse scaling bit number is determined based on the bit depth of the samples. However, it is not limited thereto, and it is determined in consideration of not only the bit depth of the samples but also the bit number scaled for the interpolation filter coefficients. That is, when performing filtering, considering the size of the registers used and the size of the buffer for storing the values generated during the filtering process, within a range where no overflow occurs, the inverse scaling bit number is also determined based on the bit depth of the samples and the bit number scaled for the interpolation filter coefficients.
[0136] The motion compensation unit 165 at the pixel group level performs motion compensation at the pixel group level for the current block and can generate a motion compensation value at the pixel group level. For example, when the motion prediction mode is the bidirectional motion prediction mode, the motion compensation unit 165 at the pixel group level can perform motion compensation at the pixel group level for the current block and generate a motion compensation value at the pixel group level.
[0137] The motion compensation unit 165 at the pixel group level can utilize the gradient values of the pixels included in the reference block of the current block and perform motion compensation at the pixel group level for the current block to generate a motion compensation value at the pixel group level.
[0138] The motion compensation unit 165 at the pixel group level can apply a filter to the first peripheral region of the first pixel among the pixels of the first reference block in the first reference picture and the second peripheral region of the second pixel among the pixels of the second reference block in the second reference picture to generate the gradient value of the first pixel and the gradient value of the second pixel.
[0139] The motion compensation unit 165 at the pixel group level can determine the pixel values and gradient values of the pixels within a first window of a predetermined size including the first reference pixel centered on the first reference pixel in the first reference picture, and can determine the pixel values and gradient values of the pixels within a second window of a predetermined size including the second reference pixel centered on the second reference pixel in the second reference picture. The motion compensation unit 165 at the pixel group level can utilize the pixel values and gradient values of the pixels within the first window and the pixel values and gradient values of the pixels within the second window to determine the displacement vector per unit time related to the current pixel.
[0140] The motion compensation unit 165 at the pixel group level can utilize the displacement vector per unit time and the gradient value of the reference pixel to perform motion compensation at the pixel group level related to the current block and generate a motion compensation value at the pixel group level.
[0141] The position of the reference pixel is not only the position of an integer pixel but also the position of a fractional pixel, not limited thereto.
[0142] When the reference position of the reference block is the position of a fractional pixel, the gradient value of the reference pixel within the reference block is also determined using the pixel values of the surrounding pixels whose predetermined direction component is an integer.
[0143] For example, the motion compensation unit 165 at the pixel group level can perform filtering on the pixel values of the surrounding pixels whose predetermined direction component is an integer using a gradient filter, and use the resulting value to determine the gradient value of the reference pixel. At this time, the filter coefficient of the gradient filter is also determined using the coefficient predetermined for the interpolation filter based on DCT.
[0144] The coefficients of the gradient filter are also filter coefficients scaled to integer coefficients in order to reduce real-number operations during filtering. At this time, the gradient filter used is also a one-dimensional gradient filter in the horizontal or vertical direction.
[0145] The motion compensation unit 165 in units of pixel groups can perform filtering on peripheral pixels whose corresponding direction components are integers by using a one-dimensional gradient filter in the horizontal or vertical direction to determine the gradient value in the horizontal or vertical direction related to the reference pixel.
[0146] For example, the motion compensation unit 165 in units of pixel groups can perform filtering on pixels whose vertical components are integers among the vertical pixels from integer pixels adjacent to the reference pixel by using a one-dimensional interpolation filter in the vertical direction, and determine the pixel values of pixels whose vertical components are decimals.
[0147] The motion compensation unit 165 in units of pixel groups can also perform filtering on pixels located in other columns adjacent to the integer pixels adjacent to the reference pixel by using a one-dimensional interpolation filter in the vertical direction for the peripheral integer pixels in the vertical direction, and determine the pixel values at the decimal pixel positions in the other columns. Here, the positions of the pixels in the other columns are decimal pixel positions in the vertical direction and integer pixel positions in the horizontal direction.
[0148] That is, when the position of the reference pixel is (x + α, y + β) (where x and y are integers and α and β are decimals), the motion compensation unit 165 in units of pixel groups can perform filtering on the peripheral integer pixels in the vertical direction from the position (x, y) by using an interpolation filter in the vertical direction, thereby determining the pixel value at the position (x, y + β).
[0149] The motion compensation unit 165 at the pixel group level can determine the horizontal gradient value at the position (x+α, y+β) by performing filtering on the pixel values of the pixels with integer horizontal components among the horizontally positioned pixels using a horizontal gradient filter based on the pixel value at the position (x, y+β) and the pixel value at the position (x, y+β).
[0150] The order of using the one-dimensional gradient filter and the one-dimensional interpolation filter is not restricted. As described above, first, the vertical interpolation filter can be used to perform filtering on the pixels at integer positions to generate the vertical interpolation filtering value, and then the one-dimensional horizontal gradient filter can be used to perform filtering on the vertical interpolation filtering value. However, it is not limited to this. First, the one-dimensional horizontal gradient filter can be used to perform filtering on the pixels at integer positions to generate the horizontal gradient filtering value, and then the one-dimensional vertical interpolation filter can be used to perform filtering on the horizontal gradient filtering value.
[0151] The above has described in detail the content of the motion compensation unit 165 at the pixel group level determining the horizontal gradient value at the position (x+α, y+β).
[0152] The motion compensation unit 165 at the pixel group level can also determine the vertical gradient value at the position (x+α, y+β) in a manner similar to where the horizontal gradient value is determined.
[0153] The motion compensation unit 165 in pixel group units uses a one-dimensional gradient filter in the vertical direction to perform filtering on the surrounding integer pixels in the vertical direction from the integer pixels around the reference pixel, and can determine the gradient value in the vertical direction related to the reference pixel. The motion compensation unit 165 in pixel group units also uses a one-dimensional gradient filter in the vertical direction to perform filtering on the surrounding integer pixels in the vertical direction for the pixels located in other columns adjacent to the reference pixel, and can determine the gradient value in the vertical direction related to the pixels located in other columns while being adjacent to the reference pixel. Here, the position of the pixel is the position of a fractional pixel in the vertical direction and is also the position of an integer pixel in the horizontal direction.
[0154] That is, when the position of the reference pixel is (x + α, y + β) (where x and y are integers and α and β are fractions), the motion compensation unit 165 in pixel group units can determine the gradient value in the vertical direction at the position (x, y + β) by performing filtering on the surrounding integer pixels in the vertical direction from the position (x, y) using a gradient filter in the vertical direction.
[0155] The motion compensation unit 165 in pixel group units can determine the gradient value in the vertical direction at the position (x + α, y + β) by performing filtering on the gradient value at the position (x, y + β) and the gradient values of the surrounding integer pixels located horizontally from the position (x, y + β) using an interpolation filter in the horizontal direction.
[0156] The order of using the one-dimensional gradient filter and the one-dimensional interpolation filter is not restricted. As described above, first, the vertical gradient filter can be used to filter the pixels at integer positions to generate the vertical gradient filtering values, and then the one-dimensional horizontal interpolation filter can be used to filter the vertical gradient filtering values. However, it is not limited to this. First, the one-dimensional horizontal interpolation filter can be used to filter the pixels at integer positions to generate the horizontal interpolation filtering values, and then the one-dimensional vertical gradient filter can be used to filter the horizontal interpolation filtering values.
[0157] The above has described in detail the content of the motion compensation unit 165 using the gradient filter and the interpolation filter to determine the gradient value at the sub-pixel position. However, it is not limited to this. The gradient filter and the interpolation filter can also be used to determine the gradient value at the integer pixel position.
[0158] In the case of integer pixels, even if the interpolation filter is not used, the pixel value can be determined. However, for the sake of consistent processing with the processing of sub-pixels, the interpolation filter may also be used to filter the integer pixels and their surrounding integer pixels. For example, the interpolation filter coefficient at the integer pixel is also {0, 0, 64, 0, 0}, and since the interpolation filter coefficient multiplied by the surrounding integer pixels is 0, filtering is performed using only the pixel value of the current integer pixel. As a result, the value generated by filtering the integer pixel and its surrounding integer pixels using the interpolation filter is the same as the pixel value of the current integer pixel.
[0159] On the other hand, when using scaled filter coefficients, the motion compensation unit 165 at the pixel group level utilizes a one-dimensional gradient filter in the horizontal direction, filters the pixels at integer positions, and can perform inverse scaling on the filtered values. At this time, the inverse scaling may include bit-shifting to the right by the number of inverse scaling bits. The number of inverse scaling bits is also determined based on the bit depth of the samples. For example, the number of inverse scaling bits is also a value obtained by subtracting 8 from the bit depth of the samples.
[0160] The motion compensation unit 165 at the pixel group level can utilize an interpolation filter in the vertical direction, filter the pixels with integer vertical components, and then perform inverse scaling. At this time, the inverse scaling may include bit-shifting to the right by the number of inverse scaling bits. The number of inverse scaling bits is also determined based on the scaled bit number of the one-dimensional interpolation filter in the vertical direction, the scaled bit number of the one-dimensional gradient filter in the horizontal direction, and the bit depth of the samples.
[0161] If the motion compensation unit 165 at the pixel group level only performs bit-shifting to the right by the number of inverse scaling bits, a rounding error may occur. Therefore, after filtering is performed using a one-dimensional interpolation filter, an offset can be added to the filtered value, and inverse scaling can be performed on the value with the offset added. At this time, the offset is also 2^(bit-shifting number - 1).
[0162] When performing motion compensation in block units, the inter prediction unit 110 can use motion vectors and store the motion vectors. At this time, the motion vector unit is also a 4x4-sized block. On the other hand, when storing motion vectors after motion compensation in block units, the motion vector storage unit is also blocks of various sizes that are not 4x4-sized (for example, blocks of RxR size; R is an integer). At this time, the motion vector storage unit is also a block larger than 4x4 size. For example, it is also a 16x16-sized block.
[0163] On the other hand, when performing motion compensation in pixel group units, together with the size of the current block, based on the window size and the interpolation filter length, the size of the target block for performing motion compensation in pixel group units is also expanded. The reason why the size of the target block is expanded based on the window size compared to the size of the current block is that in the case of pixels located at the edge of the current block, using the window, based on the pixels located at the current edge and the surrounding pixels, motion compensation in pixel group units related to the current block is performed.
[0164] Therefore, in order to reduce the number of memory accesses and the execution of multiplication operations, the motion compensation unit 120 in pixel group units adjusts the positions of the pixels outside the current block among the pixels in the window to the positions of the pixels adjacent to the inside of the current block, and determines the pixel values and gradient values at the adjusted pixel positions, so that the number of memory accesses and the number of multiplication operations are also reduced in the process of performing motion compensation in pixel group units using the window.
[0165] The motion compensation unit 120 at the pixel group level does not use the pixel values of integer-position pixels to determine the gradient value of the reference pixel, which is a value necessary for motion compensation at the pixel group level. That is, the compensation unit 120 at the pixel group level applies a horizontal or vertical gradient filter to the pixel values of fractional-position pixels to calculate the horizontal or vertical gradient value of the first corresponding reference pixel in the first reference block or the second corresponding reference pixel in the second reference block. At this time, the gradient filter length can also be 5. At this time, the coefficients of the filter can have coefficients that are symmetric about the central coefficient of the filter. The fractional-position pixel is also a pixel in which at least one of the horizontal and vertical components indicating the position of the pixel has a fractional value.
[0166] The pixel values of the fractional-position pixels in the first reference block or the second reference block are also calculated by applying a horizontal or vertical interpolation filter to the pixel values of the integer-position pixels.
[0167] The displacement vector per unit time in the horizontal or vertical direction for the current pixel group is also the displacement vector per unit time determined by using the values determined based on the first corresponding reference pixels included in the first corresponding reference pixel group in the first reference picture corresponding to the current pixel group, the second corresponding reference pixels included in the second corresponding reference pixel group in the second reference picture, and the pixel values and gradient values of their surrounding pixels, the first POC difference between the first reference picture and the current picture, and the second POC difference between the second reference picture and the current picture.
[0168] When the first corresponding reference pixel or the second corresponding reference pixel is a boundary peripheral pixel adjacent to the inside of the boundary of the first reference block or the second reference block, the compensation unit 120 at the pixel group level can use the pixel value of the boundary peripheral pixel to derive the pixel value of the pixel located outside the boundary of the first reference block or the second reference block.
[0169] The pixel group unit compensation unit 120 can determine a displacement vector per unit time in the horizontal or vertical direction of the current block based on the pixel values of the boundary peripheral pixels and the pixel values of the pixels located outside the boundary of the current block induced by using the pixel values of the boundary peripheral pixels. That is, there are pixels located outside the boundary within the pixels included in the window centered on the boundary peripheral pixels. At this time, the pixel values and gradient values of the pixels located outside the boundary are also the pixel values and gradient values of the pixels induced from the boundary peripheral pixels that are not the values stored in the memory.
[0170] The pixel group unit compensation unit 120 can calculate the value related to the current pixel by using the pixel values of the first corresponding reference pixels included in the first reference block, the pixel values of the second corresponding reference pixels included in the second reference block, the gradient values of the first corresponding reference pixels, the gradient values of the second corresponding reference pixels, the first POC difference between the first reference picture and the current picture, and the second POC difference between the second reference picture and the current picture. That is, the value related to the current pixel is also the result value of a function based on the pixel values and gradient values of the corresponding reference pixels of each reference picture and the POC differences between each reference picture and the current picture.
[0171] The pixel group unit compensation unit 120 can calculate a value related to the corresponding peripheral pixel calculated using the pixel value of the first corresponding peripheral pixel of the first corresponding reference pixel, the gradient value of the first corresponding peripheral pixel, the pixel value of the second corresponding peripheral pixel of the second corresponding reference pixel, the gradient value of the second corresponding peripheral pixel, the first POC difference between the first reference picture and the current picture, and the second POC difference between the second reference picture and the current picture. That is, the value related to the peripheral pixel is also the result value of a function based on the pixel value and gradient value of the corresponding reference pixel of each reference picture and the POC difference between each reference picture and the current picture. That is, the value related to the corresponding peripheral pixel is also the result value of a function based on the pixel value and gradient value of the corresponding peripheral pixel of each reference picture and the POC difference between each reference picture and the current picture.
[0172] The pixel group unit compensation unit 120 can calculate a weighted average value for the current pixel necessary for calculating a displacement vector per unit time in the horizontal direction, using the value related to the current pixel, the value related to the corresponding peripheral pixel, and a weighting value. At this time, the weighting value is also determined based on the distance between the current pixel and the peripheral pixel, the distance between the pixel and the block boundary, the number of pixels located outside the boundary, or whether the pixel is located inside or outside the boundary.
[0173] The weighted average value for the current pixel is also a value calculated by applying an exponential smoothing technique in the up, down, left, and right directions to the values related to the pixels included in the first reference block and the second reference block. By applying the exponential smoothing technique in the up, down, left, and right directions to the values related to the pixels, the value calculated for the current pixel has the largest weighting value related to the value related to the current pixel, and the weighting value related to the value related to the peripheral pixel is also a value that exponentially decreases according to the distance from the current pixel.
[0174] The pixel group unit compensation unit 120 can determine the displacement vector per unit time in the horizontal or vertical direction of the current pixel by using the weighted average value of the current pixel.
[0175] The inter prediction unit 155 can generate the predicted pixel value of the current block by using the motion compensation value in block units related to the current block and the motion compensation value in pixel group units related to the current block. For example, the inter prediction unit 155 can combine the motion compensation value in block units related to the current block and the motion compensation value in pixel group units to generate the predicted pixel value of the current block. In particular, when the motion prediction mode of the current block is the bidirectional motion prediction mode, the inter prediction unit 155 can use the motion compensation value in block units related to the current block and the motion compensation value in pixel group units related to the current block to generate the predicted pixel value of the current block.
[0176] When the motion prediction mode of the current block is the unidirectional motion prediction mode, the inter prediction unit 155 can use the motion compensation value in block units related to the current block to generate the predicted pixel value of the current block. Here, unidirectional means using one reference picture among the previously encoded pictures. One reference picture is not limited to being a picture before the current picture in the display order, but can also be a picture after it.
[0177] The inter prediction unit 155 can determine the motion prediction mode of the current block and output information indicating the motion prediction mode of the current block. For example, the inter prediction unit 155 can determine the bidirectional motion prediction mode in the motion prediction mode of the current block and output information indicating the bidirectional motion prediction mode. Here, the bidirectional motion prediction mode means a mode of motion prediction using reference blocks in two decoded reference pictures.
[0178] The bitstream generation unit 170 can generate a bitstream including motion vectors indicating reference blocks. The bitstream generation unit 170 can encode motion vectors indicating reference blocks and generate a bitstream including the encoded motion vectors. The bitstream generation unit 170 can encode the vector difference value of the motion vectors indicating reference blocks and generate a bitstream including the encoded vector difference value of the motion vectors. Here, the vector difference value of the motion vectors means the difference value between the motion vector and the predictor of the motion vector. At this time, the vector difference value of the motion vectors means the vector difference value of the motion vectors related to the reference pictures related to each prediction direction including the L0 direction and the L1 direction. Here, the vector difference value of the motion vectors related to the L0 direction means the vector difference value of the motion vectors indicating reference blocks in the reference picture included in the L0 reference picture list, and the vector difference value of the motion vectors related to the L1 direction means the vector difference value of the motion vectors indicating reference blocks in the reference picture included in the L1 reference picture list.
[0179] In addition, the bitstream generation unit 170 can generate a bitstream further including information indicating the motion prediction mode of the current block. The bitstream generation unit 170 can encode the reference picture index indicating the reference picture of the current block among the previously encoded pictures and generate a bitstream including the encoded reference picture index. At this time, the reference picture index means the reference picture index related to each prediction direction including the L0 direction and the L1 direction. Here, the reference picture index related to the L0 direction means the index indicating the reference picture among the pictures included in the L0 reference picture list, and the reference picture index related to the L1 direction means the index indicating the reference picture among the pictures included in the L1 reference picture list.
[0180] The video encoding device 150 may include a video encoding unit (not shown), and the video encoding unit (not shown) may include an inter prediction unit 155 and a bitstream generation unit 170. The video encoding unit will be described with reference to FIG. 1F.
[0181] FIG. 1D illustrates a flowchart of a video encoding method according to various embodiments.
[0182] Referring to FIG. 1D, at step S150, the video encoding device 150 performs motion compensation for the current block and motion compensation in units of pixel groups, and can obtain a predicted block of the current block, a first motion vector and a second motion vector, and parameters related to the motion compensation in units of pixel groups.
[0183] At step S155, the video encoding device 150 can generate a bitstream including information about the first motion vector and the second motion vector, and motion prediction mode information indicating whether the motion prediction mode related to the current block is a bidirectional motion prediction mode. Here, the first motion vector is also a motion vector indicating a first reference block of a first reference picture corresponding to the current block in the current picture from the current block, and the second motion vector is also a motion vector indicating a second reference block of a second reference picture corresponding to the current block in the current picture from the current block.
[0184] The video encoding device 150 can encode a residual block of a current block indicating a difference between pixels of a predicted block of the current block and an original block of the current block, and generate a bitstream further including the encoded residual signal. The video encoding device 150 can encode information about a prediction mode of the current block and a reference picture index, and generate a bitstream further including the encoded information about the prediction mode of the current block and the reference picture index. For example, the video encoding device 150 can encode information indicating whether the prediction mode of the current block is an inter prediction mode and a reference picture index indicating at least one of the previously decoded pictures, and generate a bitstream further including the encoded information about the prediction mode of the current block and the reference picture index.
[0185] Based on a gradient value in a horizontal or vertical direction of a first corresponding reference pixel in the first reference block corresponding to a current pixel included in a current pixel group within a current block, a gradient value in a horizontal or vertical direction of a second corresponding reference pixel in the second reference block corresponding to the current pixel, a pixel value of the first corresponding reference pixel, a pixel value of the second corresponding reference pixel, and a displacement vector per unit time in a horizontal or vertical direction of the current pixel, the video encoding device 150 can perform motion compensation in units of blocks related to the current block and motion compensation in units of pixel groups. The video encoding device 150 can perform motion compensation in units of blocks related to the current block and motion compensation in units of pixel groups, and obtain a predicted block of the current block.
[0186] At this time, the displacement vector per unit time in the horizontal or vertical direction of the pixels of the current block including the pixels adjacent to the inside of the boundary of the current block is determined by using the values related to the reference pixels included in the first reference block and the second reference block without using the values stored for the pixels located outside the boundaries of the first reference block and the second reference block. The values related to the reference pixels included in the first reference block and the second reference block are also the pixel values or gradient values of the reference pixels.
[0187] FIG. 1E illustrates a block diagram of a video decoding unit 600 according to various embodiments.
[0188] The video decoding unit 600 according to various embodiments performs operations that the video data goes through in the video decoding unit (not shown) of the video decoding apparatus 100 to encode the video data.
[0189] Referring to FIG. 1E, the entropy decoding unit 615 parses the encoded video data to be decoded and the encoding information necessary for decoding from the bitstream 605. The encoded video data is quantized transform coefficients, and the inverse quantization unit 620 and the inverse transform unit 625 restore the residual data from the quantized transform coefficients.
[0190] The intra prediction unit 640 performs intra prediction block by block. The inter prediction unit 635 performs inter prediction block by block using the reference video acquired in the restored picture buffer 630. The inter prediction unit 635 in FIG. 1E corresponds to the inter prediction unit 110 in FIG. 1A.
[0191] By adding the prediction data related to each block generated by the intra prediction unit 640 or the inter prediction unit 635 and the residual data, the data of the spatial region related to the block of the current video 605 is restored, and the deblocking unit 645 and the SAO execution unit 650 can perform loop filtering on the restored data of the spatial region and output the filtered restored video 660. Also, the restored video stored in the restored picture buffer 630 may be output as a reference video.
[0192] In the decoding unit (not shown) of the video decoding apparatus 100, in order to decode video data, the step-by-step operations of the video decoding unit 600 according to various embodiments may also be performed block by block.
[0193] FIG. 1F illustrates a block diagram of a video encoding unit according to various embodiments.
[0194] The video encoding unit 700 according to various embodiments performs the operations through which the video data is encoded in the video encoding unit (not shown) of the video encoding apparatus 150.
[0195] That is, the intra prediction unit 720 performs intra prediction block by block in the current video 705, and the inter prediction unit 715 performs inter prediction block by block using the reference video obtained from the current video 705 and the restored picture buffer 710. Here, the inter prediction unit 715 in FIG. 1E corresponds to the inter prediction unit 160 in FIG. 1C.
[0196] The prediction data related to each block output from the intra prediction unit 720 or the inter prediction unit 715 is subtracted from the data related to the block to be encoded in the current video 705 to generate residual data. The conversion unit 725 and the quantization unit 730 can perform conversion and quantization on the residual data and output the quantization coefficients for each block. The inverse quantization unit 745 and the inverse conversion unit 750 can perform inverse quantization and inverse conversion on the quantized conversion coefficients to restore the residual data in the spatial domain. The restored residual data in the spatial domain is added to the prediction data related to each block output from the intra prediction unit 720 or the inter prediction unit 715, thereby restoring the data in the spatial domain related to the block in the current video 705. The deblocking unit 755 and the SAO execution unit 760 perform in-loop filtering on the restored data in the spatial domain to generate a filtered restored video. The generated restored video is stored in the restored picture buffer 710. The restored video stored in the restored picture buffer 710 is also used as a reference video for inter prediction of other videos. The entropy encoding unit 735 performs entropy encoding on the quantized conversion coefficients, and the entropy-encoded coefficients are also output as a bitstream 740.
[0197] For the video encoding unit 700 according to various embodiments to be applied to the video encoding apparatus 150, the step-by-step operations of the video encoding unit 700 according to various embodiments may also be performed for each block.
[0198] FIG. 2 is a reference diagram for explaining the process of block-based bidirectional motion prediction and its compensation according to an embodiment. Referring to FIG. 2, the video encoding device 150 performs bidirectional motion prediction to search for the region in the first reference picture 210 and the second reference picture 220 that is most similar to the current block 201 to be encoded in the current picture 200. Here, the first reference picture 210 is a picture before the current picture 200, and the second reference picture 220 is assumed to be a picture after the current picture 200. The video encoding device 150 determines the bidirectional motion prediction result, the first corresponding region 212 in the first reference picture 210 that is most similar to the current block 201, and the second corresponding region 222 in the second reference picture 220 that is most similar to the current block 201. Here, the first corresponding region and the second corresponding region also become the reference regions of the current block.
[0199] Then, the video encoding device 150 determines the first motion vector MV1 based on the position difference between the block 211 at the same position as the current block 201 in the first reference picture 210 and the first corresponding region 212, and determines the second motion vector MV2 based on the position difference between the block 221 at the same position as the current block 201 in the second reference picture 220 and the second corresponding region 222.
[0200] The video encoding device 150 uses the first motion vector MV1 and the second motion vector MV2 to perform block-based bidirectional motion compensation on the current block 201.
[0201] For example, if the pixel value located at (i, j) (where i and j are integers) in the first reference picture 210 is P0(i, j), the pixel value located at (i, j) in the second reference picture 220 is P1(i, j), MV1 = (MVx1, MVy1), and MV2 = (MVx2, MVy2), then the block-based bidirectional motion compensation value P_BiPredBlock(i, j) for the pixel at the (i, j) position in the current block 201 may also be calculated by the following formula: P_BiPredBlock(i, j) = {P0(i + MVx1, j + MVy1) + P1(i + MVx2, j + MVy2)} / 2. In this way, the video encoding device 150 can perform block-based motion compensation for the current block 201 by using the average value or weighted sum of the pixels in the first corresponding region 212 and the second corresponding region 222 indicated by the first motion vector MV1 and the second motion vector MV2, and generate a block-based motion compensation value.
[0202] FIGS. 3A to 3C are reference diagrams for explaining the process of performing motion compensation in units of pixel groups according to an embodiment.
[0203] In FIG. 3A, it is assumed that the first corresponding region 310 and the second corresponding region 320 respectively correspond to the first corresponding region 212 and the second corresponding region 222 in FIG. 2, and are shifted using the bidirectional motion vector (MV1, MV2) so as to overlap the current block 300.
[0204] Also, define the pixel at the bidirectionally predicted (i, j) (where i and j are integers) position in the current block 300 as P(i, j), the first reference pixel value of the first reference picture corresponding to the bidirectionally predicted pixel P(i, j) in the current block 300 as P0(i, j), and the second reference pixel value of the second reference picture corresponding to the bidirectionally predicted pixel P(i, j) in the current block 300 as P1(i, j).
[0205] In other words, the first reference pixel value P0(i,j) is the pixel corresponding to the pixel P(i,j) of the current block 300 determined by the bidirectional motion vector MV1 indicating the first reference picture, and the pixel value P1(i,j) of the second reference pixel is the pixel corresponding to the pixel P(i,j) of the current block 300 determined by the bidirectional motion vector MV2 indicating the second reference picture.
[0206] Also, the horizontal gradient value of the first reference pixel is
[0207]
Number
[0208]
Number
[0209]
Number
[0210]
Number
[0211] Assume that there is a small motion determined in the video sequence. Then, the pixel in the first corresponding region 310 of the first reference picture that is most similar to the current pixel P(i, j) bidirectionally motion-compensated in pixel group units is not the first reference pixel P0(i, j), but rather the first displaced reference pixel PA obtained by moving the first reference pixel P0(i, j) by a predetermined displacement vector. As described above, since it is assumed that there is motion determined in the video sequence, in the second corresponding region 320 of the second reference picture, the pixel that is most similar to the current pixel P(i, j) can be estimated as the second displaced reference pixel PB obtained by moving the second reference pixel P1(i, j) by a predetermined displacement vector.
[0212] The displacement vector is also composed of the displacement vector Vx in the x-axis direction and the displacement vector Vy in the y-axis direction as described above. Therefore, the motion compensation unit 165 in pixel group units calculates the displacement vector Vx in the x-axis direction and the displacement vector Vy in the y-axis direction that constitute such a displacement vector, and uses them to perform motion compensation in pixel group units.
[0213] Optical flow means the pattern of apparent motion on the appearance of an object or surface induced by the relative motion between an observer (such as an eye or a video image acquisition device like a camera) and a scene. In a video sequence, the optical flow is also expressed by calculating the motion between frames acquired at arbitrary times t and t + Δt. The pixel value located at (x, y) within the frame at time t is also defined as I(x, y, t). That is, I(x, y, t) is also a value that changes spatiotemporally. Differentiating I(x, y, t) with respect to time t gives the following formula (1).
[0214]
Equation
[0215]
Number
[0216] The motion compensation unit 165 at the pixel group level calculates the displacement vector Vx in the x-axis direction and the displacement vector Vy in the y-axis direction according to Equation (2), and performs motion compensation at the pixel group level using such displacement vectors Vx and Vy. In Equation (2), since the pixel value I(x, y, t) is the value of the original signal, directly using the value of the original signal as it is will induce a lot of overhead during encoding. Therefore, the motion compensation unit 165 at the pixel group level can calculate the displacement vectors Vx and Vy according to Equation (2) by using the pixels of the first reference picture and the second reference picture determined as the result of bidirectional motion prediction in block units. That is, the motion compensation unit 165 at the pixel group level determines the displacement vector Vx in the x-axis direction and the displacement vector Vy in the y-axis direction such that Δ becomes the minimum within a window (Ωij) of a predetermined size including surrounding pixels with the currently pixel P(i, j) to be bidirectionally motion-compensated at the center. Although the case where Δ is 0 is most desirable, since there are no displacement vectors Vx in the x-axis direction and Vy in the y-axis direction that satisfy the case where Δ is 0 for all pixels within the window (Ωij), the displacement vector Vx in the x-axis direction and the displacement vector Vy in the y-axis direction that minimize Δ are determined. The process of obtaining the displacement vectors Vx and Vy will be described in detail with reference to FIG. 8A.
[0217] To determine the predicted pixel value of the current pixel, a function P(t) related to t may also be determined as in the following Equation (3).
[0218]
Equation
[0219] Assume that the time distance from the first reference picture (assuming that the first reference picture is at a position temporally prior to the current picture) to the current picture is τ0, and the time distance from the second reference picture (assuming that the second reference picture is at a position temporally subsequent to the current picture) to the current picture is τ1. Then, the reference pixel value at the first reference picture is the same as P(-τ0), and the reference pixel value at the second reference picture is the same as P(τ1). Hereinafter, for the convenience of calculation, it is assumed that both τ0 and τ1 are the same as τ.
[0220] The coefficients of each degree of P(t) are also determined by the following mathematical formula (4). Here, P0(i,j) means the pixel value at the (i,j) position of the first reference picture, and P1(i,j) means the pixel value at the (i,j) position of the second reference picture.
[0221]
Equation
[0222]
Equation
[0223]
Equation
[0224] For the sake of convenience as described above, when the time distance from the first reference picture to the current picture is τ and the time distances from the second reference picture to the current picture are all the same as τ, the process of determining the predicted pixel value of the current pixel was described. However, the time distance from the first reference picture to the current picture is τ0 and the time distance from the second reference picture to the current picture is τ1. At this time, the predicted pixel value P(0) of the current pixel is also determined as shown in the following formula (7).
[0225]
Equation
[0226]
Equation
[0227] For example, as shown in FIG. 3B, both the first reference picture including the first corresponding region and the second reference picture including the second corresponding region can be located temporally before the current picture including the current block in the display order.
[0228] In that case, in the formula (8) derived with reference to FIG. 3A, the predicted pixel value P(0) of the current pixel is also determined by the formula (9) in which τ1 indicating the temporal distance difference between the second reference picture and the current picture is replaced with -τ1.
[0229]
Number
[0230] In that case, in the formula (8) derived with reference to FIG. 3A, the predicted pixel value of the current pixel is also determined by the formula (10) in which τ0 indicating the temporal distance difference between the first reference picture and the current picture is replaced with -τ0.
[0231]
Number
[0232] FIG. 4 is a reference diagram for explaining the process of calculating gradient values in the horizontal and vertical directions according to an embodiment. Referring to FIG. 4, the horizontal gradient value
[0233] [Number] and the vertical gradient value
[0234] [Number] is also calculated by obtaining the amount of change in pixel values at the peripheral fractional pixel positions adjacent to the first reference pixel P0(i,j)410 in the horizontal direction and the amount of change in pixel values at the peripheral fractional pixel positions adjacent in the vertical direction. That is, as in the following formula (11), from P0(i,j), the amount of change in pixel values of the fractional pixel P0(i - h,j)460 and the fractional pixel P0(i + h,j)470, which are separated by h (h is a fractional value smaller than 1) in the horizontal direction, is calculated, and the gradient value
[0235] [Number] in the horizontal direction is calculated, the amount of change in pixel values of the fractional pixel P0(i,j - h)480 and the fractional pixel P0(i,j + h)490, which are separated by h in the vertical direction, is calculated, and the vertical gradient value
[0236] [Number] can be calculated.
[0237] [Number] The values of the fractional pixels P0(i - h,j)460, P0(i + h,j)470, P0(i,j - h)480, and P0(i,j + h)490 are also calculated using a general interpolation method. Also, the horizontal and vertical gradient values of the second reference pixels of other second reference pictures are also calculated similarly to formula (11).
[0238] According to one embodiment, as in formula (11), instead of calculating the amount of change in pixel values at the fractional pixel position and calculating the gradient value, a predetermined filter can be used to calculate the gradient value at each reference pixel. The filter coefficients of the predetermined filter are also determined from the coefficients of the interpolation filter used to obtain the pixel values at the fractional pixel position in consideration of the linearity of the filter.
[0239] FIG. 5 is a reference diagram for explaining the process of calculating gradient values in the horizontal and vertical directions according to another embodiment.
[0240] According to another embodiment, the gradient value is also determined by applying a predetermined filter to the pixels of the reference picture. Referring to FIG. 5, for the reference pixel P0500 for which the video decoder 100 is trying to obtain the current horizontal gradient value, M Max pixels 520 are on the left side and |M Min | pixels 510 are on the right side, and a predetermined filter is applied to calculate the horizontal gradient value of P0500. At this time, the filter coefficients used are, as shown in FIGS. 7A to 7D, M Max integer pixels used to determine the window size and M Min integer pixels, and are also determined by the α value indicating the interpolation position (fractional pel position) between the two. As an example, referring to FIG. 7A, M Min is -2, M Max is 3, and when it is about 1 / 4 away from the reference pixel P0500, that is, when α = 1 / 4, the filter coefficients {4, -17, -36, 60, -15, 4} in the second row of FIG. 7A are applied to the surrounding pixels P-2, P-1, P0, P1, P2, P3. In that case, the horizontal gradient value of the reference pixel 500
[0241]
Equation
[0242]
Equation
[0243] FIGS. 6A and 6B are diagrams for explaining the process of determining the gradient values in the horizontal and vertical directions using a one-dimensional filter according to an embodiment.
[0244] Referring to FIG. 6A, in the reference picture, in order to determine the horizontal gradient value of the reference pixel, filtering is performed on the integer pixels using a plurality of one-dimensional filters. Motion compensation in pixel group units is additional motion compensation performed after motion compensation in block units. Therefore, in the process of motion compensation in block units, the reference position of the reference block of the current block indicated by the motion vector is also a fractional pixel position, and motion compensation in pixel group units is performed on the reference pixels in the reference block at the fractional pixel position. Therefore, filtering is performed in consideration of determining the gradient value of the pixel at the fractional pixel position.
[0245] Referring to FIG. 6A, first, the video decoding device 100 can perform filtering on the pixels located horizontally or vertically from the surrounding integer pixels of the reference pixel in the reference picture using a first one-dimensional filter. Similarly, the video decoding device 100 can perform filtering on the adjacent integer pixels located in a row or column different from the reference pixel using the first one-dimensional filter. The video decoding device 100 can generate the horizontal gradient value of the reference pixel by performing filtering on the value generated by the filtering using a second one-dimensional filter.
[0246] For example, when the position of the reference pixel is the position of a fractional pixel at (x + α, y + β) (where x and y are integers and α and β are fractions), for the integer pixels (x, y) and (x - 1, y), (x + 1, y), …, (x + M Min , y), (x + M Max , y) (where M Min , M Max are all integers), a one-dimensional vertical interpolation filter is used and filtering is performed as shown in the following formula (12).
[0247]
Equation
[0248] That is, the first one-dimensional filter is also an interpolation filter for determining the fractional pixel value in the vertical direction. offset1 is an offset for preventing rounding errors, and shift1 means the number of inverse scaling bits. Temp[i, j + β] means the pixel value at the fractional pixel position (i, j + β). Temp[i’, j + β] (where i’ is an integer from i + M min to i + M max is also determined by substituting i with i’ in formula (12).
[0249] Next, the video decoding device 100 can perform filtering on the pixel value at the fractional pixel position (i, j + β) and the pixel value at the fractional pixel position (i’, j + β) using a second one-dimensional filter.
[0250]
Equation
[0251] That is, according to Equation (13), the video decoding apparatus 100 performs filtering on the pixel value (Temp[i, j+β]) at (i, j+β) and the pixel value (Temp[i’, j+β]) located vertically from the pixel position (i, j+β) using the gradient filter (gradFilterα), thereby obtaining the horizontal gradient value at (i+α, j+β)
[0252]
Equation
[0253] Previously, the interpolation filter was first applied, and then the gradient filter was applied to explain the content of determining the horizontal gradient value. However, it is not limited thereto. First, the gradient filter may be applied, and then the interpolation filter may be applied to determine the horizontal gradient value. In the following, an embodiment in which the gradient filter is applied first and then the interpolation filter is applied to determine the horizontal gradient value will be described.
[0254] For example, when the position of the reference pixel is the fractional pixel position of (x+α, y+β) (x and y are integers, and α and β are fractions), the horizontal integer pixels (x, y) and (x-1, y), (x+1, y), …, (x+M Min , y), (x+M Max, y)(M Min , M Mmax are all integers), the first one-dimensional filter is used, and filtering is performed as in the following formula (14).
[0255]
Equation
[0256] That is, the first one-dimensional filter is also an interpolation filter for determining the horizontal gradient value of the pixel at the position where the horizontal component of the pixel position is fractional. offset3 is an offset for preventing rounding errors, and shift3 means the number of inverse scaling bits. Temp[i + α, j] means the horizontal gradient value at the pixel position (i + α, j). Temp[i + α, j'] (j' is an integer from j + M min to j + M max is also determined by substituting j with j' in formula (14).
[0257] Next, the video decoding device 100 can perform filtering as in the following formula (15) using the second one-dimensional filter on the horizontal gradient value at the pixel position (i + α, j) and the horizontal gradient value at the pixel position (i + α, j').
[0258]
Equation
[0259] That is, according to Equation (15), the video decoding apparatus 100 performs filtering on the horizontal gradient value (Temp[i + α, j]) at (i + α, j) and the horizontal gradient value (Temp[i + α, j’]) of the pixel located vertically from the pixel position (i + α, j) by using the gradient filter (fracFilterβ), thereby obtaining the horizontal gradient value at (i + α, j + β)
[0260]
Number
[0261] Referring to FIG. 6B, in the reference picture, in order to determine the vertical gradient value of the reference pixel, filtering is performed on the integer pixels by using a plurality of one-dimensional filters. Motion compensation in pixel group units is additional motion compensation performed after motion compensation in block units. Therefore, in the process of motion compensation in block units, the reference position of the reference block of the current block indicated by the motion vector is also a fractional pixel position, and motion compensation in pixel group units is also performed on the reference pixels in the reference block at the fractional pixel position. Therefore, considering determining the gradient value of the pixel at the fractional pixel position, filtering may also be performed.
[0262] Referring to FIG. 6B, first, the video decoding apparatus 100 can perform filtering on pixels located horizontally or vertically from the peripheral integer pixels of the reference pixel within the reference picture using a first one-dimensional filter. Similarly, the video decoding apparatus 100 can perform filtering on adjacent pixels located in a different column or row from the reference pixel using the first one-dimensional filter. The video decoding apparatus 100 can generate a gradient value in the vertical direction of the reference pixel by performing filtering on the value generated by the filtering using a second one-dimensional filter.
[0263] For example, when the position of the reference pixel is the position of a fractional pixel of (x + α, y + β) (x and y are integers, and α and β are fractions), for the horizontal integer pixels (x, y) and (x - 1, y - 1), (x + 1, y + 1),…, (x + M Min , y + M Min ), (x + M Max , y + M max )(both MMin and MMmax are integers), the first one-dimensional filter is used, and filtering is performed as shown in the following mathematical formula (16).
[0264]
Equation
[0265] That is, the first one-dimensional filter is also an interpolation filter for determining the pixel value at the horizontal fractional pixel position α. offset5 is an offset for preventing rounding errors, and shift5 means the number of inverse scaling bits.
[0266] Temp[i + α, j] means the pixel value at the fractional pixel position (i + α, j). Temp[i + α, j’] (where j’ is an integer from j + M min to j + M max is also determined by substituting j with j’ in Equation (16).
[0267] Next, the video decoder 100 can perform filtering on the pixel value at the pixel position (i + α, j) and the pixel value at the pixel position (i + α, j’) using a second two-dimensional filter as shown in the following Equation (17).
[0268]
Equation
[0269] That is, according to Equation (17), the video decoder 100 uses a gradient filter (gradFilter β ) to perform filtering on the pixel value at (i + α, j) (Temp[i + α, j]) and the pixel value at the pixel position vertically located from (i + α, j) (Temp[i + α, j’]), thereby obtaining the vertical gradient value at (i + α, j + β)
[0270]
Equation
[0271] Previously, it was described that the interpolation filter is first applied, and then the gradient filter is applied to determine the gradient value in the vertical direction. However, it is not limited thereto. The gradient filter can also be applied first, and then the interpolation filter can be applied to determine the gradient value in the horizontal direction. Hereinafter, an embodiment in which the gradient filter is applied first, and then the interpolation filter is applied to determine the gradient value in the vertical direction will be described.
[0272] For example, when the position of the reference pixel is the position of a fractional pixel at (x + α, y + β) (x and y are integers, and α and β are fractions), the integer pixels (x, y) and (x, y - 1), (x, y + 1), …, (x, y + M Min ),(x, y + M max )(M Min 、M max are all integers), the first one-dimensional filter is used, and filtering is performed as shown in the following formula (18).
[0273]
Equation
[0274] That is, the first one-dimensional filter is also an interpolation filter for determining the gradient value in the vertical direction of the pixel at the position where the vertical component of the pixel position is a fraction. offset7 is an offset for preventing rounding errors, and shift7 means the number of inverse scaling bits.
[0275] Temp[i,j+β] represents the vertical gradient value at the pixel position (i,j+β). Temp[i’,j+β] (where i’ is an integer from i+M min to i+M max is also determined by substituting i with i’ in Equation (18).
[0276] Next, the video decoder 100 can perform filtering as shown in the following Equation (19) using a second two-dimensional filter on the vertical gradient value at the pixel position (i,j+β) and the vertical gradient value at the pixel position (i’,j+β).
[0277]
Equation
[0278] That is, according to Equation (19), the video decoder 100 performs filtering using the interpolation filter (fracFilter α ) on the vertical gradient value at (i,j+β) (Temp[i,j+β]) and the vertical gradient value of the pixel located horizontally from the pixel position (i,j+β) (Temp[i’,j+β]), thereby obtaining the vertical gradient value at (i+α,j+β)
[0279]
Equation
[0280] According to one embodiment, in the video decoder 100, the gradient values in the horizontal and vertical directions at (i + α, j + β) are also determined by the above-described various combinations of filters. For example, to determine the horizontal gradient value, an interpolation filter for determining the vertical pixel values is used with a first one-dimensional filter, and a gradient filter for the horizontal gradient value may also be used with a second one-dimensional filter. To determine the vertical gradient value, a gradient filter for determining the vertical gradient value is used with a first one-dimensional filter, and an interpolation filter for determining the horizontal pixel values may also be used with a second one-dimensional filter.
[0281] FIGS. 7A to 7E are tables showing pixel values at fractional pixel positions in fractional pixel units according to one embodiment, and filter coefficients of filters used to determine gradient values in the horizontal and vertical directions.
[0282] FIGS. 7A and 7B are tables showing filter coefficients of filters for determining horizontal or vertical gradient values at fractional pixel positions in 1 / 4 pel units.
[0283] As described above, one-dimensional gradient filters and one-dimensional interpolation filters may also be used to determine horizontal or vertical gradient values. Referring to FIG. 7A, the filter coefficients of the one-dimensional gradient filter are illustrated. At this time, a 6-tap filter may also be used as the gradient filter. The filter coefficients of the gradient are also coefficients scaled by about 2^4. M min means the difference between the position of the pixel farthest from the center integer pixel in the negative number direction applied to the filter and the position of the center integer pixel, based on the center integer pixel, and M maxmeans the difference between the position of the farthest integer pixel in the positive direction applied to the filter and the position of the central integer pixel, with the central integer pixel as a reference. For example, in the horizontal direction, the gradient filter coefficients for obtaining the horizontal gradient value of a pixel with a fractional pixel position α = 1 / 4 are also {4, -17, -36, 60, -15, -4}. The gradient filter coefficients for obtaining the horizontal gradient values of pixels with fractional pixel positions α = 0, 1 / 2, 3 / 4 in the horizontal direction are also determined with reference to FIG. 7A.
[0284] Referring to FIG. 7B, the filter coefficients of the one-dimensional interpolation filter are illustrated. At this time, a 6-tap filter may also be used as the interpolation filter. The filter coefficients of the interpolation filter are also coefficients scaled by 2^6. M min means the difference between the position of the farthest integer pixel in the negative direction applied to the filter and the position of the central integer pixel, with the central integer pixel as a reference, and M max means the difference between the position of the farthest integer pixel in the positive direction applied to the filter and the position of the central integer pixel, with the central integer pixel as a reference.
[0285] FIG. 7C is a table showing the filter coefficients of the one-dimensional interpolation filter used to determine the pixel value at a fractional pixel position in 1 / 4-pel units.
[0286] As described above, two identical one-dimensional interpolation filters may also be used in the horizontal and vertical directions to determine the pixel value at a fractional pixel position.
[0287] Referring to FIG. 7C, the filter coefficients of the one-dimensional interpolation filter are illustrated. At this time, the one-dimensional interpolation filter is also a 6-tap filter. The filter coefficients of the gradient are also coefficients scaled by 2^6. M min means the difference between the position of the farthest integer pixel in the negative direction applied to the filter and the position of the central integer pixel, with the central integer pixel as a reference, and Mmax means the difference between the position of the outermost integer pixel in the positive integer pixel direction applied to the filter and the position of the central integer pixel, with the central integer pixel as a reference.
[0288] FIG. 7D is a table showing filter coefficients of a filter used to determine horizontal or vertical gradient values at fractional pixel positions in 1 / 16 pel units.
[0289] As described above, a one-dimensional gradient filter and a one-dimensional interpolation filter may also be used to determine horizontal or vertical gradient values. Referring to FIG. 7D, the filter coefficients of the one-dimensional gradient filter are illustrated. At this time, a 6-tap filter may also be used as the gradient filter. The filter coefficients of the gradient are also coefficients scaled by about 2^4. For example, in the horizontal direction, the gradient filter coefficients for obtaining the horizontal gradient value of a pixel where the fractional pixel position α is 1 / 16 are also {8, -32, -13, 50, -18, 5}. In the horizontal direction, the gradient filter coefficients for obtaining the horizontal gradient values of pixels where the fractional pixel positions α are 0, 1 / 8, 3 / 16, 1 / 4, 5 / 16, 3 / 8, 7 / 16, 1 / 2 are also determined using FIG. 7D. On the other hand, the gradient filter coefficients for obtaining the horizontal gradient values of pixels where the fractional pixel positions α are 9 / 16, 5 / 8, 11 / 16, 3 / 4, 13 / 16, 7 / 8, 15 / 16 are also determined using the symmetry of the filter coefficients with reference to α = 1 / 2. That is, based on α = 1 / 2 disclosed in FIG. 7D, the filter coefficients of the left fractional pixel positions are used, and the filter coefficients of the remaining right fractional pixel positions are also determined based on α = 1 / 2. For example, the filter coefficients at α = 15 / 16 are also determined using the filter coefficients {8, -32, -13, 50, -18, 5} at α = 1 / 16, which is a position symmetric with reference to α = 1 / 2. That is, the filter coefficients at α = 15 / 16 are determined by arranging the filter coefficients {8, -32, -13, 50, -18, 5} in reverse order, as {5, -18, 50, -13, -32, 8}.
[0290] Referring to FIG. 7E, the filter coefficients of the one-dimensional interpolation filter are illustrated. At this time, a 6-tap filter may also be used as the interpolation filter. The filter coefficients of the interpolation filter are also coefficients scaled by about 2^6. For example, in the horizontal direction, the one-dimensional interpolation filter coefficients for obtaining the horizontal pixel value of a pixel with a fractional pixel position α of 1 / 16 are also {1, -3, 64, 4, -2, 0}. In the horizontal direction, the interpolation filter coefficients for obtaining the horizontal pixel values of pixels with fractional pixel positions α = 0, 1 / 8, 3 / 16, 1 / 4, 5 / 16, 3 / 8, 7 / 16, 1 / 2 are also determined using FIG. 7E. On the other hand, in the horizontal direction, the interpolation filter coefficients for obtaining the horizontal pixel values of pixels with fractional pixel positions α of 9 / 16, 5 / 8, 11 / 16, 3 / 4, 13 / 16, 7 / 8, 15 / 16 are also determined using the symmetry of the filter coefficients with respect to the α = 1 / 2 reference. That is, the filter coefficients of the left fractional pixel positions with respect to the α = 1 / 2 reference disclosed in FIG. 7E are used, and the filter coefficients of the remaining right fractional pixel positions with respect to the α = 1 / 2 reference are also determined. For example, the filter coefficients at α = 15 / 16 are also determined using the filter coefficients {1, -3, 64, 4, -2, 0} at α = 1 / 16, which is the position symmetric with respect to the α = 1 / 2 reference. That is, the filter coefficients at α = 15 / 16 are also determined by arranging the filter coefficients {1, -3, 64, 4, -2, 0} in reverse order as {0, -2, 4, 64, -3, 1}.
[0291] FIG. 8A is a reference diagram for explaining the process of determining the horizontal displacement vector and the vertical displacement vector related to pixels according to an embodiment.
[0292] Referring to FIG. 8A, a window (Ωij) 800 of a predetermined size has a size of (2M + 1) * (2N + 1) (M and N are integers) centered on the pixel P(i, j) to be bi-directionally predicted in the current block.
[0293] Pixels of the current block predicted bidirectionally within the window are P(i’, j’) (when i - M ≤ i’ ≤ i + M and j - N ≤ j’ ≤ j + N, (i’, j’) ∈ Ωij), the pixel value of the first reference pixel of the first reference picture 810 corresponding to the bidirectionally predicted pixel P(i’, j’) of the current block is P0(i’, j’), the pixel value of the second reference pixel of the second reference picture 820 corresponding to the bidirectionally predicted pixel P(i’, j’) of the current block is P1(i’, j’), the gradient value in the horizontal direction of the first reference pixel is
[0294]
Number
[0295]
Number
[0296]
Number
[0297]
Number
[0298]
Number
[0299] The difference value △i’j’ between the first displacement-corresponding pixel PA’ and the second displacement-corresponding pixel PB’ is determined as in the following mathematical formula (21).
[0300]
Number
[0301]
Number
[0302]
Number
[0303]
Number
[0304] [Mathematics] By using the following, two linear equations with variables Vx(i,j) and Vy(i,j) can be obtained as shown in the following equation (24).
[0305] [Mathematics] In equation (24), s1 to s6 are as shown in the following equation (25).
[0306] [Mathematics] By solving the simultaneous equations of equation (24), using Kramer's formulas, the values of Vx(i,j) and Vy(i,j) can be solved as τ*Vx(i,j)=-det1 / det and τ*Vy(i,j)=-det2 / det. Here, det1 = s3*s5 - s2*s6, det2 = s1*s6 - s3*s4, and det = s1*s5 - s2*s2.
[0307] In the horizontal direction, minimization is first performed, and then minimization is performed in the vertical direction, and a simplified solution of the above equation is also determined. That is, for example, assuming that only the horizontal displacement vector changes, in the first equation of equation (24), Vy can be assumed to be 0. Therefore, the equation: τVx = s3 / s1 is also determined.
[0308] Then, by arranging the second equation of equation (24) using the equation: τVx = s3 / s1, the equation: τVy = (s6 - τVx*S2) / s5 is also determined.
[0309] Here, the gradient value
[0310] [Mathematics] It can also be scaled without changing the resulting values of Vx(i,j) and Vy(i,j). However, it is assumed that no overflow occurs and no rounding error occurs.
[0311] In the process of obtaining Vx(i,j) and Vy(i,j), adjustment parameters r and m are also introduced to prevent multiplication operations from being performed by 0 or very small values.
[0312] For the sake of convenience, assume that Vx(i,j) and Vy(i,j) are opposite to the directions shown in FIG. 3A. For example, based on the directions of Vx(i,j) and Vy(i,j) shown in FIG. 3A, Vx(i,j) and Vy(i,j) derived by Equation (24) have the same magnitude as Vx(i,j) and Vy(i,j) determined to be opposite to the direction of FIG. 3A, except that their values differ only by a sign difference.
[0313] The first displacement corresponding pixel PA’ and the second displacement corresponding pixel PB’ are also determined as in the following Equation (26). At this time, PA’ and PB’ are also determined using the first linear term of the local Taylor expansion.
[0314]
Equation
[0315]
Equation
[0316] [Number]
[0317] [Number] Φ(Vx, Vy) is a function with Vx and Vy as intermediate variables, and the maximum or minimum value is also determined by calculating the values that make the partial derivatives of Φ(Vx, Vy) with respect to Vx and Vy equal to 0, as shown in the following mathematical formula (30).
[0318] [Number] That is, Vx and Vy are also determined as the Vx and Vy that minimize the Φ(Vx, Vy) value. To solve the optimization problem, first minimization is performed in the vertical direction, and then minimization is performed in the horizontal direction. By this minimization, Vx is also determined as shown in the following mathematical formula (31).
[0319] [Number] Here, the clip3(x, y, z) function outputs x if z < x, y if z > y, and z if x < z < y. According to mathematical formula (31), when s1 + r > m, Vx is clip3(-thBIO, thBIO, -s3 / (s1 + r)), and when s1 + r > m is not true, Vx is also 0.
[0320] Through the above minimization, Vy is also determined as in the following mathematical formula (32).
[0321]
Number
[0322] At this time, s1, s2, s3, and s5 are also determined as in the following mathematical formula (33). s4 can have the same value as s2.
[0323]
Number
[0324]
Number
[0325] However, not limited thereto, the adjustment parameters r, m, and thBIO may also have their values determined based on information about the adjustment parameters obtained from the bitstream. At this time, the information about the adjustment parameters is also included in the slice header, picture parameter set, sequence parameter set, and various forms of high-level syntax carriers.
[0326] Also, the adjustment parameters may be determined depending on the availability of temporally different bidirectional prediction. For example, thBIO when temporally different bidirectional prediction is available diff is larger than thBIO when bidirectional prediction that is temporally the same as each other is available, and the magnitude of thBIO same is also twice the size of thBIO diff same .
[0327] FIG. 8B is a reference diagram for explaining the process of determining the horizontal displacement vector and vertical displacement vector related to a pixel group according to an embodiment.
[0328] Referring to FIG. 8B, a window (Ωij) 810 of a predetermined size is centered on a pixel group 820 of KxK size of a plurality of pixels that are not pixels bi-directionally predicted in the current block, and has a size of (2M + K + 1) * (2N + K + 1) (where M and N are integers).
[0329] At this time, the difference from FIG. 8A is that the window size becomes larger. Except for this, the horizontal displacement vector and the vertical displacement vector related to the pixel group can be determined in the same manner.
[0330] FIG. 8C is a reference diagram for explaining a process of determining a horizontal displacement vector and a vertical displacement vector related to pixels according to an embodiment.
[0331] Referring to FIG. 8C, the video decoding apparatus 100 can determine a horizontal displacement vector and a vertical displacement vector for each pixel. Therefore, a displacement vector 835 per unit time for each pixel can be determined. At this time, the horizontal displacement vector Vx[i, j] and the vertical displacement vector Vy[i, j] of the displacement vector 835 for each pixel are also determined by the following mathematical formula (35). Here, i and j represent the x component and the y component of the pixel coordinates. Here, σ1[i, j], σ2[i, j], σ3[i, j], σ5[i, j] and σ6 [i, j] are also s1, s2, s3, s5 and s6 in the mathematical formula (33), respectively.
[0332] [Equation] FIG. 8D is a reference diagram for explaining a process of determining a horizontal displacement vector and a vertical displacement vector related to a pixel group according to an embodiment.
[0333] Referring to FIG. 8D, for each pixel included in each of the pixel groups 840, the video decoding apparatus 100 can determine σ1[i,j], σ2[i,j], σ3[i,j], σ5[i,j], and σ6[i,j] as in the above formula (35).
[0334] The video decoding apparatus 100 can determine the horizontal displacement vector Vx[i,j] related to the pixel group 840 by using σ1[i,j] and σ3[i,j] of the pixels as in the following formula (36). Here, i and j represent the x-component and y-component of the upper left coordinate of the pixel group.
[0335]
Equation
[0336] The video decoding apparatus 100 can determine the vertical displacement vector Vy[i,j] related to the pixel group 840 by using σ2[i,j], σ5[i,j], σ6[i,j], and Vx[i,j] of the pixels as in the following formula (37). Here, the horizontal displacement vector Vx[i,j] is also the value determined by the formula (36).
[0337]
Equation
[0338] On the other hand, although FIG. 8D was previously referred to and the case where the pixel group has a 2×2 size was described in detail, it is not limited thereto, and the pixel group can have an L×L (L is an integer) size.
[0339] At this time, the size L of the pixel group is also determined as in the following mathematical formula (38). W and H respectively represent the width and height of the current block.
[0340]
Equation
[0341]
Table 1
[0342] FIG. 9A is a drawing for explaining a process of determining a gradient value in a horizontal direction or a vertical direction by adding an offset and performing inverse scaling after filtering is performed according to an embodiment.
[0343] Referring to FIG. 9A, the video decoding apparatus 100 can perform filtering on pixels whose predetermined direction components are at integer positions using a first one-dimensional filter and a second one-dimensional filter, and determine gradient values in the horizontal or vertical direction. However, for pixels whose predetermined direction components are at integer positions, the values obtained by filtering using the first one-dimensional filter or the second one-dimensional filter may also fall outside a predetermined range. Such a phenomenon is called an overflow phenomenon. The coefficients of the one-dimensional filter are also determined as integers in order to perform integer operations instead of inaccurate and complex fractional operations. The coefficients of the one-dimensional filter may also be scaled because they are determined as integers. If filtering is performed using the scaled coefficients of the one-dimensional filter, integer operations become possible. However, when compared with the case where filtering is performed using an unscaled one-dimensional filter, the magnitude of the filtered value becomes larger, and an overflow phenomenon may occur. Therefore, in order to prevent such an overflow phenomenon, inverse scaling may also be performed after filtering is performed using the one-dimensional filter. At this time, the inverse scaling may include performing bit shifting to the right by the number of inverse scaling bits to the right. The number of inverse scaling bits is also determined in consideration of the maximum number of bits of the register for the filtering operation and the maximum number of bits of the temporary buffer for storing the filtering operation result while maximizing the accuracy of the calculation. In particular, the number of inverse scaling bits is also determined based on the internal bit depth, the number of scaling bits for the interpolation filter, and the number of scaling bits for the gradient filter.
[0344] In the following, in order to determine the gradient value in the horizontal direction, first, an interpolation filter in the vertical direction is used to perform filtering on pixels at integer positions to generate an interpolation filtering value in the vertical direction. Then, in the process of performing filtering on the interpolation filtering value in the vertical direction using a gradient filter in the horizontal direction, the content of performing inverse scaling will be described.
[0345] According to the foregoing formula (12), the video decoder 100 can first perform filtering on the pixels at integer positions using a vertical interpolation filter in order to determine the gradient value in the horizontal direction. At this time, shift1 is also determined as b - 8. At this time, b is also the internal bit depth of the input video. Hereinafter, referring to Table 2, when actual inverse scaling is performed based on the shift1, the bit depth of the register (Reg Bitdepth) and the bit depth of the temporary buffer (Temp Bitdepth) will be described.
[0346]
Table 2
[0347]
Number
[0348] For example, assuming that the 1 / 4 pel unit gradient filter fracFilter disclosed in FIG. 7C is used, FilterSumPos is also 88, and FilterSumNeg is also -24.
[0349] The Ceiling(x) function is also a function that outputs the smallest integer among the integers greater than or equal to error x. Offset1 is an offset value added to the filtered value to prevent rounding errors that may occur in the process of inverse scaling using shift1, and offset1 is also determined to be 2^(shift1 - 1).
[0350] Referring to Table 2, when the internal bit depth b is 8, the register bit depth (RegBitdpeth) is 16; when the internal bit depth b is 9, the register bit depth is 17; hereinafter, when the internal bit depths b are 10, 11, 12, and 16, the register bit depths are also 18, 19, and 24. If the register used for filtering is a 32-bit register, since the register bit depths in Table 2 do not exceed 32, no overflow phenomenon will occur.
[0351] Similarly, when the internal bit depths b are 8, 9, 10, 11, 12, and 16, the temporary buffer bit depth (TempBitDepth) is 16 in all cases. If the temporary buffer used to save the filtered and inverse-scaled value is a 16-bit buffer, since the temporary buffer bit depths in Table 2 are 16 and do not exceed 16, no overflow phenomenon will occur.
[0352] According to Equation (12), in order to determine the gradient value in the horizontal direction, the video decoding apparatus 100 first uses an interpolation filter in the vertical direction to filter the pixels at integer positions to generate the interpolation filtering value in the vertical direction. Then, according to Equation (13), a gradient filter in the horizontal direction can be used to filter the interpolation filtering value in the vertical direction. At this time, shift2 is also determined to be p + q - shift1. Here, p represents the number of bits scaled for the interpolation filter including the filter coefficients shown in FIG. 7C, and q represents the number of bits scaled for the gradient filter including the filter coefficients shown in FIG. 7A. For example, p is 6 and q is also 4. Therefore, shift2 = 18 - b.
[0353] The reason why shift2 is determined as described above is that when the filter coefficients are upscaled and when the filter coefficients are not upscaled, in order for the final filtering result value to be the same, shift1 + shift2, which is the sum of the bits to be inverse-scaled, must be the same as the sum (p + q) of the bits upscaled for the filter.
[0354] Hereinafter, with reference to Table 3, when actual inverse scaling is performed based on the shift2, the bit depth of the register (Reg Bitdepth) and the bit depth of the temporary buffer (Temp Bitdepth) will be described.
[0355]
Table 3
[0356]
Number
[0357] offset2 is an offset value added to the filtered value to prevent a rounding error that may occur in the process of performing inverse scaling using shift2, and offset2 is also determined to be 2^(shift2 - 1).
[0358] shift1 and shift2 are determined as described above, but are not limited thereto. shift1 and shift2 may be determined in various ways such that the sum of shift1 and shift2 is the same as the sum of the scaling bits for the filter. At this time, the shift1 value and the shift2 value are also determined on the premise that no overflow phenomenon occurs. shift1 and shift2 are also determined based on the internal bit depth of the input video and the scaling bits for the filter.
[0359] However, it is not always necessary to determine shift1 and shift2 such that the sum of shift1 and shift2 is the same as the sum of the scaling bits for the filter. For example, shift1 is also determined to be d - 8, but shift2 is also determined to be a fixed number.
[0360] If shift1 is the same as before and shift2 is 7, which is an integer, in Table 3 mentioned previously, OutMax, OutMin, and Temp Bitdepth may also be different. Hereinafter, referring to Table 4, the bit depth (Temp Bitdepth) of the temporary buffer will be described.
[0361]
Table 4
[0362] When shift2 is a constant, without using the scaled filter coefficients, the result value when filtering is performed and the result value when inverse scaling is performed after filtering using the scaled filter coefficients may be different. In that case, it will be easily understood by those skilled in the art that additional inverse scaling must be performed.
[0363] Previously, to determine the horizontal gradient value, first a vertical interpolation filter was used to filter the pixels at integer positions to generate the vertical interpolation filtering values. Then, the horizontal gradient filter was used, and the content of performing inverse scaling during the process of filtering the vertical interpolation filtering values was described. However, it is not limited thereto. It will be easily understood by those skilled in the art that when filtering is performed on pixels with a predetermined direction component being an integer to determine the horizontal and vertical gradient values by various one-dimensional filter combinations, similarly, inverse scaling can be performed.
[0364] FIG. 9B is a drawing for explaining a process of determining a gradient value in a horizontal or vertical direction by adding an offset and performing inverse scaling after filtering according to another embodiment.
[0365] Referring to FIG. 9B, the video decoding apparatus 100 can perform filtering with the fractional pixels and integer pixels of the reference picture as inputs. Here, it is assumed that the fractional pixels of the reference picture are values determined by applying one-dimensional filters in the horizontal and vertical directions to the integer pixels of the reference picture.
[0366] The video decoding device 100 can perform filtering using a one-dimensional filter in the horizontal or vertical direction on pixels with a fractional position component in a predetermined direction and integer pixels, and determine the gradient value in the horizontal or vertical direction. However, for pixels with a fractional predetermined direction position component and integer pixels, the value obtained by filtering using the first one-dimensional filter is also outside a predetermined range. Such a phenomenon is called an overflow phenomenon. The coefficients of the one-dimensional filter are also determined to be integers in order to perform integer operations instead of inaccurate and complex fractional operations. The one-dimensional filter coefficients are also scaled because they are determined as integers. If filtering is performed using the scaled one-dimensional filter coefficients, integer operations become possible. However, when compared with the case where filtering is performed using an unscaled one-dimensional filter, the magnitude of the value obtained by filtering becomes larger, and an overflow phenomenon may occur. Therefore, in order to prevent the overflow phenomenon, inverse scaling may also be performed after filtering is performed using the one-dimensional filter. At this time, the inverse scaling may include performing bit shifting to the right by the number of inverse scaling bits (shift1) to the right. The number of inverse scaling bits is also determined in consideration of the maximum number of bits of the register for the filtering operation and the maximum number of bits of the temporary buffer for storing the filtering operation result while maximizing the calculation accuracy. In particular, the number of inverse scaling bits is also determined based on the internal bit depth and the number of scaling bits for the gradient filter.
[0367] FIG. 9C is a drawing for explaining the range necessary for determining the horizontal displacement vector and the vertical displacement vector in the process of performing pixel unit motion compensation on the current block.
[0368] Referring to FIG. 9C, in the process of performing pixel - unit motion compensation on the reference block 910 corresponding to the current block, the video decoder 100 can use the window 920 around the pixel 915 located at the upper - left end of the reference block 910 to determine the displacement vector per unit time in the horizontal direction and the displacement vector per unit time in the vertical direction at the pixel 915. At this time, by using the pixel values and gradient values of the pixels located outside the range of the reference block 910, the displacement vector per unit time in the horizontal direction or the vertical direction can be determined. In a similar manner, in the process of determining the horizontal displacement vector and the vertical displacement vector for the pixels located at the boundary of the reference block 910, the video decoder 100 will determine the pixel values and gradient values of the pixels located outside the range of the reference block 910. Therefore, the video decoder 100 can use the block 925 with a range larger than the reference block 910 to determine the horizontal displacement vector and the displacement vector per unit time in the vertical direction. For example, if the size of the current block is AxB and the pixel - by - pixel window size is (2M + 1)x(2N + 1), then the size of the range for determining the horizontal displacement vector and the vertical displacement vector is also (A + 2M)x(B + 2N).
[0369] FIGS. 9D and 9E are diagrams for explaining the range of the area used in the process of performing pixel - unit motion compensation according to various embodiments.
[0370] Referring to FIG. 9D, in the process of performing motion compensation on a pixel-by-pixel basis, the video decoding apparatus 100 can determine a horizontal displacement vector for each pixel and a displacement vector per unit time in the vertical direction included in the reference block 930 based on the block 935 in the range extended by the size of the window of the pixels located at the boundary of the reference block 930. However, in the process of determining the displacement vector per unit time in the horizontal and vertical directions, the video decoding apparatus 100 requires the pixel value and the gradient value of the pixels located in the block 935. At this time, an interpolation filter or a gradient filter can be used to obtain the aforementioned pixel value and gradient value. In the process of using the interpolation filter or the gradient filter for the boundary pixels of the block 935, the pixel values of the surrounding pixels can be used, and thus, the pixels located outside the block boundary can be used. Therefore, motion compensation on a pixel-by-pixel basis is also performed using the block 940 in the additionally extended range by a value obtained by subtracting 1 from the number of taps of the interpolation filter or the gradient filter. Therefore, when the size of the block is NxN, the pixel-by-pixel window size is (2M + 1)x(2M + 1), and the length of the interpolation filter or the gradient filter is T, it is also (N + 2M + T - 1)x(N + 2M + T - 1).
[0371] Referring to FIG. 9E, in the process of performing motion compensation on a pixel-by-pixel basis, the video decoding apparatus 100 does not expand the reference block according to the size of the window of the pixels located at the boundary of the reference block 945, but uses the pixel values and gradient values of the pixels located within the reference block 945 to determine the horizontal displacement vector for each pixel and the displacement vector per unit time in the vertical direction. Specifically, the process for the video decoding apparatus 100 to determine the displacement vector per unit time in the horizontal direction and the displacement vector per unit time in the vertical direction without expanding the reference block will be described with reference to FIG. 9E. However, in order to obtain the pixel values and gradient values of the pixels, an interpolation filter or a gradient filter of the reference block 945 will be used, and pixel-by-pixel motion compensation may also be performed using the expanded block 950. Therefore, when the size of the block is NxN, the window size for each pixel is (2M + 1)x(2M + 1), and the length of the interpolation filter or the length of the gradient filter is T, it is also (N + T - 1)x(N + T - 1).
[0372] FIG. 9F is a drawing for explaining the process of determining the horizontal displacement vector and the vertical displacement vector without expanding the reference block.
[0373] Referring to FIG. 9F, in the case of pixels located outside the boundary of the reference block 955, the video decoding apparatus 100 adjusts the position of the pixel to the position of the available pixel at the nearest position among the pixels located within the boundary of the reference block 955, and the pixel value and the gradient value of the pixel located outside the boundary are also determined as the pixel value and the gradient value of the available pixel at the nearest position. At this time, the video decoding apparatus 100 adjusts the position of the pixel located outside the boundary of the reference block 955 according to the mathematical formula:
[0374]
Equation
[0375]
Number
[0376] Here, i' is the x - coordinate value of the pixel, j' is the y - coordinate value of the pixel, and H and W represent the height and width of the reference block. At this time, it is assumed that the position of the upper - left corner of the reference block is (0,0). If the position of the upper - left corner of the reference block is (xP,yP), the position of the final pixel is also (i'+xP,j'+yP).
[0377] Also, referring to FIG. 9D, in the block 935 expanded according to the size of the pixel - by - pixel window, the position of the pixel located outside the boundary of the reference block 930 is adjusted to the position of the pixel adjacent to the inside of the boundary of the reference block 930. As shown in FIG. 9E, the video decoding apparatus 100 can determine the horizontal displacement vector for each pixel and the displacement vector per unit time in the vertical direction within the reference block 945 with the pixel values and gradient values of the reference block 945.
[0378] Therefore, the video decoding apparatus 100 does not expand the reference block 945 according to the pixel - by - pixel window size, and performs pixel - unit motion compensation, thereby reducing the number of memory accesses for pixel - value reference, reducing the number of multiplication operations, and also reducing the complexity of the operations.
[0379] When the video decoder 100 performs motion compensation in units of blocks (as when operating according to the HEVC standard), when performing pixel - unit motion compensation together with block expansion based on the window size, and when performing pixel - unit motion compensation without block expansion, the memory access operations and multiplication operations are performed in the video decoder 100 as shown in Table 5 below for each case. At this time, assume that the signal (interpolation) filter length is 8, the gradient filter length is 6, the block size is NxN, and the pixel - by - pixel window size 2M + 1 is 5.
[0380]
Table 5
[0381] When performing pixel - level motion compensation without block expansion, since there is no block expansion, like the block - based motion compensation according to the HEVC standard, (N + 7)×(N + 7) reference samples are required. And because it is bidirectional motion prediction compensation and two reference blocks are used, in pixel - level motion compensation performed without block expansion, as shown in Table 5, 2×(N + 7)×(N + 7) memory accesses are required.
[0382] On the other hand, in the block - based motion compensation according to the HEVC standard, since an 8 - tap interpolation filter is used for each sample, the samples required for the first horizontal interpolation are (N + 7)×N samples. The samples required for the second vertical interpolation are N×N samples. On the other hand, the number of multiplication operations required per 8 - tap interpolation filter is 8. And because it is bidirectional motion prediction compensation and two reference blocks are used, in the block - based motion compensation according to the HEVC standard, as shown in Table 5, 2×8×{(N + 7)×N+N×N} multiplication operations are required.
[0383] When performing pixel - level motion compensation with block expansion, in order to perform pixel - level motion compensation, since the block size is expanded, an 8 - tap interpolation filter is used for the expanded (N + 4)×(N + 4) - sized block. To determine the pixel values at sub - pixel - level positions, as shown in Table 5, a total of 2×8×{(N + 4 + 7)×(N + 4)+(N + 4)×(N + 4)} multiplication operations are required.
[0384] On the one hand, when performing pixel - unit motion compensation along with block expansion, in order to determine the gradient value in the horizontal or vertical direction, a 6 - tap gradient filter and a 6 - tap interpolation filter are used. Since the size of the block is expanded, for the expanded (N + 4)x(N + 4) - sized block, to determine the gradient value using a 6 - tap interpolation filter and a gradient filter, as shown in Table 5, a total of 2*6*{(N + 4+5)x(N + 4)+(N + 4)x(N + 4)}*2 multiplication operations are required.
[0385] When performing pixel - unit motion compensation without block expansion, since there is no block expansion, (N + 7)x(N + 7) reference samples are required like block - unit motion compensation according to the HEVC standard. Since it is bidirectional motion prediction compensation and two reference blocks are used, in pixel - unit motion compensation performed without block expansion, for an NxN - sized block, to determine the pixel value at the position of a fractional - pixel unit using an 8 - tap interpolation filter, as shown in Table 5, 2*8*{(N + 7)xN+NxN} multiplication operations are required.
[0386] On the other hand, when performing pixel - unit motion compensation without block expansion, in order to determine the gradient value in the horizontal or vertical direction, a 6 - tap gradient filter and a 6 - tap interpolation filter are used. For an NxN - sized block, to determine the gradient value using a 6 - tap interpolation filter and a gradient filter, as shown in Table 5, a total of 2*6*{(N + 5)xN+NxN}*2 multiplication operations are required.
[0387] FIGs. 9G to 9I are diagrams for explaining the process for determining a horizontal displacement vector and a vertical displacement vector without expanding a reference block according to other embodiments.
[0388] As described above with reference to FIG. 9D, in the process of performing motion compensation on a pixel-by-pixel basis, the video decoding apparatus 100 can determine a horizontal displacement vector for each pixel and a displacement vector per unit time in the vertical direction included in the reference block 930 based on the block 935 in the range extended by the size of the window of the pixels located at the boundary of the reference block 930. For example, when the window size is (2M + 1)×(2M + 1), if the window is applied to the pixels located at the boundary of the reference block 930, the video decoding apparatus 100 can refer to the pixel values and gradient values of the pixels that are M pixels outside the reference block 930 and determine the horizontal displacement vector and the vertical displacement vector for each pixel.
[0389] Hereinafter, according to another embodiment, a method for determining values (s1 to s6 in Equation (33)) for determining a horizontal displacement vector and a vertical displacement vector for each pixel by using only the pixel values and gradient values of the reference block corresponding to the current block without referring to the pixel values and gradient values of the pixels outside the reference block by the video decoding apparatus 100 will be described. Here, it is assumed that the window size is 5×5. For convenience, the pixels located in the horizontal direction will be described with the current pixel as the center. It will be easily understood by those skilled in the art that the weighted values are also determined in the same manner for the pixels located in the vertical direction with the current pixel as the center.
[0390] For each pixel included in the window centered on the current pixel for which the video decoding apparatus 100 is to determine a displacement vector in the horizontal or vertical direction, the pixel values P0(i’, j’) and P1(i’, j’), and the horizontal or vertical gradient values
[0391]
Number
[0392] At this time, by making the weighting values the same for each pixel, the result values of the operations performed for each pixel are multiplied by the weighting values, and these values are combined to determine the values s1 to s6 for determining the horizontal displacement vector and the vertical displacement vector for each pixel.
[0393] Referring to FIG. 9G, in the current block 960, in the process of determining the horizontal and vertical displacement vectors for the current pixel 961, the video decoder 100 can determine the weighting values so that the window pixels have the same value 1. The video decoder 100 multiplies the result values calculated for each pixel by the weighting values determined for each pixel, adds all the result values, and can determine the values s1 to s6 for determining the horizontal displacement vector and the vertical displacement vector for the current pixel.
[0394] Referring to FIG. 9H, in the current block 970, when the current pixel 971 is immediately adjacent to the boundary of the current block 970, the video decoder 100 can determine that the weighting value of the pixel in contact with the boundary of the block 970 located outside the boundary of the current block 970 is 3. The video decoder 100 can determine that the weighting value of the other pixel 973 is 1.
[0395] Referring to FIG. 9H, when the current pixel 981 is located near the boundary of the current block 980 within the current block 980 (when the current pixel is about one pixel away from the boundary), the video decoder 100 may determine that the weighting value for the pixel 982 located outside the boundary of the current block 980 is 0, and the weighting value for the pixel 983 adjacent to the boundary of the current block 980 is 2. The video decoder 100 may determine that the weighting value for the other pixel 984 is 1.
[0396] As described with reference to FIGS. 9G to 9I, the video decoder 100 assigns different weighting values to pixels within the window according to the position of the current pixel, so as not to use the pixel values and gradient values of the pixels located outside the reference block corresponding to the current block, but to use the pixel values and gradient values of the pixels located inside the reference block, and to be able to determine values s1 to s6 for determining the displacement vectors in the horizontal and vertical directions for each pixel.
[0397] FIG. 9J is a drawing for explaining the process of determining the horizontal displacement vector and the vertical displacement vector for each pixel by referring to the pixel values and gradient values of the reference block and applying the exponential smoothing technique vertically and horizontally without expanding the block according to an embodiment.
[0398] Referring to FIG. 9J, the video decoder 100 calculates, for each pixel included in the current block 990, the pixel values P0(i’, j’) and P1(i’, j’) of the corresponding reference pixels included in the corresponding reference block, the horizontal or vertical gradient values of the corresponding reference pixels included in the pixels of the corresponding reference block
[0399]
Equation
[0400] The video decoding apparatus 100 combines the function operation execution result value of the current pixel and the function operation execution result values of the surrounding pixels, and determines values s1 to s6 (σ k(k=1,2,3,4,5,6) ) for determining the displacement vectors in the horizontal and vertical directions of the current pixel. That is, the values s1 to s6 for determining the horizontal displacement vector and the vertical displacement vector of the current pixel are also expressed by the weighted average of the operation execution values for the current pixel and the surrounding pixels as shown in the following mathematical formula (41). At this time, the position coordinates of the pixels included in the window Ω are (i’, j’). Also, W[i’, j’] means the weighted value related to the pixels included in the window Ω. Here, the size of the window Ω is also (2M + 1) x (2M + 1) (M is an integer). Also, the function A k [i’, j’] is the pixel value P0(i’, j’) and P1(i’, j’) (I[i’, j’](0, 1)) of the corresponding reference pixels related to the pixel at the (i’, j’) position included in the window Ω, the horizontal or vertical gradient value of the corresponding reference pixels included in the pixels of the corresponding reference block
[0401]
Equation
[0402]
Equation
[0403]
Table 6
[0404]
Equation
[0405]
Equation
[0406] Video decoder 100 is A k In order to determine the weighted average related to A k [i’, j’], the exponential smoothing technique can be applied in the up, down, left, and right directions for averaging.
[0407] Referring to FIG. 9J, first, the video decoder 100 can apply the exponential smoothing technique to A k [i’, j’] in the left - right direction. In the following, the video decoder 100 will apply the exponential smoothing technique to A k [i’, j’] in the left - right direction and will determine the weighted average value related to A k [i’, j’] in detail.
[0408] First, the video decoder 100 will explain the content of applying the exponential smoothing technique for averaging in the right - hand direction. The video decoder 100 can average A k [i’, j’] in the right - hand direction according to the following pseudo - code 1. At this time, H represents the height of the current block, W represents the width of the current block, and Stride represents the distance between one line and the next line in the one - dimensional array. That is, the two - dimensional array A[i, j] can also be represented by the one - dimensional array A[i + j*Stride]. [Pseudo - code 1]
[0409]
Equation
[0410] Next, the content in which the video decoding device 100 applies the exponential smoothing technique to perform averaging in the left direction will be described. The video decoding device 100 can perform averaging in the left direction on A k [i’, j’] using the following pseudo code 2. At this time, H represents the height of the current block, W represents the width of the current block, and Stride represents the distance between one line and the next line in the one-dimensional array. That is, the two-dimensional array A[i, j] can also be represented by the one-dimensional array A[i + j * Stride]. [Pseudo code 2]
[0411]
Number
[0412] Next, the content in which the video decoding device 100 applies the exponential smoothing technique to perform averaging in the downward direction will be described. The video decoding device 100 can perform averaging in the downward direction on A k [i’, j’] using the following pseudo code 3. At this time, H represents the height of the current block, W represents the width of the current block, and Stride represents the distance between one line and the next line in the one-dimensional array. That is, the two-dimensional array A[i, j] can also be represented by the one-dimensional array A[i + j * Stride]. [Pseudo code 3]
[0413]
Number
[0414]
Number
[0415] Therefore, the video decoder 100 uses the exponential smoothing technique for the current block 990 and performs averaging in the up, down, left, and right directions to obtain σ, which is the weighted average necessary for determining the horizontal or vertical displacement vector for each pixel. k[i, j] can be determined. That is, the video decoding apparatus 100 can determine the displacement vector in the horizontal direction or the vertical direction for each pixel without referring to the pixel values and gradient values of the reference block corresponding to the block 996 obtained by expanding the current block 990, and by referring only to the pixel values and gradient values of the reference block corresponding to the current block 990.
[0416] FIG. 9K is a diagram for explaining a process of determining pixel values of reference pixels in a reference block and gradient values in the horizontal and vertical directions by using a filter in order to perform motion compensation related to a current block according to an embodiment.
[0417] Referring to FIG. 9K, the video decoding apparatus 100 can perform pixel unit and block unit motion compensation related to the current block by using the pixel values and gradient values of the reference pixels in the reference block corresponding to the current block. Therefore, in order to perform pixel unit and block unit motion compensation related to the current block, the pixel values and gradient values of the reference pixels in the reference block corresponding to the current block must be determined. At this time, the units of the pixel values and gradient values of the reference pixels in the reference block are also in sub-pixel units. For example, the units of the pixel values and gradient values of the reference pixels in the reference block are also in 1 / 16 pixel (1 / 16 pel) units.
[0418] The video decoding apparatus 100 can perform filtering on the pixel values of the integer pixels of the reference block in order to determine the pixel values and gradient values of the reference pixels in the reference block in sub-pixel units.
[0419] First, the video decoding apparatus 100 can determine the pixel values of the reference pixels in the reference block by applying a horizontal 8-tap signal filter (also referred to as an interpolation filter) and a vertical 8-tap signal filter to the pixel values of the integer pixels of the reference pixels.
[0420] The video decoding device 100 can determine the pixel values of reference pixels having position components in units of sub-pixels in the horizontal direction by performing filtering by applying an 8-tap signal filter in the horizontal direction to the integer pixel values of the reference block, and can store it in a buffer. The video decoding device 100 can determine the pixel values of reference pixels having position components in units of sub-pixels in the vertical direction by applying an 8-tap signal filter in the vertical direction to the pixel values of reference pixels having position components in integer units in the vertical direction.
[0421]
[0422] That is, the video decoding device 100 can determine the pixel values of reference pixels having position components in units of sub-pixels in the vertical direction by performing filtering by applying a 6-tap signal filter in the vertical direction to the integer pixel values of the reference block, and can store it in a buffer.
[0423] The video decoding device 100 can determine the horizontal gradient values of reference pixels having position components in units of sub-pixels in the horizontal direction by applying a 6-tap gradient filter in the horizontal direction to the pixel values of reference pixels having position components in integer units in the horizontal direction.
[0424]
[0425] That is, the video decoder 100 can determine the gradient value of a reference pixel having a position component in units of fractional pixels in the vertical direction by performing filtering by applying a 6-tap gradient filter in the vertical direction to the integer pixel values of the reference block, and save it in a buffer.
[0426] In the horizontal direction, the video decoder 100 can apply a 6-tap signal filter in the horizontal direction to the vertical gradient value of a reference pixel having a position component in integer units, and also determine the vertical gradient value of a reference pixel having a position component in units of fractional pixels in the horizontal direction.
[0427] That is, in order for the video decoder 100 to determine the pixel value of a reference pixel in the reference block, the horizontal gradient value of the reference pixel in the reference block, and the vertical gradient value of the reference pixel in the reference block, respectively, two one-dimensional filters are applied, and at this time, a multiplication operation is also performed between the coefficients of each filter and the value related to the corresponding pixel. For example, in order to determine the horizontal gradient value of a reference pixel in the reference block, two 6-tap signal / gradient filters are used, and a total of 12 multiplication operations are also performed per pixel. Also, in order to determine the vertical gradient value of a reference pixel in the reference block, two 6-tap signal / gradient filters are used, and a total of 12 multiplication operations are also performed per pixel.
[0428] FIG. 9L is a drawing for explaining, according to another embodiment, the process of using a filter to determine the pixel value of a reference pixel in the reference block and the gradient values in the horizontal and vertical directions in order to perform motion compensation related to the current block.
[0429] Referring to FIG. 9L, the video decoding apparatus 100 can perform motion compensation in terms of pixels and blocks related to the current block by using the pixel values and gradient values of the reference pixels in the reference block corresponding to the current block. Therefore, in order to perform motion compensation in terms of pixels and blocks related to the current block, the pixel values and gradient values of the reference pixels in the reference block corresponding to the current block must be determined. At this time, the units of the pixel values and gradient values of the reference pixels in the reference block are also in fractional pixel units. For example, the units of the pixel values and gradient values of the reference pixels in the reference block are also in 1 / 16 pixel (1 / 16 pel) units.
[0430] The video decoding apparatus 100 can perform filtering on the pixel values of the integer pixels of the reference block in order to determine the pixel values and gradient values of the reference pixels in the reference block in fractional pixel units.
[0431] Different from the description with reference to FIG. 9K, the video decoding apparatus 100 first applies a horizontal 8-tap signal filter and a vertical 8-tap signal filter to the integer pixel values of the reference block to determine the pixel values of the reference pixels in the reference block, and then applies a horizontal 5-tap gradient filter to the pixel values of the reference pixels in the reference block to determine the horizontal gradient values of the reference pixels in the reference block. Also, the video decoding apparatus 100 can apply a vertical 5-tap gradient filter to the pixel values of the reference pixels in the reference block to determine the vertical gradient values of the reference pixels in the reference block.
[0432] Video decoding apparatus 100 applies two one-dimensional signal (interpolation) filters to determine the pixel value of a reference pixel having a position in units of fractional pixels, and then applies two one-dimensional gradient filters in parallel to the pixel value of the reference pixel at the position in units of fractional pixels, whereby a horizontal gradient value or a vertical gradient value of the reference pixel in the reference block can be determined.
[0433] Video decoding apparatus 100 applies a 5-tap horizontal gradient filter (filter coefficients are {9, -48, 0, 48, 9}; however, the filter coefficients are not limited thereto) to the pixel value of the reference pixel as shown in the following mathematical formula (42), and can determine the horizontal gradient value Ix(k) of the reference pixel in the reference block. Here, k can have a value of 0 or 1, and can indicate reference pictures 0 and 1, respectively. I(k)[i,j] is also the pixel value of the reference pixel in the reference block at the (i,j) position. i is the horizontal position component of the pixel, and j is the vertical position component of the pixel, and the unit thereof is also in units of fractional pixels.
[0434]
Equation
[0435] In addition, the video decoder 100 can apply a 5-tap vertical gradient filter (filter coefficients are {9, -48, 0, 48, 9}; however, it is not limited thereto) to the pixel values of the reference pixels as in the following mathematical formula (43) to determine the vertical gradient value Iy(k) of the reference pixels in the reference block. Here, k can have a value of 0 or 1, and can respectively indicate reference pictures 0 and 1. I(k)[i,j] is also the pixel value of the reference pixel in the reference block at the (i,j) position. i is the horizontal position component of the pixel, and j is the vertical position component of the pixel, and its unit is also in fractional pixel units.
[0436] [Number] Therefore, the video decoder 100 can determine the vertical gradient value Iy(k) of the reference pixels in the reference block by performing only two multiplication operations per sample.
[0437] When the video decoder 100 performs pixel unit motion compensation together with block expansion by the window size, and when performing pixel unit motion compensation, reducing the gradient filter length, the memory access operations and multiplication operations in the video decoder 100 are also performed as shown in Table 7 below for each case. At this time, assume that the length T of the signal filter is 8, the length T of the gradient filter is 6, the length T of the simplified gradient filter is 5, the size of the block is NxN, and the pixel-by-pixel window size 2M + 1 is 5.
[0438] [Table 7] That is, the video decoder 100 performs multiplication operations twice per one-dimensional gradient filter according to formulas (41) and (42), and the gradient filter is applied to two reference blocks. Since the gradient filter is applied to the (N + 4) x (N + 4) reference block extended based on the window size, a total of 2 * 2 * {(N + 4) x (N + 4)} * 2 multiplication operations are also performed to determine the gradient values of the reference pixels in the horizontal and vertical directions.
[0439] When the video decoder 100 performs pixel unit motion compensation together with block expansion according to the window size, without block expansion, the length of the gradient filter is reduced, and when performing motion compensation in pixel group units, the number of memory accesses and multiplication operations involved in each case are such that the memory access operation, multiplication operation, and multiplication operation are also performed in the video decoder 100 as shown in Table 8 below. At this time, it is assumed that the length T of the signal filter is 8, the length T of the gradient filter is 6, the length T of the reduced gradient filter is 5, the size of the pixel group is LxL, the size of the block is NxN, and the pixel-by-pixel window size 2M + 1 is 5.
[0440]
Table 8
[0441] Hereinafter, with reference to FIGS. 10 to 23, a method for determining a data unit that can be used in the process of a video decoding apparatus 100 decoding video according to an embodiment will be described. The operation of the video encoding apparatus 150 is similar to various embodiments related to the operation of the video decoding apparatus 100 described later, or may be an operation opposite thereto.
[0442] FIG. 10 illustrates, according to an embodiment, a process in which the video decoding apparatus 100 divides a current encoding unit and determines at least one encoding unit.
[0443] According to an embodiment, the video decoding apparatus 100 can use the block form information to determine the form of the encoding unit, and can use the division form information to determine how the encoding unit is divided. That is, the division method of the encoding unit indicated by the division form information is also determined by what block form the block form information used by the video decoding apparatus 100 indicates.
[0444] According to an embodiment, the video decoding apparatus 100 can use the block form information indicating that the current encoding unit is square. For example, the video decoding apparatus 100 can determine whether to divide the square encoding unit, divide it vertically, divide it horizontally, or divide it into four encoding units according to the division form information. Referring to FIG. 10, when the block form information of the current encoding unit 1000 indicates a square form, the video decoding apparatus 100 can determine not to divide the encoding unit 1010a having the same size as the current encoding unit 1000 according to the division form information indicating non-division, or can determine the encoding units 1010b, 1010c, 1010d divided based on the division form information indicating a predetermined division method.
[0445] Referring to FIG. 10, according to one embodiment, the video decoding apparatus 100 can determine two coded units 1010b obtained by vertically splitting the current coded unit 1000 based on split form information indicating that it is split in the vertical direction. The video decoding apparatus 100 can determine two coded units 1010c obtained by horizontally splitting the current coded unit 1000 based on split form information indicating that it is split in the horizontal direction. The video decoding apparatus 100 can determine four coded units 1010d obtained by vertically and horizontally splitting the current coded unit 1000 based on split form information indicating that it is split in both the vertical and horizontal directions. However, the split form in which a square coded unit is split is not construed as being limited to the foregoing form, and may include various forms that can be indicated by the split form information. A predetermined split form in which a square coded unit is split will be specifically described below through various embodiments.
[0446] FIG. 11 illustrates, according to one embodiment, the process by which the video decoding apparatus 100 splits a coded unit having a non-square form and determines at least one coded unit.
[0447] According to an embodiment, the video decoder 100 can utilize block form information indicating that the current coding unit is non-square. The video decoder 100 can determine whether to divide the non-square current coding unit or divide it in a predetermined method based on the division form information. Referring to FIG. 11, when the block form information of the current coding unit 1100 or 1150 indicates a non-square form, the video decoder 100 can determine whether to divide the coding unit 1110 or 1160 having the same size as the current coding unit 1100 or 1150 according to the division form information indicating non-division, or determine the divided coding units 1120a, 1120b, 1130a, 1130b, 1130c, 1170a, 1170b, 1180a, 1180b, 1180c based on the division form information indicating a predetermined division method. The predetermined division method for dividing the non-square coding unit will be specifically described through various embodiments below.
[0448] According to an embodiment, the video decoder 100 can utilize the division form information to determine the form in which the coding unit is divided. In that case, the division form information can indicate the number of at least one coding unit generated by dividing the coding unit. Referring to FIG. 11, when the division form information indicates that the current coding unit 1100 or 1150 is divided into two coding units, the video decoder 100 can divide the current coding unit 1100 or 1150 based on the division form information and determine the two coding units 1120a, 11420b or 1170a, 1170b included in the current coding unit.
[0449] According to one embodiment, when the video decoding device 100 divides the current encoded unit 1100 or 1150 in a non-square form based on the division form information, it can consider the position of the long side of the non-square current encoded unit 1100 or 1150 and divide the current encoded unit. For example, the video decoding device 100 can consider the form of the current encoded unit 1100 or 1150 and divide the current encoded unit 1100 or 1150 in the direction of dividing the long side of the current encoded unit 1100 or 1150 to determine a plurality of encoded units.
[0450] According to one embodiment, when the division form information indicates dividing the encoded unit into an odd number of blocks, the video decoding device 100 can determine the odd number of encoded units included in the current encoded unit 1100 or 1150. For example, when the division form information indicates dividing the current encoded unit 1100 or 1150 into three encoded units, the video decoding device 100 can divide the current encoded unit 1100 or 1150 into three encoded units 1130a, 1130b, 1130c, 1180a, 1180b, 1180c. According to one embodiment, the video decoding device 100 can determine the odd number of encoded units included in the current encoded unit 1100 or 1150, and the sizes of the determined encoded units are not all the same. For example, among the determined odd number of encoded units 1130a, 1130b, 1130c, 1180a, 1180b, 1180c, the size of a predetermined encoded unit 1130b or 1180b can also be different from the sizes of the other encoded units 1130a, 1130c, 1180a, 1180c. That is, the encoded units determined by dividing the current encoded unit 1100 or 1150 can have multiple types of sizes, and in some cases, the odd number of encoded units 1130a, 1130b, 1130c, 1180a, 1180b, 1180c can each have different sizes from each other.
[0451] According to one embodiment, when the segmentation form information indicates that the encoding unit is segmented into an odd number of blocks, the video decoding apparatus 100 can determine the odd number of encoding units included in the current encoding unit 1100 or 1150. Further, the video decoding apparatus 100 can impose a predetermined restriction on at least one of the odd number of encoding units generated by segmentation. Referring to FIG. 11, the video decoding apparatus 100 can make the decoding process for the centrally located encoding units 1130b and 1180b among the three encoding units 1130a, 1130b, 1130c, 1180a, 1180b, and 1180c generated by dividing the current encoding unit 1100 or 1150 different from those of the other encoding units 1130a, 1130c, 1180a, and 1180c. For example, the video decoding apparatus 100 can, for the centrally located encoding units 1130b and 1180b, unlike the other encoding units 1130a, 1130c, 1180a, and 1180c, restrict them so that they are not further segmented, or restrict them so that they are segmented only a predetermined number of times.
[0452] FIG. 12 illustrates a process in which the video decoding apparatus 100 segments an encoding unit based on at least one of block form information and segmentation form information according to one embodiment.
[0453] According to one embodiment, the video decoding apparatus 100 can determine whether to divide the square first encoding unit 1200 into encoding units or not based on at least one of the block form information and the division form information. According to one embodiment, when the division form information indicates dividing the first encoding unit 1200 in the horizontal direction, the video decoding apparatus 100 can divide the first encoding unit 1200 in the horizontal direction and determine the second encoding unit 1210. According to one embodiment, the first encoding unit, the second encoding unit, and the third encoding unit used are terms used to understand the pre- and post-division relationship between the encoding units. For example, if the first encoding unit is divided, the second encoding unit is determined, and if the second encoding unit is divided, the third encoding unit is also determined. In the following, the relationship among the first encoding unit, the second encoding unit, and the third encoding unit used is also understood to be based on the aforementioned features.
[0454] According to an embodiment, the video decoding apparatus 100 can determine whether to divide the determined second encoded unit 1210 into encoded units or not based on at least one of the block form information and the division form information. Referring to FIG. 12, the video decoding apparatus 100 divides the first encoded unit 1200 based on at least one of the block form information and the division form information, and determines whether to divide the non-square second encoded unit 1210 into at least one third encoded unit 1220a, 1220b, 1220c, 1220d or not to divide the second encoded unit 1210. The video decoding apparatus 100 can obtain at least one of the block form information and the division form information. The video decoding apparatus 100 divides the first encoded unit 1200 based on at least one of the obtained block form information and the division form information, and can divide a plurality of second encoded units (for example, 1210) in various forms. The second encoded unit 1210 can also be divided by the way the first encoded unit 1200 is divided based on at least one of the block form information and the division form information. According to an embodiment, when the first encoded unit 1200 is divided into the second encoded unit 1210 based on at least one of the block form information and the division form information related to the first encoded unit 1200, the second encoded unit 1210 is also divided into a third encoded unit (for example, 1220a, 1220b, 1220c, 1220d) based on at least one of the block form information and the division form information related to the second encoded unit 1210. That is, the encoded units can also be recursively divided based on at least one of the division form information and the block form information related to each encoded unit. Therefore, in the non-square encoded unit, a square encoded unit is determined, and such a square encoded unit is recursively divided, and a non-square encoded unit is also determined. Referring to FIG. 12, in the odd-numbered third encoded units 1220b, 1220c, 1220d in which the non-square second encoded unit 1210 is divided and determined, a predetermined encoded unit (for example, the encoded unit located in the middle, or the square encoded unit) can also be recursively divided.According to one embodiment, a square-shaped third encoding unit 1220c, which is one of an odd number of third encoding units 1220b, 1220c, 1220d, may be divided in the horizontal direction and also divided into a plurality of fourth encoding units. A non-square-shaped fourth encoding unit 1240, which is one of the plurality of fourth encoding units, may also be divided into a plurality of encoding units. For example, the non-square-shaped fourth encoding unit 1240 is further divided into an odd number of encoding units 1250a, 1250b, 1250c.
[0455] The method used for the recursive division of the encoding unit will be described later through various embodiments.
[0456] According to one embodiment, the video decoding device 100 can determine whether to divide each of the third encoding units 1220a, 1220b, 1220c, and 1220d into encoding units or not divide the second encoding unit 1210 based on at least one of the block form information and the division form information. According to one embodiment, the video decoding device 100 can divide the non-square second encoding unit 1210 into an odd number of third encoding units 1220b, 1220c, and 1220d. The video decoding device 100 can impose a predetermined restriction on a predetermined third encoding unit among the odd number of third encoding units 1220b, 1220c, and 1220d. For example, the video decoding device 100 can restrict that the encoding unit 1220c located in the middle among the odd number of third encoding units 1220b, 1220c, and 1220d is not further divided, or restrict that it must be divided a settable number of times. Referring to FIG. 12, the video decoding device 100 can restrict that the encoding unit 1220c located in the middle among the odd number of third encoding units 1220b, 1220c, and 1220d included in the non-square second encoding unit 1210 is not further divided, or is divided into a predetermined division form (for example, only divided into four encoding units, or divided in a form corresponding to the divided form of the second encoding unit 1210), or is divided only a predetermined number of times (for example, divided only n times, n>0). However, the restriction on the encoding unit 1220c located in the middle is only a simple embodiment, and thus should not be construed as being limited to the foregoing embodiments, and should be construed as including various restrictions such that the encoding unit 1220c located in the middle is decoded differently from the different encoding units 1220b and 1220d.
[0457] According to one embodiment, the video decoding device 100 can obtain at least one of the block form information and the division form information used to divide the current encoding unit at a predetermined position within the current encoding unit.
[0458] FIG. 13 illustrates, according to one embodiment, a method for a video decoding apparatus 100 to determine a predetermined encoding unit among an odd number of encoding units. Referring to FIG. 13, at least one of the block form information and the segmentation form information of the current encoding unit 1300 may also be obtained from a sample at a predetermined position (for example, the sample 1340 located in the middle) among a plurality of samples included in the current encoding unit 1300. However, the predetermined position within the current encoding unit 1300 from which at least one of such block form information and segmentation form information is obtained should not be construed as being limited to the middle position illustrated in FIG. 13, and the predetermined position should be interpreted to include various positions within the current encoding unit 1300 (for example, the uppermost end, the lowermost end, the left side, the right side, the upper left end, the lower left end, the upper right end, or the lower right end, etc.). The video decoding apparatus 100 can obtain at least one of the block form information and the segmentation form information obtained from the predetermined position, and determine whether to divide the current encoding unit into encoding units of various forms and sizes or not to divide it.
[0459] According to one embodiment, the video decoding apparatus 100 can select one of the encoding units when the current encoding unit is divided into a predetermined number of encoding units. There are various methods for selecting one of the plurality of encoding units, and the description related to such methods will be described later through various embodiments below.
[0460] According to one embodiment, the video decoding apparatus 100 can divide the current encoding unit into a plurality of encoding units and determine the encoding unit at a predetermined position.
[0461] FIG. 13 illustrates, according to one embodiment, a method for a video decoding apparatus 100 to determine an encoding unit at a predetermined position among an odd number of encoding units.
[0462] According to an embodiment, the video decoding apparatus 100 can utilize information indicating the position of each of an odd number of encoding units to determine the encoding unit located in the middle of the odd number of encoding units. Referring to FIG. 13, the video decoding apparatus 100 can divide the current encoding unit 1300 and determine an odd number of encoding units 1320a, 1320b, 1320c. The video decoding apparatus 100 can utilize information related to the positions of the odd number of encoding units 1320a, 1320b, 1320c to determine the middle encoding unit 1320b. For example, the video decoding apparatus 100 can determine the position of the encoding unit 1320b located in the middle by determining the positions of the encoding units 1320a, 1320b, 1320c based on information indicating the positions of predetermined samples included in the encoding units 1320a, 1320b, 1320c. Specifically, the video decoding apparatus 100 can determine the position of the encoding unit 1320b located in the middle by determining the positions of the encoding units 1320a, 1320b, 1320c based on information indicating the positions of the samples 1330a, 1330b, 1330c at the upper left ends of the encoding units 1320a, 1320b, 1320c.
[0463] According to one embodiment, the information indicating the positions of the left upper samples 1330a, 1330b, and 1330c included in the encoding units 1320a, 1320b, and 1320c respectively may include information related to the positions or coordinates within the pictures of the encoding units 1320a, 1320b, and 1320c. According to one embodiment, the information indicating the positions of the left upper samples 1330a, 1330b, and 1330c included in the encoding units 1320a, 1320b, and 1320c respectively may include information indicating the widths or heights of the encoding units 1320a, 1320b, and 1320c included in the current encoding unit 1300, and such widths or heights correspond to the information indicating the differences between the coordinates within the pictures of the encoding units 1320a, 1320b, and 1320c. That is, the video decoding apparatus 100 can determine the centrally located encoding unit 1320b by directly using the information related to the positions or coordinates within the pictures of the encoding units 1320a, 1320b, and 1320c, or by using the information related to the widths or heights of the encoding units corresponding to the difference values between the coordinates.
[0464] According to one embodiment, the information indicating the position of the sample 1330a at the upper left end of the upper encoding unit 1320a can indicate (xa, ya) coordinates, the information indicating the position of the sample 1330b at the upper left end of the middle encoding unit 1320b can indicate (xb, yb) coordinates, and the information indicating the position of the sample 1330c at the upper left end of the lower encoding unit 1320c can indicate (xc, yc) coordinates. The video decoding device 100 can determine the middle encoding unit 1320b by using the coordinates of the samples 1330a, 1330b, and 1330c at the upper left ends respectively included in the encoding units 1320a, 1320b, and 1320c. For example, when the coordinates of the samples 1330a, 1330b, and 1330c at the upper left end are sorted in ascending or descending order, the encoding unit 1320b including the coordinates (xb, yb) of the sample 1330b located in the middle can be determined as the encoding unit located in the middle among the encoding units 1320a, 1320b, and 1320c determined by dividing the current encoding unit 1300. However, the coordinates indicating the positions of the samples 1330a, 1330b, and 1330c at the upper left end can indicate the coordinates of the absolute positions within the picture, and further, based on the position of the sample 1330a at the upper left end of the upper encoding unit 1320a, the (dxb, dyb) coordinates indicating the relative position of the sample 1330b at the upper left end of the middle encoding unit 1320b and the (dxc, dyc) coordinates indicating the relative position of the sample 1330c at the upper left end of the lower encoding unit 1320c can also be used. Also, as the information indicating the position of the sample included in the encoding unit, the method of determining the encoding unit at a predetermined position by using the coordinates of the sample should not be construed as being limited to the above-described method, but must be construed as various arithmetic methods that can use the coordinates of the sample.
[0465] According to one embodiment, the video decoding apparatus 100 can divide the current encoding unit 1300 into a plurality of encoding units 1320a, 1320b, 1320c, and can select an encoding unit according to a predetermined criterion from among the encoding units 1320a, 1320b, 1320c. For example, the video decoding apparatus 100 can select the encoding unit 1320b having a different size among the encoding units 1320a, 1320b, 1320c.
[0466] According to one embodiment, the video decoding apparatus 100 can use the (xa, ya) coordinates, which are information indicating the position of the sample 1330a at the upper left end of the upper-end encoding unit 1320a, the (xb, yb) coordinates, which are information indicating the position of the sample 1330b at the upper left end of the middle encoding unit 1320b, and the (xc, yc) coordinates, which are information indicating the position of the sample 1330c at the upper left end of the lower-end encoding unit 1320c, to determine the width or height of each of the encoding units 1320a, 1320b, 1320c. The video decoding apparatus 100 can use the coordinates (xa, ya), (xb, yb), (xc, yc), which are coordinates indicating the positions of the encoding units 1320a, 1320b, 1320c, to determine the size of each of the encoding units 1320a, 1320b, 1320c.
[0467] According to an embodiment, the video decoding apparatus 100 can determine the width of the upper encoding unit 1320a as xb - xa and the height as yb - ya. According to an embodiment, the video decoding apparatus 100 can determine the width of the middle encoding unit 1320b as xc - xb and the height as yc - yb. According to an embodiment, for the lower encoding unit, the video decoding apparatus 100 can determine the width or height using the width or height of the current encoding unit and the widths and heights of the upper encoding unit 1320a and the middle encoding unit 1320b. Based on the determined widths and heights of the encoding units 1320a, 1320b, and 1320c, the video decoding apparatus 100 can determine an encoding unit having a size different from that of other encoding units. Referring to FIG. 13, the video decoding apparatus 100 can determine the middle encoding unit 1320b, which has a size different from those of the upper encoding unit 1320a and the lower encoding unit 1320c, as the encoding unit at a predetermined position. However, the process by which the video decoding apparatus 100 determines an encoding unit having a size different from that of different encoding units is only one embodiment of determining the encoding unit at a predetermined position using the size of the encoding unit determined based on sample coordinates. Thus, various processes of comparing the sizes of the encoding units determined by predetermined sample coordinates and determining the encoding unit at a predetermined position may also be used.
[0468] However, the sample positions considered for determining the position of the encoding unit should not be construed as being limited to the upper left side described above, and it may also be construed that information related to any sample position included in the encoding unit may be used.
[0469] According to an embodiment, the video decoding apparatus 100 can select an encoding unit at a predetermined position from among an odd number of encoding units determined by dividing the current encoding unit, considering the form of the current encoding unit. For example, if the current encoding unit is a non-square shape whose width is longer than its height, the video decoding apparatus 100 can determine an encoding unit at a predetermined position along the horizontal direction. That is, the video decoding apparatus 100 can determine one of the encoding units having different positions in the horizontal direction and place a restriction on the encoding unit. If the current encoding unit is a non-square shape whose height is longer than its width, the video decoding apparatus 100 can determine an encoding unit at a predetermined position along the vertical direction. That is, the video decoding apparatus 100 can determine one of the encoding units having different positions in the vertical direction and place a restriction on the encoding unit.
[0470] According to an embodiment, the video decoding apparatus 100 can use information indicating the position of each of the even number of encoding units to determine an encoding unit at a predetermined position from among the even number of encoding units. The video decoding apparatus 100 can divide the current encoding unit to determine an even number of encoding units, and use the information related to the positions of the even number of encoding units to determine an encoding unit at a predetermined position. The specific process related thereto is also a process corresponding to the process of determining an encoding unit at a predetermined position (for example, the middle position) from among the odd number of encoding units described with reference to FIG. 13, and thus will be omitted.
[0471] According to an embodiment, when a non-square current encoding unit is divided into a plurality of encoding units, in order to determine an encoding unit at a predetermined position among the plurality of encoding units, predetermined information about the encoding unit at the predetermined position can be used in the division process. For example, the video decoding apparatus 100 can use at least one of the block form information and the division form information stored in the samples included in the middle encoding unit in the division process in order to determine the encoding unit located in the middle among the plurality of encoding units into which the current encoding unit is divided.
[0472] Referring to FIG. 13, the video decoder 100 can divide the current encoding unit 1300 into a plurality of encoding units 1320a, 1320b, 1320c based on at least one of the block form information and the division form information, and can determine the encoding unit 1320b located in the middle among the plurality of encoding units 1320a, 1320b, 1320c. Furthermore, the video decoder 100 can determine the encoding unit 1320b located in the middle in consideration of the position where at least one of the block form information and the division form information is obtained. That is, at least one of the block form information and the division form information of the current encoding unit 1300 is obtained from the sample 1340 located in the middle of the current encoding unit 1300, and based on at least one of the block form information and the division form information, when the current encoding unit 1300 is divided into a plurality of encoding units 1320a, 1320b, 1320c, the encoding unit 1320b including the sample 1340 can be determined as the encoding unit located in the middle. However, the information used to determine the encoding unit located in the middle is not construed as being limited to at least one of the block form information and the division form information, and various types of information may also be used in the process of determining the encoding unit located in the middle.
[0473] According to an embodiment, the predetermined information for identifying the encoding unit at a predetermined position is also obtained from predetermined samples included in the encoding unit to be determined. Referring to FIG. 13, in a plurality of encoding units 1320a, 1320b, 1320c determined by dividing the current encoding unit 1300, the video decoding apparatus 100 can use at least one of the block form information and the division form information obtained from the samples at the predetermined position (for example, the sample located in the middle among the encoding units divided into a plurality) within the current encoding unit 1300 in order to determine the encoding unit at the predetermined position. That is, the video decoding apparatus 100 can determine the sample at the predetermined position in consideration of the block form of the current encoding unit 1300, and the video decoding apparatus 100 can determine the encoding unit 1320b including the sample from which predetermined information (for example, at least one of the block form information and the division form information) is obtained among the plurality of encoding units 1320a, 1320b, 1320c determined by dividing the current encoding unit 1300, and can impose a predetermined restriction. Referring to FIG. 13, according to an embodiment, the video decoding apparatus 100 can determine the sample 1340 located in the middle of the current encoding unit 1300 as the sample from which the predetermined information is obtained, and the video decoding apparatus 100 can impose a predetermined restriction on the encoding unit 1320b including such a sample 1340 during the decoding process. However, the position of the sample from which the predetermined information is obtained is not construed as being limited to the aforementioned position, and is also construed as the sample at any position included in the encoding unit 1320b to be determined for imposing the restriction.
[0474] According to one embodiment, the position of the sample from which predetermined information is acquired is also determined by the form of the current encoding unit 1300. According to one embodiment, the block form information can determine whether the form of the current encoding unit is square or non-square, and based on the form, the position of the sample from which predetermined information is acquired can be determined. For example, the video decoding device 100 uses at least one of the information related to the width and the information related to the height of the current encoding unit, and determines a sample located on the boundary that divides at least one of the width and the height of the current encoding unit into half as the sample from which predetermined information is acquired. As another example, when the block form information related to the current encoding unit indicates that it is non-square, the video decoding device 100 can determine one of the samples adjacent to the boundary that divides the long side of the current encoding unit into half as the sample from which predetermined information is acquired.
[0475] According to one embodiment, when the video decoding device 100 divides the current encoding unit into a plurality of encoding units, in order to determine the encoding unit at a predetermined position among the plurality of encoding units, it can use at least one of the block form information and the division form information. According to one embodiment, the video decoding device 100 can acquire at least one of the block form information and the division form information from a sample at a predetermined position included in the encoding unit, and the video decoding device 100 can divide the plurality of encoding units generated by dividing the current encoding unit by using at least one of the division form information and the block form information acquired from the samples at the predetermined positions included in each of the plurality of encoding units. That is, the encoding unit is also recursively divided by using at least one of the block form information and the division form information acquired from the samples at the predetermined positions included in each encoding unit. The recursive division process of the encoding unit has been described with reference to FIG. 12, so detailed description will be omitted.
[0476] According to one embodiment, the video decoding apparatus 100 can divide a current encoding unit and determine at least one encoding unit, and can determine the order in which such at least one encoding unit is decoded by a predetermined block (e.g., the current encoding unit).
[0477] FIG. 14 illustrates, according to one embodiment, the order in which a plurality of encoding units are processed when the video decoding apparatus 100 divides a current encoding unit and determines a plurality of encoding units.
[0478] According to one embodiment, the video decoding apparatus 100 can divide the first encoding unit 1400 vertically to determine the second encoding units 1410a and 1410b according to block form information and division form information, or divide the first encoding unit 1400 horizontally to determine the second encoding units 1430a and 1430b, or divide the first encoding unit 1400 both vertically and horizontally to determine the second encoding units 1450a, 1450b, 1450c, and 1450d.
[0479] Referring to FIG. 14, the video decoding apparatus 100 can determine the order such that the second encoding units 1410a and 1410b determined by vertically dividing the first encoding unit 1400 are processed in the horizontal direction (1410c). The video decoding apparatus 100 can determine the processing order of the second encoding units 1430a and 1430b determined by horizontally dividing the first encoding unit 1400 in the vertical direction (1430c). For the second encoding units 1450a, 1450b, 1450c, and 1450d determined by dividing the first encoding unit 1400 both vertically and horizontally, the video decoding apparatus 100 can determine a predetermined order (e.g., raster scan order or z scan order (1450e)) in which the encoding units located in the next row are processed after the encoding units located in one row are processed.
[0480] According to one embodiment, the video decoder 100 can recursively divide an encoding unit. Referring to FIG. 14, the video decoder 100 can divide a first encoding unit 1400 to determine a plurality of encoding units 1410a, 1410b, 1430a, 1430b, 1450a, 1450b, 1450c, 1450d, and can recursively divide each of the determined plurality of encoding units 1410a, 1410b, 1430a, 1430b, 1450a, 1450b, 1450c, 1450d. The method of dividing the plurality of encoding units 1410a, 1410b, 1430a, 1430b, 1450a, 1450b, 1450c, 1450d also corresponds to the method of dividing the first encoding unit 1400. Thereby, the plurality of encoding units 1410a, 1410b, 1430a, 1430b, 1450a, 1450b, 1450c, 1450d are each independent and can also be divided into a plurality of encoding units. Referring to FIG. 14, the video decoder 100 can divide the first encoding unit 1400 in the vertical direction to determine second encoding units 1410a, 1410b, and can further determine whether to independently divide each of the second encoding units 1410a, 1410b or not divide them.
[0481] According to one embodiment, the video decoder 100 can divide the left second encoding unit 1410a in the horizontal direction into third encoding units 1420a, 1420b, and the right second encoding unit 1410b is not divided.
[0482] According to one embodiment, the processing order of the encoding units is also determined based on the division process of the encoding units. In other words, the processing order of the divided encoding units is also determined based on the processing order of the encoding unit immediately before being divided. The video decoding apparatus 100 can determine the order in which the third encoding units 1420a and 1420b determined by dividing the second encoding unit 1410a on the left are processed independently of the second encoding unit 1410b on the right. Since the second encoding unit 1410a on the left is divided in the horizontal direction and the third encoding units 1420a and 1420b are determined, the third encoding units 1420a and 1420b are also processed in the vertical direction (1420c). Also, since the order in which the second encoding unit 1410a on the left and the second encoding unit 1410b on the right are processed corresponds to the horizontal direction 1410c, after the third encoding units 1420a and 1420b included in the second encoding unit 1410a on the left are processed in the vertical direction (1420c), the second encoding unit 1410b on the right is also processed. Since the foregoing content is for explaining the process in which the processing order of the encoding units is determined by the encoding units before division respectively, it should not be construed as being limited to the foregoing embodiment, and it should be construed that the encoding units divided and determined in various forms are used in various methods of being processed independently in a predetermined order.
[0483] FIG. 15 illustrates, according to one embodiment, the process in which the video decoding apparatus 100 determines that when an encoding unit cannot be processed in a predetermined order, the current encoding unit is divided into an odd number of encoding units.
[0484] According to one embodiment, the video decoding apparatus 100 can determine that the current coding unit is divided into an odd number of coding units based on the acquired block form information and division form information. Referring to FIG. 15, a square-shaped first coding unit 1500 is divided into non-square-shaped second coding units 1510a and 1510b, and the second coding units 1510a and 1510b are independent of each other and may also be divided into third coding units 1520a, 1520b, 1520c, 1520d, and 1520e. According to one embodiment, the video decoding apparatus 100 can horizontally divide the left coding unit 1510a among the second coding units to determine a plurality of third coding units 1520a and 1520b, and the right coding unit 1510b can be divided into an odd number of third coding units 1520c, 1520d, and 1520e.
[0485] According to an embodiment, the video decoding apparatus 100 can determine whether the third encoding units 1520a, 1520b, 1520c, 1520d, 1520e are processed in a predetermined order, and can determine whether there are encoding units divided into an odd number. Referring to FIG. 15, the video decoding apparatus 100 can recursively divide the first encoding unit 1500 and determine the third encoding units 1520a, 1520b, 1520c, 1520d, 1520e. The video decoding apparatus 100 can determine whether the first encoding unit 1500, the second encoding units 1510a, 1510b, or the third encoding units 1520a, 1520b, 1520c, 1520d, 1520e are divided into an odd number of encoding units among the forms in which they are divided based on at least one of the block form information and the division form information. For example, in the second encoding units 1510a, 1510b, the encoding unit located on the right side may also be divided into an odd number of third encoding units 1520c, 1520d, 1520e. The order in which a plurality of encoding units included in the first encoding unit 1500 are processed is also a predetermined order (for example, z-scan order (1530)), and the video decoding apparatus 100 can determine whether the third encoding units 1520c, 1520d, 1520e determined by dividing the right second encoding unit 1510b into an odd number satisfy the condition of being processed in the predetermined order.
[0486] According to an embodiment, the video decoding apparatus 100 can determine whether the third encoding units 1520a, 1520b, 1520c, 1520d, 1520e included in the first encoding unit 1500 satisfy the condition that they are processed in a predetermined order. The condition is related to whether at least one of the width and height of the second encoding units 1510a, 1510b is divided in half by the boundaries of the third encoding units 1520a, 1520b, 1520c, 1520d, 1520e. For example, the third encoding units 1520a, 1520b determined by dividing the height of the non-square-shaped left second encoding unit 1510a in half satisfy the condition. However, since the boundaries of the third encoding units 1520c, 1520d, 1520e determined by dividing the right second encoding unit 1510b into three encoding units cannot divide the width or height of the right second encoding unit 1510b in half, it is determined that the third encoding units 1520c, 1520d, 1520e cannot satisfy the condition. In such a case where the condition is not satisfied, the video decoding apparatus 100 determines that there is a disconnection in the scan order, and based on the determination result, it can be determined that the right second encoding unit 1510b is divided into an odd number of encoding units. According to an embodiment, when the video decoding apparatus 100 is divided into an odd number of encoding units, it can impose a predetermined restriction on the encoding unit at a predetermined position among the divided encoding units. Since such restriction details or predetermined positions have been described through various embodiments, detailed descriptions will be omitted.
[0487] FIG. 16 illustrates a process in which a video decoding apparatus 100 divides a first encoding unit 1600 and determines at least one encoding unit according to an embodiment. According to an embodiment, the video decoding apparatus 100 can divide the first encoding unit 1600 based on at least one of block form information and division form information acquired via an acquisition unit 110. The square first encoding unit 1600 can be divided into four square encoding units or a plurality of non-square encoding units. For example, referring to FIG. 16, if the block form information indicates that the first encoding unit 1600 is square and the division form information indicates that it is divided into non-square encoding units, the video decoding apparatus 100 divides the first encoding unit 1600 into a plurality of non-square encoding units. Specifically, if the division form information indicates that the first encoding unit 1600 is divided in the horizontal or vertical direction to determine an odd number of encoding units, the video decoding apparatus 100 can divide the square first encoding unit 1600 into the second encoding units 1610a, 1610b, 1610c determined by being divided vertically as odd-numbered encoding units, or the second encoding units 1620a, 1620b, 1620c determined by being divided horizontally.
[0488] According to an embodiment, the video decoding apparatus 100 can determine whether the second encoding units 1610a, 1610b, 1610c, 1620a, 1620b, and 1620c included in the first encoding unit 1600 satisfy the condition that they are processed in a predetermined order. The condition is related to whether at least one of the width and height of the first encoding unit 1600 is divided in half by the boundaries of the second encoding units 1610a, 1610b, 1610c, 1620a, 1620b, and 1620c. Referring to FIG. 16, since the boundaries of the second encoding units 1610a, 1610b, and 1610c determined by vertically dividing the square-shaped first encoding unit 1600 cannot divide the width of the first encoding unit 1600 in half, it is also determined that the first encoding unit 1600 cannot satisfy the condition of being processed in a predetermined order. Also, since the boundaries of the second encoding units 1620a, 1620b, and 1620c determined by horizontally dividing the square-shaped first encoding unit 1600 cannot divide the width of the first encoding unit 1600 in half, it is also determined that the first encoding unit 1600 cannot satisfy the condition of being processed in a predetermined order. When such a condition is not satisfied, the video decoding apparatus 100 determines that there is a break in the scan order, and based on the determination result, it can be determined that the first encoding unit 1600 is divided into an odd number of encoding units. According to an embodiment, when the video decoding apparatus 100 is divided into an odd number of encoding units, a predetermined restriction can be placed on the encoding unit at a predetermined position among the divided encoding units. Since such restriction details or predetermined positions have been described through various embodiments, detailed descriptions will be omitted.
[0489] According to an embodiment, the video decoding apparatus 100 can divide the first encoding unit and determine encoding units in various forms.
[0490] Referring to FIG. 16, the video decoding apparatus 100 can divide the square-shaped first encoding unit 1600, the non-square-shaped first encoding unit 1630, or 1650 into encoding units in various forms.
[0491] FIG. 17 illustrates, according to one embodiment, that when a non-square second encoded unit determined by dividing a first encoded unit 1700 of a video decoding apparatus 100 satisfies a predetermined condition, a form in which the second encoded unit can be divided is restricted.
[0492] According to one embodiment, the video decoding apparatus 100 can determine to divide a square first encoded unit 1700 into non-square second encoded units 1710a, 1710b, 1720a, and 1720b based on at least one of block form information and division form information acquired via an acquisition unit 105. The second encoded units 1710a, 1710b, 1720a, and 1720b can also be independently divided. Thereby, the video decoding apparatus 100 can determine whether to divide into a plurality of encoded units or not based on at least one of block form information and division form information related to each of the second encoded units 1710a, 1710b, 1720a, and 1720b. According to one embodiment, the video decoding apparatus 100 can divide a non-square left second encoded unit 1710a determined by dividing the first encoded unit 1700 in the vertical direction in the horizontal direction to determine third encoded units 1712a and 1712b. However, when the video decoding apparatus 100 divides the left second encoded unit 1710a in the horizontal direction, it can be restricted such that the right second encoded unit 1710b is not divided in the horizontal direction in the same direction as the direction in which the left second encoded unit 1710a is divided. If the right second encoded unit 1710b is divided in the same direction and third encoded units 1714a and 1714b are determined, the third encoded units 1712a, 1712b, 1714a, and 1714b are also determined by independently dividing the left second encoded unit 1710a and the right second encoded unit 1710b in the horizontal direction. However, that is the same result as when the video decoding apparatus 100 divides the first encoded unit 1700 into four square second encoded units 1730a, 1730b, 1730c, and 1730d based on at least one of block form information and division form information, and that is inefficient in terms of video decoding.
[0493] According to one embodiment, the video decoding apparatus 100 can divide the non-square second encoding units 1720a or 1720b determined by dividing the first encoding unit 11300 in the horizontal direction in the vertical direction to determine the third encoding units 1722a, 1722b, 1724a, and 1724b. However, when the video decoding apparatus 100 divides one of the second encoding units (for example, the upper-end second encoding unit 1720a) in the vertical direction, for the above-described reasons, the other second encoding unit (for example, the lower-end encoding unit 1720b) can be restricted so as not to be divided in the vertical direction in the same direction as the direction in which the upper-end second encoding unit 1720a is divided.
[0494] FIG. 18 illustrates a process in which the video decoding apparatus 100 divides a square encoding unit when the division form information cannot indicate that it is divided into four square encoding units according to one embodiment.
[0495] According to one embodiment, the video decoding apparatus 100 can divide the first encoding unit 1800 based on at least one of the block form information and the division form information to determine the second encoding units 1810a, 1810b, 1820a, and 1820b. The division form information may include information related to various forms in which the encoding unit is divided. However, the information related to various forms may not include information for dividing into four square encoding units. According to such division form information, the video decoding apparatus 100 cannot divide the square first encoding unit 1800 into four square second encoding units 1830a, 1830b, 1830c, and 1830d. Based on the division form information, the video decoding apparatus 100 can determine non-square second encoding units 1810a, 1810b, 1820a, and 1820b.
[0496] According to one embodiment, the video decoding device 100 can independently divide non-square-shaped second encoding units 1810a, 1810b, 1820a, and 1820b respectively. Through a recursive method, each of the second encoding units 1810a, 1810b, 1820a, and 1820b is divided in a predetermined order, which is also a division method corresponding to the method by which the first encoding unit 1800 is divided based on at least one of the block form information and the division form information.
[0497] For example, the video decoding device 100 can horizontally divide the left second encoding unit 1810a to determine square-shaped third encoding units 1812a and 1812b, and can horizontally divide the right second encoding unit 1810b to determine square-shaped third encoding units 1814a and 1814b. Furthermore, the video decoding device 100 can horizontally divide both the left second encoding unit 1810a and the right second encoding unit 1810b to determine square-shaped third encoding units 1816a, 1816b, 1816c, and 1816d. In such a case, the encoding unit may also be determined in the same form as when the first encoding unit 1800 is divided into four square-shaped second encoding units 1830a, 1830b, 1830c, and 1830d.
[0498] For another example, the video decoding device 100 can vertically divide the upper second encoding unit 1820a to determine square-shaped third encoding units 1822a and 1822b, and can vertically divide the lower second encoding unit 1820b to determine square-shaped third encoding units 1824a and 1824b. Furthermore, the video decoding device 100 can vertically divide both the upper second encoding unit 1820a and the lower second encoding unit 1820b to determine square-shaped third encoding units 1822a, 1822b, 1824a, and 1824b. In such a case, the encoding unit may also be determined in the same form as when the first encoding unit 1800 is divided into four square-shaped second encoding units 1830a, 1830b, 1830c, and 1830d.
[0499] FIG. 19 illustrates that, according to one embodiment, the processing order among a plurality of encoding units may be different depending on the splitting process of the encoding units.
[0500] According to one embodiment, the video decoding apparatus 100 can split the first encoding unit 1900 based on the block form information and the splitting form information. When the block form information indicates a square shape and the splitting form information indicates that the first encoding unit 1900 is split in at least one of the horizontal direction and the vertical direction, the video decoding apparatus 100 can split the first encoding unit 1900 and determine, for example, the second encoding units 1910a, 1910b, 1920a, 1920b, 1930a, 1930b, 1930c, 1930d. Referring to FIG. 19, the non-square second encoding units 1910a, 1910b, 1920a, 1920b determined by splitting the first encoding unit 1900 only in the horizontal direction or the vertical direction are also independently split based on the relevant block form information and splitting form information. For example, the video decoding apparatus 100 can split the second encoding units 1910a, 1910b generated by splitting the first encoding unit 1900 in the vertical direction in the horizontal direction respectively to determine the third encoding units 1916a, 1916b, 1916c, 1916d, and can split the second encoding units 1920a, 1920b generated by splitting the first encoding unit 1900 in the horizontal direction in the horizontal direction respectively to determine the third encoding units 1926a, 1926b, 1926c, 1926d. Since the splitting process of such second encoding units 1910a, 1910b, 1920a, 1920b has been described in connection with FIG. 17, detailed description will be omitted.
[0501] According to an embodiment, the video decoding device 100 can process encoding units in a predetermined order. The features related to the processing of encoding units in the predetermined order have been described with reference to FIG. 14, so the detailed description will be omitted. Referring to FIG. 19, the video decoding device 100 can divide a square-shaped first encoding unit 1900 and determine four square-shaped third encoding units 1916a, 1916b, 1916c, 1916d, 1926a, 1926b, 1926c, 1926d. According to an embodiment, the video decoding device 100 can determine the processing order of the third encoding units 1916a, 1916b, 1916c, 1916d, 1926a, 1926b, 1926c, 1926d according to the form in which the first encoding unit 1900 is divided.
[0502] According to an embodiment, the video decoding device 100 can divide the second encoding units 1910a, 1910b generated by being divided in the vertical direction in the horizontal direction respectively to determine the third encoding units 1916a, 1916b, 1916c, 1916d. The video decoding device 100 can process the third encoding units 1916a, 1916b, 1916c, 1916d in the order (1917) of first processing the third encoding units 1916a, 1916b included in the left second encoding unit 1910a in the vertical direction and then processing the third encoding units 1916c, 1916d included in the right second encoding unit 1910b in the vertical direction.
[0503] According to an embodiment, the video decoding device 100 can divide the second encoding units 1920a, 1920b generated by being divided in the horizontal direction in the vertical direction respectively to determine the third encoding units 1926a, 1926b, 1926c, 1926d. The video decoding device 100 can process the third encoding units 1926a, 1926b, 1926c, 1926d in the order (1927) of first processing the third encoding units 1926a, 1926b included in the upper second encoding unit 1920a in the horizontal direction and then processing the third encoding units 1926c, 1926d included in the lower second encoding unit 1920b in the horizontal direction.
[0504] Referring to FIG. 19, the second encoding units 1910a, 1910b, 1920a, 1920b are each divided, and square-shaped third encoding units 1916a, 1916b, 1916c, 1916d, 1926a, 1926b, 1926c, 1926d may also be determined. The second encoding units 1910a, 1910b determined by being divided in the vertical direction and the second encoding units 1920a, 1920b determined by being divided in the horizontal direction are divided into different forms, but according to the third encoding units 1916a, 1916b, 1916c, 1916d, 1926a, 1926b, 1926c, 1926d determined later, ultimately, in the encoding units of the same form, it results from the division of the first encoding unit 1900. Thereby, the video decoding apparatus 100 recursively divides the encoding units through different processes based on at least one of the block form information and the division form information, and as a result, even if encoding units of the same form are determined, a plurality of encoding units determined to be of the same form can be processed in different orders.
[0505] FIG. 20 illustrates, according to an embodiment, the process of determining the depth of an encoding unit when the encoding unit is recursively divided and a plurality of encoding units are determined, due to the change in the form and size of the encoding unit.
[0506] According to an embodiment, the video decoding apparatus 100 can determine the depth of the encoding unit according to a predetermined criterion. For example, the predetermined criterion can also be the length of the long side of the encoding unit. When the length of the long side of the current encoding unit is divided into 2n (n>0) times the length of the long side of the encoding unit before division, the video decoding apparatus 100 can determine that the depth of the current encoding unit is increased by n compared to the depth of the encoding unit before division. Hereinafter, the encoding unit with an increased depth will be expressed as an encoding unit with a lower depth.
[0507] Referring to FIG. 20, according to one embodiment, based on block form information indicating a square shape (for example, the block form information can indicate "0:SQUARE"), the video decoder 100 can divide the first encoded unit 2000 having a square shape and determine second encoded units 2002, third encoded units 2004, etc. with a lower depth. If the size of the square first encoded unit 2000 is 2Nx2N, then the width and height of the first encoded unit 2000 are divided by 1 / 2 1 The second encoded unit 2002 determined by doubling can have a size of NxN. Furthermore, the third encoded unit 2004 determined by dividing the width and height of the second encoded unit 2002 by 1 / 2 size can have a size of N / 2xN / 2. In that case, the width and height of the third encoded unit 2004 correspond to 1 / 2 2 times that of the first encoded unit 2000. When the depth of the first encoded unit 2000 is D, the depth of the second encoded unit 2002 which is 1 / 2 1 times the width and height of the first encoded unit 2000 is also D + 1, and the depth of the third encoded unit 2004 which is 1 / 2 2 times the width and height of the first encoded unit 2000 is also D + 2.
[0508] According to one embodiment, based on block form information indicating a non-square shape (for example, the block form information can indicate "1:NS_VER" indicating a non-square with a height longer than the width, or "2:NS_HOR" indicating a non-square with a width longer than the height), the video decoder 100 can divide the non-square first encoded unit 2010 or 2020 and determine second encoded units 2012 or 2022, third encoded units 2014 or 2024 with a lower depth.
[0509] The video decoder 100 can divide at least one of the width and height of the first encoding unit 2010 with an Nx2N size, and for example, can determine the second encoding units 2002, 2012, 2022. That is, the video decoder 100 can divide the first encoding unit 2010 horizontally to determine the second encoding unit 2002 with an NxN size or the second encoding unit 2022 with an NxN / 2 size, or can divide it both horizontally and vertically to determine the second encoding unit 2012 with an N / 2xN size.
[0510] According to one embodiment, the video decoder 100 can also divide at least one of the width and height of the first encoding unit 2020 with a 2NxN size, and for example, can determine the second encoding units 2002, 2012, 2022. That is, the video decoder 100 can divide the first encoding unit 2020 vertically to determine the second encoding unit 2002 with an NxN size or the second encoding unit 2012 with an N / 2xN size, or can divide it both horizontally and vertically to determine the second encoding unit 2022 with an NxN / 2 size.
[0511] According to one embodiment, the video decoder 100 can also divide at least one of the width and height of the second encoding unit 2002 with an NxN size, and for example, can determine the third encoding units 2004, 2014, 2024. That is, the video decoder 100 can divide the second encoding unit 2002 both vertically and horizontally to determine the third encoding unit 2004 with an N / 2xN / 2 size, 2 or determine the third encoding unit 2014 with an N / 2 2 xN / 2 size, or determine the third encoding unit 2024 with an N / 2xN / 2
[0512] According to one embodiment, the video decoding apparatus 100 can also divide at least one of the width and height of the second encoding unit 2012 of size N / 2xN, and for example, determine the third encoding units 2004, 2014, 2024. That is, the video decoding apparatus 100 can divide the second encoding unit 2012 horizontally to determine the third encoding unit 2004 of size N / 2xN / 2, or the third encoding unit 2024 of size N / 2xN / 2 2 or divide it both vertically and horizontally to determine the third encoding unit 2014 of size N / 2 2 xN / 2.
[0513] According to one embodiment, the video decoding apparatus 100 can also divide at least one of the width and height of the second encoding unit 2014 of size NxN / 2, and for example, determine the third encoding units 2004, 2014, 2024. That is, the video decoding apparatus 100 can divide the second ...
Claims
1. A video decoding method, comprising: obtaining information related to a first motion vector and information related to a second motion vector from a bitstream; obtaining a first extended reference block in a first reference picture using the first motion vector, and obtaining a second extended reference block in a second reference picture using the second motion vector, wherein the first extended reference block includes a first reference block and a first part extended from the first reference block, and the second extended reference block includes a second reference block and a second part extended from the second reference block; determining a displacement vector of a pixel group including at least one pixel adjacent to the inside of the boundary of a current block using gradient values of at least one reference pixel in the first extended reference block and gradient values of at least one reference pixel in the second extended reference block; wherein the first part of the first extended reference block is used to calculate gradient values of at least one reference pixel in the first reference block, and the second part of the second extended reference block is used to calculate gradient values of at least one reference pixel in the second reference block; obtaining predicted pixel values of the current block by performing optical flow-based compensation on the current block using the gradient values of the at least one reference pixel in the first reference block, the gradient values of the at least one reference pixel in the second reference block, and the displacement vector of the pixel group; restoring the current block based on the predicted pixel values; A video decoding method comprising the above steps.
2. A video encoding method, comprising: determining a first motion vector and a second motion vector; obtaining predicted pixel values of a current block by performing optical flow-based compensation on the current block using gradient values of at least one reference pixel in a first reference block, gradient values of at least one reference pixel in a second reference block, and a displacement vector of a pixel group of the current block, wherein the pixel group includes at least one pixel adjacent to the inside of the boundary of the current block; Based on the predicted pixel value, generating a bitstream including information related to the first motion vector and information related to the second motion vector as a result of encoding the current block; Using the first motion vector, a first extended reference block is obtained within a first reference picture; Using the second motion vector, a second extended reference block is obtained within a second reference picture; The first extended reference block includes the first reference block and a first part extended from the first reference block, and the second extended reference block includes the second reference block and a second part extended from the second reference block; The displacement vector of the pixel group is determined using gradient values of at least one reference pixel within the first extended reference block and gradient values of at least one reference pixel within the second extended reference block; The first part of the first extended reference block is used to calculate the gradient value of the at least one reference pixel within the first reference block; The second part of the second extended reference block is used to calculate the gradient value of the at least one reference pixel within the second reference block; Video encoding method.
3. A method of transmitting the bitstream generated by the video encoding method according to Claim 2.
Citation Information
Patent Citations
Motion compensation method and device for encoding and decoding scalable video
US20150350671A1
Cited By
Matrix-based intra prediction using upsampling
US12519945B2
Context coding for matrix-based intra prediction
US12563225B2
Syntax signaling and parsing based on colour component
US12568230B2
Matrix-based intra prediction using filtering
US12621487B2