Video signal encoding / decoding method and apparatus therefor
The proposed video signal encoding/decoding method employs an affine model for inter prediction and derives affine seed vectors from sub-block motion vectors, improving prediction and encoding efficiency and overcoming the limitations of current video compression standards.
Patent Information
- Application Number
- JP2024029968
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-09-21
- Filing Date
- 2024-02-29
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2039-09-20
AI Technical Summary
The increasing demand for high-definition video services has led to a need for improved video compression ratios, as existing standards like HEVC are reaching performance limitations.
A video signal encoding/decoding method that uses an affine model for inter prediction, deriving affine seed vectors from translational motion vectors of sub-blocks, and converting distances between adjacent blocks into power series of 2.
This method enhances prediction efficiency and encoding efficiency by effectively utilizing affine motion models and power series conversions, addressing the limitations of existing compression standards.
Smart Images

Figure 0007693872000047 
Figure 0007693872000048 
Figure 0007693872000049
Abstract
Description
Technical Field
[0001] The present invention relates to a video signal encoding / decoding method and an apparatus therefor.
Background Art
[0002] As display panels are becoming larger and larger, there is a need for higher-quality video services. The biggest problem with high-definition video services is the significant increase in data volume. To solve such problems, studies on improving the video compression ratio are being actively carried out. As a typical example, in 2009, the Motion Picture Experts Group (MPEG) and the Video Coding Experts Group (VCEG) under the umbrella of the International Telecommunication Union - Telecommunication (ITU-T) established the Joint Collaborative Team on Video Coding (JCT-VC). JCT-VC submitted the video compression standard HEVC (High Efficiency Video Coding), and the standard was approved on January 25, 2013. Its compression performance is about twice that of H.264 / AVC. With the rapid growth of high-definition video services, the performance limitations of HEVC are gradually emerging.
Summary of the Invention
Problems to be Solved by the Invention
[0003] An object of the present invention is to provide an inter prediction method using an affine model when encoding / decoding a video signal and an apparatus used for the inter prediction method.
[0004] An object of the present invention is to provide a method for deriving an affine seed vector using a translational motion vector of a sub-block when encoding / decoding a video signal, and an apparatus for executing the method.
[0005] Another object of the present invention is to provide a method for deriving an affine seed vector by converting the distance between an adjacent block and a current block into a power series of 2 when encoding / decoding a video signal, and an apparatus for executing the method.
[0006] The technical problems to be solved by the present invention are not limited to the above-mentioned technical problems, and those having ordinary knowledge in the technical field to which the present invention belongs can clearly understand other technical problems not mentioned from the following description.
Means for Solving the Problems
[0007] A video signal decoding / encoding method according to the present invention includes: generating a merge candidate list of a current block; designating any one of a plurality of merge candidates included in the merge candidate list; deriving a first affine seed vector and a second affine seed vector of the current block based on the first affine seed vector and the second affine seed vector of the designated merge candidate; deriving an affine vector of a sub-block within the current block by using the first affine seed vector and the second affine seed vector of the current block; and performing motion compensation prediction on the sub-block based on the affine vector. In this case, the sub-block is a region having a size smaller than the size of the current block. Note that the first affine seed vector and the second affine seed vector of the merge candidate can be derived based on motion information of an adjacent block adjacent to the current block.
[0008] In the video signal decoding / encoding method according to the present invention, when the adjacent block is included in an encoding tree unit different from the encoding tree unit of the current block, based on the motion vectors of the lower left sub-block and the lower right sub-block of the adjacent block, the first affine seed vector and the second affine seed vector of the merge candidate can be derived.
[0009] In the video signal decoding / encoding method according to the present invention, the lower left sub-block may include a lower left reference sample located at the lower left corner of the adjacent block, and the lower right sub-block may include a lower right reference sample located at the lower right corner of the adjacent block.
[0010] In the video signal decoding / encoding method according to the present invention, based on a value obtained by performing a shift operation on a difference value between the motion vectors of the lower left sub-block and the lower right sub-block using a scale factor, the first affine seed vector and the second affine seed vector of the merge candidate can be derived, and the scale factor can be derived based on a value obtained by adding a horizontal distance and an offset between the lower left reference sample and the lower right reference sample.
[0011] In the video signal decoding / encoding method according to the present invention, based on a value obtained by performing a shift operation on a difference value between the motion vectors of the lower left sub-block and the lower right sub-block using a scale factor, the first affine seed vector and the second affine seed vector of the merge candidate can be derived, and the scale factor can be derived based on a distance between an adjacent sample adjacent to the right side of the lower right reference sample and the lower left reference sample.
[0012] In the video signal decoding / encoding method according to the present invention, the merge candidate list includes a first merge candidate and a second merge candidate. The first merge candidate is derived based on a first available block among upper adjacent blocks located above the current block and the determined upper adjacent block. The second merge candidate is derived based on a first available block among left adjacent blocks located to the left of the current block and the determined left adjacent block.
[0013] In the video signal decoding / encoding method according to the present invention, when the adjacent block is included in an encoding tree unit that is the same as the encoding tree unit of the current block, the first affine seed vector and the second affine seed vector of the merge candidate can be derived based on the first affine seed vector and the second affine seed vector of the adjacent block.
[0014] The features briefly summarized above about the present invention are only exemplary embodiments of the detailed description of the present invention to be described later, and do not limit the scope of the present invention.
Effects of the Invention
[0015] According to the present invention, there is an effect of improving prediction efficiency by using an inter prediction method using an affine model.
[0016] According to the present invention, there is an effect of improving encoding efficiency by deriving an affine seed vector using the translational motion vector of a sub-block.
[0017] According to the present invention, there is an effect of improving encoding efficiency by deriving an affine seed vector by converting the distance between an adjacent block and the current block into a power of 2.
[0018] The effects achievable by the present invention are not limited to the above effects, and those having ordinary knowledge in the technical field to which the present invention pertains can clearly understand other effects not mentioned from the following description.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Embodiments for Carrying Out the Invention
[0020] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0021] Video encoding and decoding are performed in units of blocks. For example, encoding / decoding processes such as conversion, quantization, prediction, loop filtering, or reconstruction can be performed on an encoding block, a conversion block, or a prediction block.
[0022] Hereinafter, a block to be encoded / decoded will be referred to as a "current block". For example, according to the current encoding / decoding process step, the current block can represent an encoding block, a conversion block, or a prediction block.
[0023] Note that the term "unit" used in this specification represents a basic unit for executing a specific encoding / decoding process, and "block" may be understood to represent an array of samples of a predetermined size. Unless otherwise specified, "block" and "unit" are used interchangeably. For example, in the embodiments described later, it may be understood that an encoding block and an encoding unit have the same meaning.
[0024] FIG. 1 is a block diagram showing a video encoder according to an embodiment of the present invention.
[0025] Referring to FIG. 1, the video encoding device 100 may include an image segmentation unit 110, prediction units 120, 125, a conversion unit 130, a quantization unit 135, a rearrangement unit 160, an entropy encoding unit 165, an inverse quantization unit 140, an inverse conversion unit 145, a filter unit 150, and a memory 155.
[0026] Each member shown in FIG. 1 is shown separately and represents different characteristic functions in the video encoding device, but does not represent that each member is composed of separate hardware or a single software assembly. That is, for ease of explanation, it is shown as a representative component and includes each component, and at least two components may be combined into one component or one component may be divided into a plurality of components to perform functions thereby. Examples of integrating such components and examples of separating such components also belong to the scope of the present invention as long as the essence of the present invention is not deviated from.
[0027] Note that some components are not essential components for performing the essential functions in the present invention, but are only selectable components for improving performance. The present invention may be implemented by including only the members necessary for realizing the essence of the present invention (excluding the components for improving performance), and a structure including only the necessary components (excluding the components for improving performance) also belongs to the scope of the present invention.
[0028] The image segmentation unit 110 can divide the input image into at least one processing unit. In this case, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The image segmentation unit 110 divides one image into a combination of a plurality of coding units, prediction units, and transform units. Based on a predetermined criterion (for example, a cost function), a combination of a coding unit, a prediction unit, and a transform unit can be selected to encode the image.
[0029] For example, one image can be divided into a plurality of coding units. To divide an image into coding units, a recursive tree structure such as a quad tree structure can be used to divide one video or the largest coding unit as the root into other coding units. The coding unit may have the same number of child nodes as the number of divided coding units. Coding units not divided by some restrictions become leaf nodes. That is, assuming that one coding unit can only achieve square division, one coding unit can be divided into at most four other coding units.
[0030] Hereinafter, in the embodiments of the present invention, the coding unit may refer to a unit that performs coding or a unit that performs decoding.
[0031] At least one prediction unit in one coding unit can be divided into squares or rectangles with the same size, and one prediction unit in one coding unit can also be divided into one with a shape and / or size different from another prediction unit.
[0032] When the prediction unit that performs intra prediction based on the coding unit is not the minimum coding unit, it is not necessary to divide it into a plurality of prediction units N×N, and intra prediction can be performed.
[0033] The prediction units 120 and 125 may include an inter prediction unit 120 that performs inter prediction and an intra prediction unit 125 that performs intra prediction. It is possible to determine whether to use inter prediction or intra prediction for the prediction unit, and based on each prediction method, specific information (for example, intra prediction mode, motion vector, reference image, etc.) can be determined. In this case, the processing unit that executes the prediction may be different from the processing unit that determines the prediction method and the specific content. For example, the prediction unit can determine the prediction method, prediction mode, etc., and the conversion unit can execute the prediction. The residual value (residual block) between the generated prediction block and the original block can be input to the conversion unit 130. Note that prediction mode information, motion vector information, etc. for prediction can be encoded in the entropy encoding unit 165 together with the residual value and transmitted to the decoder. When using a specific encoding mode, it is also possible to directly encode the original block without generating a prediction block by the prediction units 120 and 125 and transmit it to the decoder.
[0034] The inter prediction unit 120 can predict the prediction unit based on the information of at least one image in the image immediately before or immediately after the current image. In some cases, the prediction unit can also be predicted based on the information of a partially encoded region in the current image. The inter prediction unit 120 may include a reference image interpolation unit, a motion prediction unit, and a motion compensation unit.
[0035] The reference image interpolation unit can receive reference image information from the memory 155 and generate pixel information of integer pixels or fractional pixels from the reference image. For luminance pixels, a DCT-based 8-tap interpolation filter with different filter coefficients can be used to generate pixel information of fractional pixels in units of 1 / 4 pixels. For chrominance signals, a DCT-based 4-tap interpolation filter with different filter coefficients can be used to generate pixel information of fractional pixels in units of 1 / 8 pixels.
[0036] The motion prediction unit can perform motion prediction based on the reference image interpolated by the reference image interpolation unit. As methods for calculating motion vectors, multiple methods such as the Full search-based Block Matching Algorithm (FBMA), the Three Step Search (TSS), and the New Three-Step Search Algorithm (NTS) can be used. According to the interpolated pixels, the motion vector may have a motion vector value in units of 1 / 2 pixels or 1 / 4 pixels. In the motion prediction unit, the current prediction unit can be predicted by using different motion prediction methods. As motion prediction methods, multiple methods such as the Skip method, the Merge method, the Advanced Motion Vector Prediction (AMVP) method, and the Intra Block Copy method can be used.
[0037] The intra prediction unit 125 can generate a prediction unit based on the reference pixel information around the current block (the reference pixel information is the pixel information in the current image). When the adjacent block of the current prediction unit is a block that has performed inter prediction and the reference pixel is a pixel that has performed inter prediction, the reference pixels included in the block that has performed inter prediction can be used as the reference pixel information of the block that has performed intra prediction on the periphery. That is, when the reference pixel is unavailable, at least one of the available reference pixels can be used instead of the unavailable reference pixel information.
[0038] In intra prediction, the prediction mode may have an angular prediction mode that uses reference pixel information based on the prediction direction and a non-angular mode that does not use direction information when performing prediction. The mode for predicting luminance information may be different from the mode for predicting chrominance information. To predict chrominance information, the intra prediction mode information for predicting luminance information or the predicted luminance signal information can be used.
[0039] When performing intra prediction, if the size of the prediction unit is the same as the size of the conversion unit, intra prediction can be performed on the prediction unit based on the pixels located on the left side, upper left side, and upper side of the prediction unit. However, when performing intra prediction, if the size of the prediction unit is different from the size of the conversion unit, intra prediction can be performed based on the reference pixels of the conversion unit. Note that intra prediction using N×N division can be applied only to the minimum coding unit.
[0040] After applying an Adaptive Intra Smoothing (AIS) filter to a reference pixel based on a prediction mode, an intra prediction method can generate a prediction block. The type of the adaptive intra smoothing filter applied to the reference pixel may be different. To execute the intra prediction method, the intra prediction mode of the current prediction unit can be predicted based on the intra prediction modes of the prediction units located around the current prediction unit. When predicting the prediction mode of the current prediction unit using the mode information predicted from the surrounding prediction units, if the intra prediction mode of the current prediction unit is the same as that of the surrounding prediction units, information indicating that the prediction mode of the current prediction unit is the same as that of the surrounding prediction units can be transmitted using predetermined flag information. If the prediction mode of the current prediction unit is different from that of the surrounding prediction units, the prediction mode information of the current block can be encoded by performing entropy encoding.
[0041] In addition, a residual block including residual information can be generated. The residual information is a difference value between an original block of a prediction unit that performs prediction based on the prediction units generated by the prediction units 120 and 125 and the prediction unit. The generated residual block can be input to the conversion unit 130.
[0042] The conversion unit 130 can convert the residual block by a conversion method such as a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), or conversion skip. The residual block includes residual information between the original block and the prediction units generated by the prediction units 120 and 125. Regarding whether to apply DCT, DST, or KLT to convert the residual block, it can be determined based on the intra prediction mode information of the prediction unit for generating the residual block.
[0043] The quantization unit 135 can quantize the values converted into the frequency domain by the conversion unit 130. The quantization coefficient may vary depending on the importance of the block or the image. The values calculated by the quantization unit 135 may be provided to the inverse quantization unit 140 and the rearrangement unit 160.
[0044] The rearrangement unit 160 can perform rearrangement of coefficient values on the quantized residual values.
[0045] The rearrangement unit 160 can change the two-dimensional block-shaped coefficients into a one-dimensional vector form by a coefficient scanning method. For example, the rearrangement unit 160 can scan the DC coefficient to the coefficients in the high-frequency region by a zig-zag scan method and convert it into a one-dimensional vector form. Depending on the size of the conversion unit and the intra prediction mode, instead of the zig-zag scan, a vertical scan that scans the two-dimensional block-shaped coefficients along the column direction and a horizontal scan that scans the two-dimensional block-shaped coefficients along the row direction can also be used. That is, based on the size of the conversion unit and the intra prediction mode, it is possible to determine which of the zig-zag scan, the vertical direction scan, and the horizontal direction scan to use.
[0046] The entropy coding unit 165 can perform entropy coding based on the values calculated by the rearrangement unit 160. For example, the entropy coding can use a plurality of coding methods such as exponential Golomb code, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC).
[0047] The entropy encoding unit 165 can encode a plurality of information such as the residual value coefficient information and block type information of the encoding units from the rearrangement unit 160 and the prediction units 120 and 125, prediction mode information, division unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, filtering information, etc.
[0048] The entropy encoding unit 165 can perform entropy encoding on the coefficient values of the encoding units input from the rearrangement unit 160.
[0049] The inverse quantization unit 140 and the inverse transformation unit 145 perform inverse quantization on the plurality of values quantized by the quantization unit 135, and perform inverse transformation on the values transformed by the transformation unit 130. By merging the residual values generated by the inverse quantization unit 140 and the inverse transformation unit 145 and the prediction units predicted by the motion prediction unit, motion compensation unit, and intra prediction unit included in the prediction units 120 and 125, a reconstructed block can be generated.
[0050] The filter unit 150 may include at least one of a deblocking filter, an offset correction unit, and an adaptive loop filter (ALF).
[0051] The deblocking filter can remove the block distortion generated in the reconstructed image due to the boundaries between blocks. To determine whether to perform deblocking, it is possible to determine whether to apply the deblocking filter to the current block based on the pixels included in several columns or rows contained in the block. When applying the deblocking filter to a block, a Strong Filter or a Weak Filter can be applied based on the required deblocking filter strength. In addition, in the process of using the deblocking filter, when performing vertical filtering or horizontal filtering, horizontal direction filtering and vertical direction filtering can be executed synchronously.
[0052] The offset correction unit can correct the offset between the deblocking image and the original image on a pixel-by-pixel basis. Offset correction can be performed on the specified image in the following manner. After dividing the pixels included in the image into a predetermined number of regions, the regions that require offset correction are determined, and offset correction is applied to the corresponding regions or offset correction is applied considering the edge information of each pixel.
[0053] Adaptive Loop Filtering (ALF) can be executed based on the value obtained by comparing the filtered reconstructed image and the original image. After dividing the pixels included in the image into predetermined groups, one filter used for the corresponding group is determined, and filtering can be performed differentially for each group. For each Coding Unit (CU), information regarding whether to apply adaptive loop filtering can be transmitted via the luminance signal. The shape and filter coefficients of the applied adaptive loop filter may vary depending on each block. Note that regardless of the characteristics of the applied block, the same type (a certain type) of ALF can also be applied.
[0054] The memory 155 can store the reconstructed block or blocks calculated by the filter unit 150, and can provide the stored reconstructed block or image to the prediction units 120 and 125 when performing inter prediction.
[0055] FIG. 2 is a block diagram showing a video decoder according to an embodiment of the present invention.
[0056] Referring to FIG. 2, the video decoder 200 may include an entropy decoding unit 210, a rearrangement unit 215, an inverse quantization unit 220, an inverse transform unit 225, a prediction unit 230, a prediction unit 235, a filter unit 240, and a memory 245.
[0057] When inputting a video bitstream from a video encoder, in a step opposite to that of the video encoder, the input bitstream can be decoded.
[0058] The entropy decoding unit 210 can perform entropy decoding in a step opposite to the step of performing entropy encoding by the entropy encoding unit of the video encoder. For example, a plurality of methods such as Exponential Golomb code, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) can be applied so as to correspond to the method executed by the video encoder.
[0059] The entropy decoding unit 210 can decode information related to intra prediction and inter prediction executed by the encoder.
[0060] The rearrangement unit 215 can perform rearrangement by a method of rearranging the bit stream entropy-decoded by the entropy decoding unit 210 in the encoding unit. Rearrangement can be performed by reconstructing a plurality of coefficients represented in a one-dimensional vector form into coefficients in a two-dimensional block shape. The rearrangement unit 215 can perform rearrangement in the following manner. It receives information related to the coefficient scanning executed by the encoding unit and performs reverse scanning according to the scanning order executed by the corresponding encoding unit.
[0061] The inverse quantization unit 220 can perform inverse quantization based on the quantization parameter provided by the encoder and the coefficient values of the rearranged blocks.
[0062] Regarding the quantization result executed by the video encoder, the inverse transform unit 225 can perform inverse transforms for the DCT, DST, and KLT, which are the transforms executed by the transform unit. That is, it performs inverse DCT, inverse DST, and inverse KLT. The inverse transform may be executed by the transmission unit determined in the video encoder. In the inverse transform unit 225 of the video decoder, the transform method (for example, DCT, DST, KLT) can be selectively executed based on a plurality of information such as the prediction method, the size of the current block, and the prediction direction.
[0063] The prediction units 230, 235 can generate a prediction block based on the information related to the generation of the prediction block provided by the entropy decoding unit 210 and the previously decoded block or image information provided by the memory 245.
[0064] As described above, when performing intra prediction in the same manner as the operation method in the video encoder, if the size of the prediction unit is the same as the size of the conversion unit, intra prediction can be performed on the prediction unit based on the pixels located on the left side, upper left side, and upper side of the prediction unit. When performing intra prediction, if the size of the prediction unit is different from the size of the conversion unit, intra prediction can be performed based on the reference pixels of the conversion unit. Note that intra prediction using N×N partitioning can be applied only to the minimum coding unit.
[0065] The prediction units 230 and 235 may include a prediction unit determination unit, an inter prediction unit, and an intra prediction unit. The prediction unit determination unit receives a plurality of pieces of information such as prediction unit information input from the entropy decoding unit 210, prediction mode information of the intra prediction method, and motion prediction related information of the inter prediction method, and classifies the prediction unit based on the current coding unit, and determines whether the prediction unit is performing inter prediction or intra prediction. The inter prediction unit 230 can perform inter prediction on the current prediction unit based on information included in at least one of the immediately preceding image or the immediately following image of the current image to which the current prediction unit belongs, using the information necessary for performing inter prediction of the current prediction unit provided by the video encoder. Alternatively, inter prediction can also be performed based on information of a reconstructed partial region in the current image to which the current prediction unit belongs.
[0066] To perform inter prediction, based on the coding unit, it can be determined which of the skip mode, merge mode, advanced motion vector prediction mode (AMVP mode), and intra block copy mode is the motion prediction method of the prediction unit included in the corresponding coding unit.
[0067] The intra prediction unit 235 can generate a prediction block based on the pixel information in the current image. When the prediction unit is a prediction unit that has performed intra prediction, intra prediction can be performed based on the intra prediction mode information of the prediction unit provided from the video encoder. The intra prediction unit 235 may include an Adaptive Intra Smoothing (AIS) filter, a reference pixel interpolation unit, and a DC filter. The adaptive intra smoothing filter is a part that performs filtering on the reference pixels of the current block, and can also determine whether to apply the filter based on the prediction mode of the current prediction unit. Using the prediction mode of the prediction unit and the adaptive intra smoothing filter information provided from the video encoder, adaptive intra smoothing filtering can be performed on the reference pixels of the current block. If the prediction mode of the current block is a mode that does not perform adaptive intra smoothing filtering, the adaptive intra smoothing filter may not be applied.
[0068] When the prediction mode of the prediction unit is a prediction unit that performs intra prediction based on the pixel values for interpolating the reference pixels, the reference pixel interpolation unit can generate reference pixels in pixel units of integer values or fractional values by interpolating the reference pixels. If the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixels, interpolation of the reference pixels may not be performed. When the prediction mode of the current block is the DC mode, the DC filter can generate a prediction block by filtering.
[0069] The reconstructed block or image can be provided to the filter unit 240. The filter unit 240 may include a deblocking filter, an offset correction unit, and an ALF.
[0070] It is possible to receive information regarding whether to apply a deblocking filter to a corresponding block or image from a video encoder, and information regarding whether to use a strong filter or a weak filter when applying the deblocking filter. The deblocking filter of the video decoder can receive information regarding the deblocking filter provided from the video encoder, and the video decoder can perform deblocking filtering on the corresponding block.
[0071] The offset correction unit can perform offset correction on the reconstructed image based on the type of offset correction applied to the image and the offset information when performing encoding.
[0072] Based on information regarding whether to apply ALF provided from the encoder, ALF coefficient information, etc., ALF can be applied to the encoding unit. Such ALF information may be provided by being included in a specific parameter set.
[0073] The memory 245 stores the reconstructed image or block, enables the image or block to be used as a reference image or reference block, and can provide the reconstructed image to the output unit.
[0074] FIG. 3 is a diagram showing a basic encoding tree unit according to an embodiment of the present invention.
[0075] The coding block of the maximum size can be defined as a coding tree block. One image may be divided into a plurality of coding tree units (CTUs). The coding tree unit is the coding unit of the maximum size and may be called the largest coding unit (LCU). FIG. 3 shows an example of dividing one image into a plurality of coding tree units.
[0076] The size of the coded tree unit may be defined at the picture level or the sequence level. Therefore, information indicating the size of the coded tree unit can be transmitted using a signal by a set of picture parameters or a set of sequence parameters.
[0077] For example, the size of the coded tree unit for the entire picture in a sequence can be set to 128×128. Alternatively, one of 128×128 or 256×256 at the picture level can be determined as the size of the coded tree unit. For example, the size of the coded tree unit in the first picture can be set to 128×128, and the size of the coded tree unit in the second picture can be set to 256×256.
[0078] Coded blocks can be generated by dividing the coded tree unit. The coded block represents a basic unit for performing encoding / decoding processing. For example, prediction or transformation can be executed according to different coded blocks, or a predictive coding mode can be determined according to different coded blocks. Here, the predictive coding mode represents a method for generating a predicted picture. For example, the predictive coding mode may include intra prediction, inter prediction, current picture referencing (CPR), intra block copy (IBC), or combined prediction. For a coded block, a predicted block related to the coded block can be generated using at least one of the predictive coding modes of intra prediction, inter prediction, current picture referencing, or combined prediction.
[0079] Through a bitstream, information representing the prediction coding mode of the current block can be transmitted in a signal. For example, the information may be a 1-bit flag indicating whether the prediction coding mode is an intra mode or an inter mode. Only when it is determined that the prediction coding mode of the current block is an inter mode, current picture reference or combined prediction can be used.
[0080] Current picture reference uses the current picture as a reference picture and is used to obtain a predicted block of the current block from the encoded / decoded region in the current picture. Here, the current picture means the picture including the current block. Through a bitstream, information representing whether to apply current picture reference to the current block can be transmitted in a signal. For example, the information may be a 1-bit flag. When the flag is true, the prediction coding mode of the current block can be determined as current picture reference. When the flag is false, the prediction mode of the current block can be determined as inter prediction.
[0081] Alternatively, based on a reference picture index, the prediction coding mode of the current block can be determined. For example, when the reference picture index indicates the current picture, the prediction coding mode of the current block can be determined as current picture reference. When the reference picture index indicates another picture that is not the current picture, the prediction coding mode of the current block can be determined as inter prediction. That is, current picture reference is a prediction method that uses information of the encoded / decoded region in the current picture, and inter prediction is a prediction method that uses information of other encoded / decoded pictures.
[0082] Combined prediction is an encoding mode formed by combining two or more of intra prediction, inter prediction, and current picture reference. For example, when applying combined prediction, a first prediction block can be generated based on one of intra prediction, inter prediction, or current picture reference, and a second prediction block can be generated based on another one. When generating the first prediction block and the second prediction block, the final prediction block can be generated by performing an average operation and a weighted addition on the first prediction block and the second prediction block. Information indicating whether to apply combined prediction can be transmitted as a signal through the bitstream. The information may be a 1-bit flag.
[0083] FIG. 4 is a diagram showing a plurality of division types of an encoding block.
[0084] Based on quadtree division, binary tree division, or ternary tree division, an encoding block can be divided into a plurality of encoding blocks. Also, based on quadtree division, binary tree division, or ternary tree division, the divided encoding blocks can be further divided into a plurality of encoding blocks.
[0085] Quadtree division is a division technique that divides a current block into four blocks. As a result of quadtree division, the current block can be divided into four square partitions (refer to "SPLIT_QT" in FIG. 4(a)).
[0086] Binary tree splitting is a splitting technique that divides the current block into two blocks. The process of dividing the current block into two blocks along the vertical direction (i.e., using a vertical line that crosses the current block) is called vertical binary tree splitting, and the process of dividing the current block into two blocks along the horizontal direction (i.e., using a horizontal line that crosses the current block) can be called horizontal binary tree splitting. As a result of binary tree splitting, the current block can be divided into two non-square partitions. "SPLIT_BT_VER" in Fig. 4(b) represents the vertical binary tree splitting result, and "SPLIT_BT_HOR" in Fig. 4(c) represents the horizontal binary tree splitting result.
[0087] Ternary tree splitting is a splitting technique that divides the current block into three blocks. The process of dividing the current block into three blocks along the vertical direction (i.e., using two vertical lines that cross the current block) is called vertical ternary tree splitting, and the process of dividing the current block into three blocks along the horizontal direction (i.e., using two horizontal lines that cross the current block) can be called horizontal ternary tree splitting. As a result of ternary tree splitting, the current block can be divided into three non-square partitions. In this case, the width / height of the partition located at the center of the current block may be twice the width / height of the other partitions. "SPLIT_TT_VER" in Fig. 4(d) represents the vertical ternary tree splitting result, and "SPLIT_TT_HOR" in Fig. 4(e) represents the horizontal ternary tree splitting result.
[0088] The number of splitting times of the coding tree unit can be defined as the partitioning depth. At the sequence or image level, the maximum partitioning depth of the coding tree unit can be determined. Therefore, the maximum partitioning depth of the coding tree unit may vary depending on the sequence or image.
[0089] Alternatively, the maximum split depth can be determined individually for each of the plurality of splitting techniques. For example, the maximum split depth that allows quadtree splitting may be different from the maximum split depth that allows binary tree splitting and / or ternary tree splitting.
[0090] The encoder can transmit, via a bitstream, information representing at least one of the split type or split depth of the current block in the signal. The decoder can determine the split type and split depth of the coding tree unit based on the information parsed from the bitstream.
[0091] FIG. 5 is a diagram showing an example of splitting of a coding tree unit.
[0092] The process of splitting an encoded block using a splitting technique such as quadtree splitting, binary tree splitting, and / or ternary tree splitting can be called multi-tree partitioning.
[0093] The encoded blocks generated by applying multi-tree partitioning to an encoded block can be called downstream encoded blocks. When the split depth of the encoded block is k, the split depth of the plurality of downstream encoded blocks is set to k + 1.
[0094] On the other hand, for an encoded block with a split depth of k + 1, the encoded block with a split depth of k can be called an upstream encoded block.
[0095] The split type of the current encoded block can be determined based on at least one of the split type of the upstream encoded block or the split type of the adjacent encoded block. Here, the adjacent encoded block is adjacent to the current encoded block, and it may include at least one of the upper adjacent block, left adjacent block, or adjacent block adjacent to the upper left corner of the current encoded block. Here, the split type may include at least one of whether to perform quadtree splitting, whether to perform binary tree splitting, binary tree splitting direction, whether to perform ternary tree splitting, or ternary tree splitting direction.
[0096] To determine the splitting type of the coded block, information indicating whether the coded block is split can be transmitted via a bitstream in a signal. The information is a 1-bit flag "split_cu_flag", and when the flag is true, it indicates that the coded block is split using the quadtree splitting technique.
[0097] When "split_cu_flag" is true, information indicating whether the coded block is split into a quadtree can be transmitted via a bitstream in a signal. The information is a 1-bit flag "split_qt_flag", and when the flag is true, the coded block may be split into four blocks.
[0098] For example, in the example shown in FIG. 5, it is illustrated that four coded blocks with a splitting depth of 1 are generated when the coding tree unit is split into a quadtree. It is also illustrated that the quadtree splitting is applied again to the first and fourth coded blocks among the four coded blocks generated as the quadtree splitting result. Finally, four coded blocks with a splitting depth of 2 can be generated.
[0099] Note that by applying the quadtree splitting again to the coded block with a splitting depth of 2, a coded block with a splitting depth of 3 can be generated.
[0100] If the quadtree splitting is not applied to the coded block, it is possible to determine whether to perform binary tree splitting or ternary tree splitting on the coded block by considering at least one of the size of the coded block, whether the coded block is located at the boundary of the image, the maximum splitting depth, or the splitting type of the adjacent blocks. When it is determined to perform binary tree splitting or ternary tree splitting on the coded block, information representing the splitting direction can be transmitted as a signal via the bitstream. The information may be a 1-bit flag "mtt_split_cu_vertical_flag". Based on the flag, it is possible to determine whether the splitting direction is vertical or horizontal. Note that information representing which of binary tree splitting or ternary tree splitting is applied to the coded block can be transmitted as a signal via the bitstream. The information may be a 1-bit flag "mtt_split_cu_binary_flag". Based on the flag, it is possible to determine whether to apply binary tree splitting or ternary tree splitting to the coded block.
[0101] For example, in the example shown in FIG. 5, it is illustrated that vertical binary tree splitting is applied to a coded block with a splitting depth of 1. Vertical ternary tree splitting is applied to the left coded block in the coded block generated as the splitting result, and vertical binary tree splitting is applied to the right coded block.
[0102] Inter prediction is a predictive coding mode that predicts the current block using the information of the previous image. For example, a block at the same position as the current block in the previous image (hereinafter referred to as a collocated block) can be used as the prediction block of the current block. Hereinafter, a prediction block generated based on a block having the same position as the current block can be referred to as a collocated prediction block.
[0103] On the other hand, when the object existing in the previous image moves to another position in the current image, the current block can be effectively predicted by symmetrical movement. For example, if the movement direction and size of the object can be understood by comparing the previous image and the current image, the predicted block (or predicted image) of the current block can be generated by considering the movement information of the object. Hereinafter, the predicted block generated by using the movement information can be called a motion prediction block.
[0104] A residual block can be generated by subtracting the predicted block from the current block. In this case, when there is movement of the object, by using a collocated prediction block instead of the motion prediction block, the energy of the residual block can be reduced, and the compression performance of the high-residual block can be improved.
[0105] As described above, the process of generating a predicted block by using movement information can be called motion compensation prediction. In most inter-predictions, a predicted block can be generated based on motion compensation prediction.
[0106] The movement information may include at least one of a motion vector, a reference image index, a prediction direction, or a bidirectional weight value index. The motion vector represents the movement direction and size of the object. The reference image index designates the reference image of the current block among a plurality of reference images included in the reference image list. The prediction direction refers to any one of unidirectional L0 prediction, unidirectional L1 prediction, or bidirectional prediction (L0 prediction and L1 prediction). Based on the prediction direction of the current block, at least one of the motion information in the L0 direction or the motion information in the L1 direction can be used. The bidirectional weight value index designates the weight value applied to the L0 prediction block and the weight value applied to the L1 prediction block.
[0107] FIG. 6 is a flowchart showing an inter-prediction method according to an embodiment of the present invention.
[0108] Referring to FIG. 6, the inter prediction method includes a step of determining an inter prediction mode of a current block (S601), a step of obtaining motion information of the current block based on the determined inter prediction mode (S602), and a step of performing motion compensation prediction on the current block based on the obtained motion information (S603).
[0109] Here, the inter prediction mode represents a plurality of techniques for determining the motion information of the current block, and may include an inter prediction mode using translational motion information and an inter prediction mode using affine motion information. For example, the inter prediction mode using translational motion information may include a merge mode and an advanced motion vector prediction mode. The inter prediction mode using affine motion information may include an affine merge mode and an affine motion vector prediction mode. According to the inter prediction mode, the motion information of the current block can be determined based on information analyzed from adjacent blocks adjacent to the current block or a bit stream.
[0110] Hereinafter, the inter prediction method using affine motion information will be described in detail.
[0111] FIG. 7 is a diagram showing non-linear motion of an object.
[0112] Symmetric motion within a video may be non-linear motion. For example, as shown in the example of FIG. 7, there may be symmetric non-linear motion such as camera zoom-in, zoom-out, rotation, or affine transformation. When non-linear motion of an object occurs, the translational motion vector cannot effectively represent the motion of the object. Therefore, in the portion where non-linear motion of the object occurs, by using affine motion instead of translational motion, the coding efficiency can be improved.
[0113] FIG. 8 is a flowchart showing an inter prediction method based on affine motion according to an embodiment of the present invention.
[0114] Based on the information analyzed from the bitstream, it is possible to determine whether to apply an inter prediction technique based on affine motion to the current block. Specifically, based on at least one of a flag indicating whether to apply an affine merge mode to the current block or a flag indicating whether to apply an affine motion vector prediction mode to the current block, it is possible to determine whether to apply an inter prediction technique based on affine motion to the current block.
[0115] When applying an inter prediction technique based on affine motion to the current block, an affine motion model of the current block can be determined (S801). The affine motion model may be determined to be at least one of a six-parameter affine motion model or a four-parameter affine motion model. The six-parameter affine motion model represents affine motion with six parameters, and the four-parameter affine motion model represents affine motion with four parameters.
[0116] Equation 1 is the case where affine motion is represented by six parameters. The affine motion represents the translational motion of a predetermined region determined by an affine seed vector.
Equation
[0117] When affine motion is represented by six parameters, complex motion can be represented, but since the number of bits required for encoding each parameter increases, the encoding efficiency is reduced. Therefore, affine motion can also be represented by four parameters. Equation 2 is the case where affine motion is represented by four parameters.
Equation
[0118] Encoding can be performed on information for determining the affine motion model of the current block, and it can be transmitted as a signal via a bitstream. For example, the information may be a 1-bit flag "affine_type_flag". When the value of the flag is 0, it indicates that a 4-parameter affine motion model is applied. When the value of the flag is 1, it indicates that a 6-parameter affine motion model is applied. The flag can be encoded in units of a slice, a tile, or a block (e.g., an encoded block or an encoded tree unit). At the slice level, when transmitting the flag using a signal, the affine motion model determined at the slice level can be applied to all blocks belonging to the slice.
[0119] Alternatively, based on the affine inter-prediction mode of the current block, the affine motion model of the current block can be determined. For example, when applying the affine merge mode, the affine motion model of the current block can be determined as a 4-parameter motion model. On the other hand, when applying the affine motion vector prediction mode, encoding can be performed on the information for determining the affine motion model of the current block, and it can be transmitted as a signal via a bitstream. For example, when applying the affine motion vector prediction mode to the current block, the affine motion model of the current block can be determined based on a 1-bit flag "affine_type_flag".
[0120] Next, the affine seed vector of the current block can be derived (S802). When a 4-parameter affine motion model is selected, the motion vectors at two control points of the current block can be derived. When a 6-parameter affine motion model is selected, the motion vectors at three control points of the current block can be derived. The motion vectors at the control points can be referred to as affine seed vectors. The control points may include at least one of the upper left corner, the upper right corner, or the lower left corner of the current block.
[0121] FIG. 9 is a diagram showing examples of the affine seed vectors of each affine motion model.
[0122] In the 4-parameter affine motion model, the affine seed vectors related to two of the upper left corner, upper right corner, or lower left corner can be derived. For example, as shown in the example of FIG. 9(a), when the 4-parameter affine motion model is selected, the affine vector can be derived using the affine seed vector sv0 related to the upper left corner of the current block (e.g., the upper left sample (x0, y0)) and the affine seed vector sv1 related to the upper right corner of the current block (e.g., the upper right sample (x1, y1)). Also, instead of the affine seed vector related to the upper left corner, the affine seed vector related to the lower left corner can be used. Or, instead of the affine seed vector related to the upper right corner, the affine seed vector related to the lower left corner can be used.
[0123] In the 6-parameter affine motion model, the affine seed vectors related to the upper left corner, upper right corner, and lower left corner can be derived. For example, as shown in the example of FIG. 9(b), when the 6-parameter affine motion model is selected, the affine vector can be derived using the affine seed vector sv0 related to the upper left corner of the current block (e.g., the upper left sample (x0, y0)), the affine seed vector sv1 related to the upper right corner of the current block (e.g., the upper right sample (x1, y1)), and the affine seed vector sv2 related to the lower left corner of the current block (e.g., the lower left sample (x2, y2)).
[0124] In the embodiments described below, in the four-parameter affine motion model, the affine seed vectors of the upper left control point and the upper right control point are referred to as the first affine seed vector and the second affine seed vector, respectively. In the embodiments described below that use the first affine seed vector and the second affine seed vector, at least one of the first affine seed vector and the second affine seed vector can be replaced with the affine seed vector (the third affine seed vector) of the lower left control point or the affine seed vector (the fourth affine seed vector) of the lower right control point.
[0125] Note that in the six-parameter affine motion model, the affine seed vectors of the upper left control point, the upper right control point, and the lower left control point are referred to as the first affine seed vector, the second affine seed vector, and the third affine seed vector, respectively. In the embodiments described below that use the first affine seed vector, the second affine seed vector, and the third affine seed vector, at least one of the first affine seed vector, the second affine seed vector, and the third affine seed vector can be replaced with the affine seed vector (the fourth affine seed vector) of the lower right control point.
[0126] Using the affine seed vector, an affine vector can be derived for different sub-blocks (S803). Here, the affine vector represents a translational motion vector derived based on the affine seed vector. The affine vector of a sub-block can be referred to as an affine sub-block motion vector or a sub-block motion vector.
[0127] FIG. 10 is a diagram showing an example of the affine vector of a sub-block in a four-parameter motion model.
[0128] Based on the position of the control point, the position of the sub-block, and the affine seed vector, the affine vector of the sub-block can be derived. For example, Equation 3 shows an example of the derivation of the affine sub-block vector.
Equation
[0129] In Equation 3, (x, y) represents the position of the sub-block. Here, the position of the sub-block represents the position of the reference sample included in the sub-block. The reference sample may be a sample located at the upper left corner of the sub-block, or may be a sample located at the center position of at least one of its x-axis or y-axis coordinates. (x0, y0) represents the position of the first control point, and (sv 0x , sv 0y ) represents the first affine seed vector. Note that (x1, y1) represents the position of the second control point, and (sv 1x , sv 1y ) represents the second affine seed vector.
[0130] When the first control point and the second control point respectively correspond to the upper left corner and the upper right corner of the current block, x1 - x0 can be set to a value equal to the width of the current block.
[0131] Subsequently, motion compensation prediction can be performed for each sub-block using the affine vector of each sub-block (S804). After performing the motion compensation prediction, a prediction block related to each sub-block can be generated. The prediction block of the sub-block can be set as the prediction block of the current block.
[0132] Based on the affine seed vectors of the adjacent blocks adjacent to the current block, the affine seed vector of the current block can be derived. When the inter prediction mode of the current block is the affine merge mode, the affine seed vector of the merge candidate included in the merge candidate list can be determined as the affine seed vector of the current block. Note that when the inter prediction mode of the current block is the affine merge mode, the motion information including at least one of the reference picture index, the specific direction prediction flag, or the bidirectional weight value of the current block can be set to be the same as that of the merge candidate.
[0133] Merge candidates can be derived based on adjacent blocks of the current block. The adjacent blocks may include at least one of spatial adjacent blocks that are spatially adjacent to the current block and temporal adjacent blocks included in an image different from the current image.
[0134] FIG. 11 is a diagram showing adjacent blocks that can be used to derive merge candidates.
[0135] The adjacent blocks of the current block may include at least one of (A) an adjacent block adjacent to the left side of the current block, (B) an adjacent block adjacent to the information of the current block, (C) an adjacent block adjacent to the upper right corner of the current block, (D) an adjacent block adjacent to the lower left corner of the current block, or an adjacent block adjacent to the upper left corner of the current block. If the coordinates of the upper left sample of the current block are (x0, y0), the left adjacent block A includes samples at the position (x0 - 1, y0 + H - 1), and the upper adjacent block B includes samples at the position (x0 + W - 1, y0 - 1). Here, W and H represent the width and height of the current block, respectively. The upper right adjacent block C includes samples at the position (x0 + W, y0 - 1), and the lower left adjacent block D includes samples at the position (x0 - 1, y0 + H). The upper left adjacent block E includes samples at the position (x0 - 1, y0 - 1).
[0136] In the affine inter prediction mode, when encoding an adjacent block, an affine seed vector of a merge candidate can be derived based on the affine seed vector of the corresponding adjacent block. Hereinafter, an adjacent block encoded in the affine inter prediction mode is called an affine adjacent block.
[0137] By searching for adjacent blocks according to a pre-defined scanning order, merge candidates for the current block can be generated. The scanning order can be pre-defined in the encoder and the decoder. For example, adjacent blocks can be searched according to the order of A, B, C, D, E. Note that merge candidates can be sequentially derived from the searched affine adjacent blocks. Alternatively, the scanning order can be adaptively determined based on at least one of the size, shape, or affine motion model of the current block. That is, the scanning order of blocks having at least one of different sizes, shapes, or affine motion models is different.
[0138] Alternatively, search for the blocks located above the current block in sequence, derive merge candidates from the first discovered affine adjacent block, and also search for the blocks located on the left side of the current block in sequence, and derive merge candidates from the first discovered affine adjacent block. Here, the plurality of adjacent blocks located above the current block may include at least one of adjacent block E, adjacent block B, or adjacent block C, and the plurality of blocks located on the left side of the current block may include at least one of block A or block D. In this case, adjacent block E may be classified as a block located on the left side of the current block.
[0139] Although not shown, merge candidates can be derived from the temporal adjacent blocks of the current block. Here, the temporal adjacent blocks may include blocks located at the same position as the current block in the collocated image or blocks adjacent thereto. Specifically, in the affine inter prediction mode, when encoding the temporal adjacent blocks of the current block, merge candidates can be derived based on the affine seed vectors of the temporal merge candidates.
[0140] A merge candidate list including merge candidates can be generated, and one affinity seed vector among the merge candidates included in the merge candidate list can be determined as the affinity seed vector of the current block. For this purpose, encoding can be performed on the index information that labels any one of the plurality of merge candidates, and it can be transmitted via a bit stream.
[0141] As another example, a plurality of adjacent blocks can be searched according to the scanning order, and the affinity seed vector of the current block can be derived from the affinity seed vector of the first discovered affine adjacent block.
[0142] As described above, in the affine merge mode, the affinity seed vector of the current block can be derived using the affinity seed vector of the adjacent block.
[0143] When the inter prediction mode of the current block is the affine motion vector prediction mode, the affinity seed vector of the motion vector prediction candidate included in the motion vector prediction candidate list can be determined as the predicted value of the affinity seed vector of the current block. By adding the affinity seed vector difference value to the predicted value of the affinity seed vector, the affinity seed vector of the current block can be derived.
[0144] Based on the adjacent blocks of the current block, an affinity seed vector prediction candidate can be derived. Specifically, according to a predetermined scanning order, a plurality of adjacent blocks located above the current block are searched, and a first affinity seed vector prediction candidate can be derived from the first discovered affine adjacent block. In addition, according to a predetermined scanning order, a plurality of adjacent blocks located on the left side of the current block are searched, and a second affinity seed vector prediction candidate can be derived from the first discovered affine adjacent block.
[0145] Encoding can be performed on information for determining an affinity seed vector difference value, and it can be transmitted via a bit stream. The information may include size information representing the size of the affinity seed vector difference value and symbol information representing the symbol of the affinity seed vector difference value. The affinity seed vector difference value related to each control point can be set to be the same. Alternatively, for each control point, the affinity seed vector difference value can be set to be different.
[0146] As described above, an affinity seed vector of a merge candidate or an affinity seed vector prediction candidate can be derived from the affinity seed vector of an affinity adjacent block, and the affinity seed vector of the current block can be derived using the derived affinity seed vector of the merge candidate or the affinity seed vector prediction candidate. Alternatively, after searching for a plurality of affinity adjacent blocks according to a predetermined scanning order, the affinity seed vector of the current block can also be derived from the affinity seed vector of the first found affinity adjacent block.
[0147] In the following, a method for deriving the affinity seed vector of the current block, a merge candidate, or an affinity seed vector prediction candidate from the affinity seed vector of the affinity adjacent block will be described in detail. In the embodiments described below, the derivation of the affinity seed vector of the current block may be understood as the derivation of the affinity seed vector of the merge candidate or the derivation of the affinity seed vector of the affinity seed vector prediction candidate.
[0148] FIG. 12 is a diagram showing the derivation of the affinity seed vector of the current block based on the affinity seed vector of the affinity adjacent block.
[0149] When the first affine seed vector nv0 related to the upper left control point and the second affine seed vector nv1 related to the upper right control point are stored in the affine adjacent block, based on the first affine seed vector and the second affine seed vector, the third affine seed vector nv2 related to the lower left control point of the affine adjacent block can be derived. Equation 4 shows an example of the derivation of the third affine seed vector.
Equation
[0150] In Equation 4, (nv 0x , nv 0y ) represents the first affine seed vector nv0, (nv 1x , nv 1y ) represents the second affine seed vector nv1, and (nv 2x , nv 2y ) represents the third affine seed vector nv2. Note that (x n0 , x n0 ) represents the position of the first control point, (x n1 , x n1 ) represents the position of the second control point, and (x n2 , x n2 ) represents the position of the third control point.
[0151] Subsequently, the affine seed vector of the current block can be derived using the first affine seed vector, the second affine seed vector, and the third affine seed vector. Equation 5 shows an example of the derivation of the first affine seed vector v0 of the current block, and Equation 6 shows an example of the derivation of the second affine seed vector v1 of the current block.
Equation
Equation
[0152] In Equation 5 and Equation 6, (v 0x , v0y ) represents the first affinity seed vector sv0 of the current block, and (v 1x , v 1y ) represents the second affinity seed vector sv1 of the current block. Here, (x0, y0) represents the position of the first control point, and (x1, y1) represents the position of the second control point. For example, the first control point represents the upper left corner of the current block, and the second control point represents the upper right corner of the current block.
[0153] In the above example, it was explained that a plurality of affinity seed vectors of the current block are derived using three affinity seed vectors related to the affinity adjacent blocks. As another example, it is also possible to derive the affinity seed vector of the current block using only two of the plurality of affinity seed vectors of the affinity adjacent blocks.
[0154] Alternatively, without using the first affinity seed vector at the upper left corner, the second affinity seed vector at the upper right corner, or the third affinity seed vector at the lower left corner related to the affinity adjacent blocks, a plurality of affinity seed vectors of the current block can be derived using the fourth affinity seed vector related to the lower right corner.
[0155] In particular, when the upper boundary of the current block is in contact with the upper boundary of the coding tree unit and an affine adjacent block adjacent above the current block (hereinafter referred to as the upper affine adjacent block) attempts to use the affine seed vector of the upper control point (for example, the upper left corner or the upper right corner), it is necessary to pre-store these in the memory. This may cause the problem of an increase in the number of line buffers. Therefore, when the upper boundary of the current block is in contact with the upper boundary of the coding tree unit, it can be set to use the affine seed vector of the lower control point (for example, the lower left corner or the lower right corner) for the upper affine adjacent block without using the affine seed vector of the upper control point. For example, a plurality of affine seed vectors of the current block can be derived using the third affine seed vector related to the lower left corner and the fourth affine seed vector related to the lower right corner of the upper affine adjacent block. In this case, the affine seed vector related to the lower corner can be derived by replicating the affine seed vector related to the upper corner, or can be derived from a plurality of affine seed vectors related to the upper corner. For example, the first affine seed vector, the second affine seed vector, or the third affine seed vector can be converted / replaced with the fourth affine seed vector related to the lower right corner.
[0156] Equations 7 and 8 show an example of deriving the first affine seed vector and the second affine seed vector of the current block using the third affine seed vector related to the lower left control point and the fourth affine seed vector related to the lower right control point of the adjacent affine vector.
Number
Number
[0157] In Equations 7 and 8, (x n2 , y n2 ) represents the coordinates of the lower left control point of the affine adjacent block, and also, (x n3 , y n3) represents the coordinates of the lower-right control point of the affine adjacent block. (x0, y0) represents the coordinates of the upper-left control point of the current block, and (x1, y1) represents the coordinates of the upper-right control point of the current block. (nv 2x , nv 2y ) represents the affine seed vector (i.e., the third affine seed vector) of the lower-left control point of the affine adjacent block, and (nv 3x , nv 3y ) represents the affine seed vector (i.e., the fourth affine seed vector) of the lower-right control point of the affine adjacent block. (v 0x , v 0y ) represents the affine seed vector (i.e., the first affine seed vector) of the upper-left control point of the current block, and (v 1x , v 1y ) represents the affine seed vector (i.e., the second affine seed vector) of the upper-right control point of the current block.
[0158] The division operations included in Equation 7 and Equation 8 can also be changed to shift operations. The shift operations can be executed based on the value derived from the width between the lower-left control point and the lower-right control point (i.e., (x n3 -x n2 ).
[0159] In the above example, a plurality of affine seed vectors of the current block can be derived based on a plurality of affine seed vectors of the encoded / decoded affine adjacent blocks. For this purpose, a plurality of affine seed vectors of the encoded / decoded affine adjacent blocks can be stored in the memory. However, in addition to the plurality of translational motion vectors (i.e., a plurality of affine vectors) of the plurality of sub-blocks included in the affine adjacent block, storing a plurality of affine seed vectors of the affine adjacent block in the memory further causes a problem of an increase in the amount of memory used. To solve such a problem, the affine seed vector of the current block can be derived using the motion vector of the sub-block adjacent to the control point of the affine adjacent block. This replaces the affine seed vector of the affine adjacent block. That is, the motion vector of the sub-block adjacent to the control point of the affine adjacent block can be set as the affine seed vector of the affine adjacent block. Here, the sub-block is a block having a size / shape predefined by the encoder and the decoder, and may also be a block having a basic size / shape for storing the motion vector. For example, the sub-block may be a square block having a size of 4×4. Alternatively, the motion vector specifying the sample position can be set as the affine seed vector of the affine adjacent block.
[0160] FIG. 13 is a diagram showing an example in which the motion vector of the sub-block is used as the affine seed vector of the affine adjacent block.
[0161] The motion vector of the sub-block adjacent to the control point can be set as the affine seed vector of the corresponding control point. For example, in the example shown in FIG. 13, the motion vector (nv 4x , nv 4y ) of the sub-block (lower left sub-block) adjacent to the lower left corner of the affine adjacent block can be set as the affine seed vector (nv 2x , nv 2y ) of the lower left control point, and the motion vector (nv5x , nv 5y ) as the affine seed vector (nv 3x , nv 3y ) of the control point in the lower right corner. Here, the lower left sub-block is a sub-block that includes samples adjacent to the lower left control point (x n2 , y n2 ) of the adjacent affine block (for example, samples at the position (x n2 , y n2 - 1)). Also, the lower right sub-block is a block that includes samples adjacent to the lower right control point (x n3 , y n3 ) of the adjacent affine block (for example, samples at the position (x n3 - 1, y n3 - 1)). When deriving the affine seed vector of the current block based on Equation 7 and Equation 8, instead of the third affine seed vector of the affine adjacent block, the motion vector of the lower left sub-block can be used, and instead of the fourth affine seed vector, the motion vector of the lower right sub-block can be used.
[0162] Hereinafter, in the embodiments described later, the sub-block used as the affine seed vector of the affine adjacent block is called an affine sub-block.
[0163] According to an embodiment of the present invention, an affine sub-block can be determined based on samples at a specific position. For example, a sub-block that includes samples at a specific position can also be used as an affine sub-block. Hereinafter, the samples at a specific position are called affine reference samples. Note that the reference sample for determining the affine sub-block of the lower left control point is called the lower left reference sample, and the reference sample for determining the affine sub-block of the lower right control point is called the lower right reference sample.
[0164] The lower left reference sample and the lower right reference sample may be selected from a plurality of samples included in the affine adjacent blocks. For example, at least one of the upper left sample, the lower left sample, the upper right sample, or the lower left sample of the lower left sub-block may be used as the lower left reference sample, and at least one of the upper left sample, the lower left sample, the upper right sample, or the lower left sample of the lower right sub-block may be used as the lower right reference sample. Therefore, the motion vectors of the lower left sub-block including the lower left reference sample and the lower right sub-block including the lower right reference sample can be the affine seed vectors related to the lower left control point and the affine seed vectors related to the lower right control point, respectively.
[0165] As another example, at least one of the lower left reference sample or the lower right reference sample can be a sample located outside the affine adjacent block. This will be described in detail with reference to FIGS. 14 to 16.
[0166] FIGS. 14 to 16 are diagrams showing the positions of the reference samples.
[0167] For example, as shown in FIG. 14(a), for the lower left control point, the upper left sample of the lower left sub-block can be used as the reference sample (x n4 , y n4 ). Therefore, the lower left sub-block including the reference sample (x n4 , y n4 ) can be used as the affine sub-block related to the lower left control point.
[0168] For the lower right control point, the sample located on the right side of the upper right sample of the lower right sub-block can be used as the reference sample (x n5 , y n5 ). Therefore, the sub-block adjacent to the right side of the lower right sub-block including the reference sample (x n5 , y n5 ) can be used as the affine sub-block related to the lower right control point.
[0169] Alternatively, as in the example shown in Fig. 14(b), for the lower left control point, a sample located on the left side of the upper left sample of the lower left sub-block can be used as the reference sample (x n4 , y n4 ). Therefore, the sub-block adjacent to the left side of the lower left sub-block containing the reference sample (x n4 , y n4 ) can be defined as the affine sub-block related to the lower left control point.
[0170] For the lower right control point, the upper right sample of the lower right sub-block can be used as the reference sample (x n5 , y n5 ). Therefore, the lower right sub-block containing the reference sample (x n5 , y n5 ) can be defined as the affine sub-block related to the lower right control point.
[0171] Alternatively, as in the example shown in Fig. 15(a), for the lower left control point, the lower left sample of the lower left sub-block can be used as the reference sample (x n4 , y n4 ). Therefore, the lower left sub-block containing the reference sample (x n4 , y n4 ) can be defined as the affine sub-block related to the lower left control point.
[0172] For the lower right control point, a sample located on the right side of the lower right sample of the lower right sub-block can be used as the reference sample (x n5 , y n5 ). Therefore, the sub-block adjacent to the right side of the lower right sub-block containing the reference sample (x n5 , y n5 ) can be defined as the affine sub-block related to the lower right control point.
[0173] Alternatively, as in the example shown in Fig. 15(b), for the lower left control point, a sample located on the left side of the lower left sample of the lower left sub-block can be used as the reference sample (x n4 , y n4 ). Therefore, the reference sample (xn4 , y n4 The sub-block adjacent to the left side of the lower left sub-block including ) can be set as an affine sub-block related to the lower left control point.
[0174] For the lower right control point, the lower right sample of the lower right sub-block can be set as the reference sample (x n5 , y n5 ). Therefore, the lower right sub-block including the reference sample (x n5 , y n5 ) can be set as an affine sub-block related to the lower right control point.
[0175] Alternatively, as in the example shown in (a) of FIG. 16, for the lower left control point, the sample located between the upper left sample and the lower left sample of the lower left sub-block (for example, the left intermediate sample) can be set as the reference sample (x n4 , y n4 ). Therefore, the lower left sub-block including the reference sample (x n4 , y n4 ) can be set as an affine sub-block related to the lower left control point.
[0176] For the lower right control point, the sample located on the right side of the sample between the upper right sample and the lower right sample of the lower right sub-block (for example, the right intermediate sample) can be set as the reference sample (x n5 , y n5 ). Therefore, the sub-block adjacent to the right side of the lower right sub-block including the reference sample (x n5 , y n5 ) can be set as an affine sub-block related to the lower right control point.
[0177] Alternatively, as in the example shown in (b) of FIG. 16, for the lower left control point, the sample located on the left side of the sample between the upper left sample and the lower left sample of the lower left sub-block can be set as the reference sample (x n4 , y n4 ). Therefore, the reference sample (x n4 , y n4The sub-block adjacent to the left side of the lower left sub-block including can be an affine sub-block related to the lower left control point.
[0178] For the lower right control point, a sample located between the upper right sample and the lower right sample of the lower right sub-block can be used as the reference sample (x n5 , y n5 ). Therefore, the lower right sub-block including the reference sample (x n5 , y n5 ) can be an affine sub-block related to the lower right control point.
[0179] When deriving a plurality of affine seed vectors of the current block based on Equation 7 and Equation 8, instead of the third affine seed vector of the affine adjacent block, the motion vector of the affine sub-block related to the lower left control point can be used, and instead of the fourth affine seed vector, the motion vector of the affine sub-block related to the lower right control point can be used. Note that instead of the position of the lower left control point, the position of the lower left reference sample can be used, and instead of the position of the lower right control point, the position of the lower right reference sample can be used.
[0180] Different from the description in FIGS. 14 to 16, a sub-block including a sample adjacent to the reference sample can also be an affine sub-block. Specifically, a sample located outside the affine adjacent sub-block can be used as the reference sample, and a sub-block included in the affine adjacent block can be an affine sub-block. For example, in the example shown in FIG. 14(a), a sample located to the right of the upper right sample of the lower right sub-block can be used as the reference sample (x n5 , y n5 ), and the lower right sub-block can be an affine sub-block related to the lower right corner. Alternatively, in the example shown in FIG. 14(b), a sample located to the left of the upper left sample of the lower left sub-block can be used as the reference sample (x n4 , y n4It can be set as such, and the lower left sub-block can be an affine sub-block related to the lower left corner.
[0181] The embodiments described in FIGS. 15 and 16 can be similarly applied. That is, in the example shown in (a) of FIG. 15 or (a) of FIG. 16, the reference sample is the sample located on the right side of the lower right sample of the lower right sub-block or the right side intermediate sample (x n5 , y n5 ) It can be set as such, and the lower right sub-block can be an affine sub-block related to the lower right corner. Alternatively, in the example shown in (b) of FIG. 15 or (b) of FIG. 16, the reference sample is the sample located on the left side of the lower left sample of the lower left sub-block or the left side intermediate sample (x n4 , y n4 ) It can be set as such, and the lower left sub-block can be an affine sub-block related to the lower left corner.
[0182] In the above-described example, the affine seed vector of the affine adjacent block can be derived by using the motion vector of the affine sub-block. For this purpose, for the encoded / decoded block, the motion vector can be stored in units of sub-blocks.
[0183] As another example, after storing the minimum number of affine seed vectors in the affine adjacent block, the motion vector of the affine sub-block can be derived by using the plurality of stored affine seed vectors.
[0184] Equations 9 and 10 show an example of deriving the motion vector of the affine sub-block by using the affine seed vector of the affine adjacent block.
Number
Number
[0185] In Equation 9 and Equation 10, (nv 4x , nv 4y ) represents the motion vector of the affine sub-block related to the lower left control point, and (nv 5x , nv 5y ) represents the motion vector of the affine sub-block related to the lower right control point. Since the motion vector of the affine sub-block is set to be the same as the affine seed vector of the control point, instead of (nv 4x , nv 4y ), the affine seed vector (nv 2x , nv 2y ) related to the lower left control point can be used, or instead of (nv 5x , nv 5y ), the affine seed vector (nv 3x , nv 3y ) related to the lower right control point can be used.
[0186] (x n4 , y n4 ) represents the position of the reference sample of the lower left sub-block. Alternatively, instead of this position, the center position of the lower left sub-block or the position of the lower left control point can also be used. (x n5 , y n5 ) represents the position of the reference sample of the lower right sub-block. Alternatively, instead of this position, the center position of the lower right sub-block or the position of the lower right control point can also be used.
[0187] Equation 9 and Equation 10 can be applied when the current block does not touch the boundary of the coding tree unit. When the current block touches the upper boundary of the coding tree unit, instead of applying Equation 9 and Equation 10, the translational motion vector of the affine sub-block determined based on the lower left reference sample can be used as the third affine seed vector, and the translational motion vector of the affine sub-block determined based on the lower right reference sample can be used as the fourth affine seed vector.
[0188] In Equation 7 and Equation 8, (x n3 - x n2) represents the width between the lower left control point and the lower right control point. As described above, instead of x n3 , the position x n5 of the lower right reference sample can be used, and instead of x n2 , the position x n4 of the lower left reference sample can be used. Hereinafter, (x n3 - x n2 ), or the value obtained by using the position of the reference sample instead of the above equation (for example, (x n5 - x n4 )) is defined as the variable W seed , and this variable is called the sub-seed vector width.
[0189] Depending on the position of the reference sample, it is possible that the sub-seed vector width is not a power of 2 (for example, 2 n ). For example, when the lower left sample of the lower left sub-block is used as the lower left reference sample and the lower right sample of the lower right sub-block is used as the lower right reference sample, the sub-seed vector width is not a multiple of 2. As described above, when the sub-seed vector width is not a power of 2, the sub-seed vector width can be converted to a power of 2. The conversion may include adding / subtracting only an offset to the sub-seed vector width, or using the position of the sample adjacent to the reference sample instead of the position of the reference sample. For example, by adding 1 to the width between the lower left reference sample and the lower right reference sample, the converted sub-seed vector width can be derived. Alternatively, the width between the adjacent reference sample adjacent to the right of the lower right reference sample and the lower left reference sample can be used as the converted sub-seed vector width. Subsequently, by substituting the converted sub-seed vector width into Equation 7 and Equation 8, the affine seed vector of the current block can be derived.
[0190] The division included in Equation 7 and Equation 8 can also be changed to a shift operation. The shift operation can be executed based on a value derived from the converted sub-seed vector width (i.e., a value represented as a power of 2).
[0191] When the reference sample for determining the affine sub-block does not belong to the affine adjacent block, the affine seed vector of the affine adjacent block can be derived based on the sample adjacent to the reference sample among the plurality of samples included in the affine adjacent block. Specifically, the translational motion vector of the sub-block including the sample adjacent to the reference sample (hereinafter referred to as the adjacent reference sample) in the affine adjacent block can be used as the affine seed vector of the affine adjacent block. As described above, the method of deriving the affine seed vector using the adjacent reference sample can be defined as a modified affine merge vector derivation method.
[0192] FIG. 17 is a diagram showing an application example of the modified affine merge vector derivation method.
[0193] If the lower-right reference sample (x n5 , y n5 ) of the affine adjacent block E does not belong to the affine adjacent block, the affine seed vector can be derived based on the sample adjacent to the left of the lower-right reference sample among the samples included in the affine adjacent block (x n5 -1, y n5 ). Specifically, the translational motion vector of the sub-block including the adjacent reference sample (x n5 -1, y n5 ) can be used as the affine seed vector of the lower-right control point.
[0194] In the example shown in FIG. 17, the sample adjacent to the right of the upper-right sample of the lower-right sub-block is shown as the lower-right reference sample. When the sample adjacent to the right of the lower-right sample of the lower-right sub-block or the sample adjacent to the right of the right-middle sample of the lower-right sub-block is used as the lower-right reference sample, the affine seed vector can be derived based on the sample adjacent to the left of the adjacent reference sample.
[0195] In addition, when the lower-left reference sample does not belong to the affine adjacent block, according to the described embodiment, an affine seed vector can also be derived based on the sample adjacent to the right side of the lower-left reference sample.
[0196] By setting the position of the reference sample and the sub-block for deriving the affine seed vector in different ways, the sub-seed vector width can be made a power series of 2.
[0197] By using the adjacent blocks around the current block that are not encoded in the affine inter mode, the merge candidate, affine seed vector prediction candidate, or affine seed vector of the current block can be derived. Specifically, non-affine adjacent blocks can be combined, and the combination can be used as a merge candidate or an affine seed vector prediction candidate. For example, at least one combination of the motion vectors of any one of the adjacent blocks adjacent to the upper-left corner of the current block, the motion vectors of any one of the adjacent blocks adjacent to the upper-right corner of the current block, and the motion vectors of any one of the adjacent blocks adjacent to the lower-left corner of the current block can be used as a merge candidate or an affine seed vector prediction candidate. In this case, the motion vectors of the adjacent blocks adjacent to the upper-left corner, the upper-right corner, and the lower-left corner can be respectively set as the first affine seed vector of the upper-left control point, the second affine seed vector of the upper-right control point, and the third affine seed vector of the lower-left control point.
[0198] Alternatively, in the above-described modified affine merge vector derivation method, the merge candidate, affine seed vector prediction candidate, or affine seed vector of the current block can be derived by using the adjacent blocks that are not encoded in the affine inter mode. Hereinafter, the adjacent blocks that are not encoded in the affine inter mode are referred to as non-affine adjacent blocks.
[0199] FIG. 18 is a diagram showing an example of deriving an affine seed vector of a current block based on non-affine adjacent blocks.
[0200] In the example shown in FIG. 18, it is assumed that all adjacent blocks adjacent to the current block are non-affine adjacent blocks.
[0201] When attempting to derive the affine seed vector of the current block from non-affine adjacent block A among the adjacent blocks adjacent to the current block, the lower left reference sample and the lower right reference sample of A can be set. For example, the sample adjacent to the left side of the lower left sample of block A can be used as the lower left reference sample, and the lower right sample of block A can be used as the lower right reference sample. Since the lower left reference sample is outside block A, the motion vector of the sub-block including the sample adjacent to the right side of the lower left reference sample can be used as the third affine seed vector of block A. Note that the motion vector of the sub-block including the lower right reference sample can be used as the fourth affine seed vector of block A. Subsequently, based on equations 9 and 10, the first affine seed vector and the second affine seed vector of the current block can be derived from block A.
[0202] The method of deriving an affine seed vector from a non-affine adjacent block can be used only when performing motion compensation prediction for the non-affine adjacent block in units of sub-blocks. Here, the prediction technique for performing motion compensation prediction as a sub-block may include at least one of STMVP, ATMVP, bidirectional optical flow (BIO), overlapped block motion compensation (OBMC), and decoder-side motion vector refinement (DMVR).
[0203] In the above embodiment, when the upper boundary of the current block touches the boundary of the coding tree unit, the third affine seed vector of the lower left control point and the fourth affine seed vector of the lower right control point of the affine adjacent block located above the current block are used to derive the merge candidate, affine seed vector prediction candidate, or affine seed vector of the current block, which has been described.
[0204] As another example, when the upper boundary of the current block touches the boundary of the coding tree unit and the adjacent block located above the current block belongs to a coding tree unit different from the coding tree unit of the current block, without using the adjacent block, among these blocks included in the coding tree unit to which the current block belongs, the adjacent block closest to the adjacent block is used to derive the merge candidate, affine seed vector prediction candidate, or affine seed vector of the current block.
[0205] In the example shown in FIG. 19, it is shown that the current block touches the upper boundary of the coding tree unit and the blocks B, C, and E located above the current block belong to coding tree units different from the coding tree unit of the current block. Therefore, instead of block E, among these blocks included in the coding tree unit to which the current block belongs, block F adjacent to block E is used to derive the affine seed vector of the current block.
[0206] For motion compensation prediction of the current block, the affine seed vectors of a plurality of blocks can be used. For example, a plurality of merge candidates can be selected from the merge candidate list, and based on the affine seed vectors of the selected merge candidates, the affine seed vector or sub-block vector of the current block can be derived. Performing encoding / decoding on the current block using the affine seed vectors of a plurality of blocks may be referred to as a multi-affine merge encoding method.
[0207] Information indicating whether the multi-affine merge coding method is applied to the current block can be coded and transmitted via a bit stream. Alternatively, based on at least one of the number of affine adjacent blocks among adjacent blocks adjacent to the current block, the number of merge candidates included in the merge candidate list, and the affine motion model of the current block, it is possible to determine whether to apply the multi-affine merge coding method to the current block.
[0208] FIGS. 20 and 21 are flowcharts showing a motion compensation prediction method using a plurality of merge candidates.
[0209] FIG. 20 is a diagram showing an example of deriving the affine seed vector of the current block by using the affine seed vectors of a plurality of merge candidates. FIG. 21 is a diagram showing an example of deriving the motion vector of each sub-block by using the affine seed vectors of a plurality of merge candidates.
[0210] The affine seed vector of the current block can be generated based on the sum, difference, average value, or weighted addition of the affine seed vectors of two merge candidates.
[0211] The following equations 11 and 12 show examples of deriving the affine seed vector of the current block by adding the affine seed vectors of the merge candidates.
Equation
Equation
[0212] In Equations 11 and 12, sv4 represents the first affinity seed vector of the current block, sv0 represents the first affinity seed vector of the first merge candidate, and sv2 represents the first affinity seed vector of the second merge candidate. Note that sv5 represents the second affinity seed vector of the current block, sv1 represents the second affinity seed vector of the first merge candidate, and sv3 represents the second affinity seed vector of the second merge candidate.
[0213] Note that the following Equations 13 and 14 show an example of deriving the affinity seed vector of the current block by weighted addition of the affinity seed vectors of the merge candidates.
Number
Number
[0214] As another example, based on the affinity seed vector of the first merge candidate and the affinity seed vector of the second merge candidate, for each sub-block in the current block, a first sub-block motion vector and a second sub-block motion vector can be generated. Subsequently, based on the sum, difference, average value, or weighted addition of the first sub-block motion vector and the second sub-block motion vector, a final sub-block motion vector can be generated.
[0215] The following Equation 15 shows an example of obtaining a final sub-block motion vector by adding the first sub-block motion vector and the second sub-block motion vector.
Number
[0216] In Equation 15, V0 represents the first sub-block motion vector, V1 represents the second sub-block motion vector, and V2 represents the final sub-block motion vector.
[0217] Note that the following Equation 16 shows an example of deriving a final sub-block motion vector by weighted addition of a first sub-block motion vector and a second sub-block motion vector.
Number
[0218] Intra prediction is to predict the current block by using the encoded / decoded reconstructed samples around the current block. In this case, the intra prediction of the current block can use the reconstructed samples before applying the in-loop filter.
[0219] Intra prediction techniques include intra prediction based on a matrix and general intra prediction considering the directionality with the surrounding reconstructed samples. Information indicating the intra prediction technique of the current block can be transmitted in the signal via the bitstream. The information may be a 1-bit flag. Alternatively, the intra prediction technique of the current block can be determined based on at least one of the position, size, shape of the current block, or the intra prediction technique of the adjacent blocks. For example, when the current block exists across the boundary of the image, the current block can be set not to apply the intra prediction based on the matrix.
[0220] The intra prediction based on the matrix is a method of obtaining a predicted block of the current block based on the matrix multiplication of the matrix stored in the encoder and decoder and the reconstructed samples around the current block. Information indicating any one of the plurality of stored matrices can be transmitted in the signal via the bitstream. The decoder can determine the matrix used for the intra prediction of the current block based on the information and the size of the current block.
[0221] General intra prediction is a method of obtaining a predicted block of a current block based on a non - angular intra prediction mode or an angular intra prediction mode. Hereinafter, with reference to the drawings, the process of performing intra prediction based on general intra prediction will be described in detail.
[0222] FIG. 22 is a flowchart showing an intra prediction method according to an embodiment of the present invention.
[0223] A reference sample line of the current block can be determined (S2201). The reference sample line is a set of reference samples included in the K - th line shifted from above and / or to the left of the current block. The reference samples can be derived from the encoded / decoded reconstructed samples around the current block.
[0224] Via a bitstream, an index information for labeling, with a signal, the reference sample line of the current block among a plurality of reference sample lines can be transmitted. The plurality of reference sample lines may be included in at least one of the first line, the second line, the third line, or the fourth line located above and / or to the left of the current block. Table 1 shows the indexes assigned to each reference sample line. In Table 1, it is assumed that the first line, the second line, and the fourth line are used as reference sample line candidates. [Table 1]
[0225] The reference sample line of the current block can also be determined based on at least one of the position, size, shape of the current block, or the prediction coding mode of an adjacent block. For example, when the current block is in contact with the boundary of an image, tile, slice, or coding tree unit, the first reference sample line can be determined as the reference sample line of the current block.
[0226] The reference sample line may include an upper reference sample located above the current block and a left reference sample located to the left of the current block. The upper reference sample and the left reference sample can be derived from the reconstructed samples around the current block. The reconstructed samples may be in a state before applying the in-loop filter.
[0227] FIG. 23 is a diagram showing the reference samples included in each reference sample line.
[0228] According to the intra prediction mode of the current block, at least one of the reference samples belonging to the reference sample line can be used to obtain a prediction sample.
[0229] Next, the intra prediction mode of the current block can be determined (S2202). Regarding the intra prediction mode of the current block, at least one of the non-angle intra prediction mode or the angle intra prediction mode can be determined as the intra prediction mode of the current block. The non-angle intra prediction mode includes Planer and DC, and the angle intra prediction mode includes 33 or 65 modes from the lower left diagonal direction to the upper right diagonal.
[0230] FIG. 24 is a diagram showing the intra prediction mode.
[0231] FIG. 24(a) shows 35 intra prediction modes, and FIG. 24(b) shows 67 intra prediction modes.
[0232] It is also possible to define a number of intra prediction modes that is more or less than the number shown in FIG. 24.
[0233] Based on the intra prediction mode of adjacent blocks adjacent to the current block, the most probable mode (MPM) can be set. Here, the adjacent blocks may include a left adjacent block adjacent to the left side of the current block and an upper adjacent block adjacent to the upper side of the current block. When the coordinates of the upper left sample of the current block are (0, 0), the left adjacent block may include samples at positions (-1, 0), (-1, H - 1), or (-1, (H - 1) / 2). Here, H represents the height of the current block. The upper adjacent block may include samples at positions (0, -1), (W - 1, -1), or ((W - 1) / 2, -1). Here, W represents the width of the current block.
[0234] When performing encoding on adjacent blocks by general intra prediction, the MPM can be derived based on the intra prediction mode of the adjacent blocks. Specifically, the intra prediction mode of the left adjacent block can be set as the variable candIntraPredModeA, and the intra prediction mode of the upper adjacent block can be set as the variable candIntraPredModeB.
[0235] In this case, when the adjacent blocks are unavailable (for example, when the adjacent blocks have not been encoded / decoded or the positions of the adjacent blocks are shifted from the image boundary), when the adjacent blocks are encoded by matrix-based intra prediction, when the adjacent blocks are encoded by inter prediction, or when the adjacent blocks are included in a different coding tree unit from the current block, the variable candIntraPredModeX (where X is A or B) derived based on the intra prediction mode of the adjacent blocks can be set as the default mode. Here, the default mode may include at least one of the planar mode, the DC mode, the vertical mode, or the horizontal mode.
[0236] Alternatively, when encoding adjacent blocks by matrix-based intra prediction, the intra prediction mode corresponding to the index value for specifying any one of the matrices can be set as candIntraPredModeX. For this purpose, a look-up table indicating the mapping relationship between the index value for specifying the matrix and the intra prediction mode can be stored in advance in the encoder and the decoder.
[0237] The MPM can be derived based on the variables candIntraPredModeA and candIntraPredModeB. In the encoder and the decoder, the number of MPMs included in the MPM list can be predefined. For example, the number of MPMs may be 3, 4, 5, or 6. Alternatively, information representing the number of MPMs can be transmitted as a signal via the bitstream. Alternatively, the number of MPMs can be determined based on at least one of the prediction coding modes of adjacent blocks, the size, or the shape of the current block.
[0238] In the embodiments described below, it is assumed that the number of MPMs is 3, and these 3 MPMs are referred to as MPM[0], MPM[1], and MPM[2]. When the number of MPMs is more than 3, the MPM may include the 3 MPMs described in the embodiments below.
[0239] When candIntraPredA is the same as candIntraPredB and candIntraPredA is in the planar mode or the DC mode, MPM[0] and MPM[1] can be set as the planar mode and the DC mode, respectively. MPM[2] can be set as the vertical intra prediction mode, the horizontal intra prediction mode, or the diagonal intra prediction mode. The diagonal intra prediction mode may be the lower left diagonal intra prediction mode, the upper left intra prediction mode, or the upper right intra prediction mode.
[0240] When candIntraPredA is the same as candIntraPredB and candIntraPredA is in an intra prediction mode, MPM[0] can be set to be the same as candIntraPredA. MPM[1] and MPM[2] can be set to intra prediction modes similar to candIntraPredA. The intra prediction mode similar to candIntraPredA may be an intra prediction mode whose index difference value from candIntraPredA is ±1 or ±2. Using modulo operation (%) and offset, the intra prediction mode similar to candIntraPredA can be derived.
[0241] When candIntraPredA is different from candIntraPredB, MPM[0] can be set to be the same as candIntraPredA and MPM[1] can be set to be the same as candIntraPredB. In this case, when both candIntraPredA and candIntraPredB are non - angular intra prediction modes, MPM[2] can be set to the vertical intra prediction mode, the horizontal intra prediction mode or the diagonal intra prediction mode. Alternatively, when at least one of candIntraPredA and candIntraPredB is an angular intra prediction mode, MPM[2] can be set to the intra prediction mode derived by adding or subtracting an offset to the larger value of planar, DC, candIntraPredA or candIntraPredB. Here, the offset may be 1 or 2.
[0242] An MPM list including a plurality of MPMs can be generated, and information indicating whether an MPM that is the same as the intra prediction mode of the current block is included in the MPM list can be transmitted via a bitstream as a signal. The said information is a 1-bit flag and may be called an MPM flag. When the MPM flag indicates that an MPM that is the same as the current block is included in the MPM list, index information identifying one of the MPMs can be transmitted via the bitstream as a signal. The MPM specified by the said index information can be used as the intra prediction mode of the current block. When the MPM flag indicates that an MPM that is the same as the current block is not included in the MPM list, residual mode information indicating any one of the residual intra prediction modes other than MPM can be transmitted via the bitstream as a signal. The residual mode information represents an index value corresponding to the intra prediction mode of the current block when reassigning an index to the residual intra prediction modes other than MPM. The decoder can determine the intra prediction mode of the current block by arranging the MPMs in ascending order and comparing the residual mode information with the MPMs. For example, when the residual mode information is the same as or smaller than the MPM, the intra prediction mode of the current block can be derived by adding 1 to the residual mode information.
[0243] Instead of the operation of setting the default mode to MPM, information indicating whether the intra prediction mode of the current block is the default mode can be transmitted by a signal via a bitstream. The information is a 1-bit flag, and the flag may be referred to as a default mode flag. The MPM flag can transmit the default mode flag by a signal only when the MPM that is the same as the current block is included in the MPM list. As described above, the default mode may include at least one of a planar, DC, vertical direction mode, or horizontal direction mode. For example, when the planar is the default mode, the default mode flag can indicate whether the intra prediction mode of the current block is planar. When the default mode flag indicates that the intra prediction mode of the current block is not the default mode, one of the MPMs indicated by the index information can be set as the intra prediction mode of the current block.
[0244] When multiple intra prediction modes are set as the default mode, index information indicating any one of the default modes can be further transmitted by a signal. The intra prediction mode of the current block can be set to the default mode indicated by the index information.
[0245] When the index of the reference sample line of the current block is not 0, it is set not to use the default mode. Therefore, when the index of the reference sample line is not 0, the default mode flag can be set to a predefined value (i.e., false) without transmitting the default mode flag by a signal.
[0246] When the intra prediction mode of the current block is determined, prediction samples related to the current block can be obtained based on the determined intra prediction mode (S2203).
[0247] When the DC mode is selected, prediction samples related to the current block can be generated based on the average value of the reference samples. Specifically, based on the average value of the reference samples, the values of all samples in the prediction block can be generated. The average value can be derived using at least one of the upper reference samples located above the current block and the left reference samples located to the left of the current block.
[0248] The number or range of reference samples for deriving the average value may vary depending on the shape of the current block. For example, if the current block is a non-square block with a width larger than its height, the average value can be calculated using only the upper reference samples. On the other hand, if the current block is a non-square block with a width smaller than its height, the average value can be calculated using only the left reference samples. That is, when the width and height of the current block are different, the average value can be calculated using only the reference samples adjacent to the longer length. Alternatively, based on the ratio of the width to the height of the current block, it can be determined whether to calculate the average value using only the upper reference samples or only the left reference samples.
[0249] When the planar mode is selected, prediction samples can be obtained using horizontal prediction samples and vertical prediction samples. Here, horizontal prediction samples are obtained based on left and right reference samples located on a horizontal line that is the same as the prediction sample, and vertical prediction samples are obtained based on upper and lower reference samples located on a vertical line that is the same as the prediction sample. Here, the right reference sample can be generated by replicating the reference sample adjacent to the upper right corner of the current block, and the lower reference sample can be generated by replicating the reference sample adjacent to the lower left corner of the current block. Based on the weighted addition of the left and right reference samples, horizontal prediction samples can be obtained, and based on the weighted addition of the upper and lower reference samples, vertical prediction samples can be obtained. In this case, based on the position of the prediction sample, the weighted value assigned to each reference sample can be determined. Based on the average operation or weighted addition of the horizontal prediction sample and the vertical prediction sample, a prediction sample can be obtained. When performing weighted addition, based on the position of the prediction sample, the weighted values assigned to the horizontal prediction sample and the vertical prediction sample can be determined.
[0250] When the angular prediction mode is selected, a parameter representing the prediction direction (or prediction angle) of the selected angular prediction mode can be determined. Table 2 below shows the intra prediction parameter intraPredAng for each intra prediction mode.
Table 2
[0251] Table 2 shows the intra direction parameters of each intra prediction mode having any one index from 2 to 34 when 35 intra prediction modes are defined. When more than 33 angular intra prediction modes are defined, Table 2 further details the intra direction parameters of each angular intra prediction mode.
[0252] After arranging the upper reference sample and the left reference sample of the current block in a row, a prediction sample can be obtained based on the value of the intra-direction parameter. In this case, when the value of the intra-direction parameter is negative, the left reference sample and the upper reference sample can be arranged in a row.
[0253] FIG. 25 and FIG. 26 are diagrams showing examples of one-dimensional arrays in which reference samples are arranged in a row.
[0254] FIG. 25 shows an example of a one-dimensional vertical array in which reference samples are arranged in the vertical direction, and FIG. 26 shows an example of a one-dimensional horizontal array in which reference samples are arranged in the horizontal direction. Assuming that 35 intra prediction modes are defined, the examples of FIGS. 25 and 26 will be described.
[0255] When the intra prediction mode index is any one of 11 to 18, a one-dimensional horizontal array in which the upper reference sample is rotated counterclockwise can be applied. When the intra prediction mode index is any one of 19 to 25, a one-dimensional vertical array in which the left reference sample is rotated clockwise can be applied. When the reference samples are arranged in a row, the intra prediction mode angle can be considered.
[0256] Based on the intra-direction parameter, a reference sample determination parameter can be determined. The reference sample determination parameter may include a reference sample index for specifying a reference sample and a weight value parameter for determining a weight value applied to the reference sample.
[0257] The reference sample index iIdx and the weight value parameter ifact are obtained by the following equations 17 and 18, respectively.
Equation
Equation
[0258] In Equations 17 and 18, P ang represents an intra-direction parameter. The reference sample specified by the reference sample index iIdx corresponds to an integer pel.
[0259] To derive a prediction sample, one or more reference samples can be specified. Specifically, considering the gradient of the prediction mode, the position of the reference sample for deriving the prediction sample can be specified. For example, using the reference sample index iIdx, the reference sample for deriving the prediction sample can be specified.
[0260] In this case, when the gradient of the intra prediction mode cannot be represented by one reference sample, the prediction sample can be generated by performing interpolation on a plurality of reference samples. For example, when the gradient of the intra prediction mode is a value between the gradient between the prediction sample and the first reference sample and the gradient between the prediction sample and the second reference sample, the prediction sample can be obtained by performing interpolation on the first reference sample and the second reference sample. That is, when the angular line following the intra prediction angle does not pass through the reference sample located at the integer pixel, the prediction sample can be obtained by performing interpolation on the reference samples adjacent to the left or right or above or below the position where the angular line has passed.
[0261] The following Equation 19 shows an example of obtaining a prediction sample based on a reference sample.
Equation
[0262] In Equation 19, P represents a prediction sample, and Ref_1D represents any one of the reference samples in the one-dimensional array. In this case, based on the position (x, y) of the prediction sample and the reference sample index iIdx, the position of the reference sample can be determined.
[0263] When the gradient of the intra prediction mode is represented as one reference sample, the weighted value parameter ifact can be set to 0. Therefore, Equation 19 can be simplified as Equation 20 below.
Number
[0264] Intra prediction can also be performed on the current block based on multiple intra prediction modes. For example, an intra prediction mode can be derived for different prediction samples, and a prediction sample can be derived based on the intra prediction mode assigned to each prediction sample.
[0265] Alternatively, an intra prediction mode can be derived for different regions, and intra prediction can be performed on each region based on the intra prediction mode assigned to each region. Here, the region may include at least one sample. Based on at least one of the size, shape, or intra prediction mode of the current block, at least one of the size or shape of the region can be adaptively determined. Alternatively, in the encoder and decoder, at least one of the size or shape of the region can be predefined regardless of the size or shape of the current block.
[0266] Alternatively, intra prediction can be executed based on a plurality of intra predictions respectively, and a final prediction sample can be derived based on the average operation or weighted addition of a plurality of prediction samples obtained by performing the intra prediction a plurality of times. For example, by performing intra prediction based on the first intra prediction mode, a first prediction sample can be obtained, and by performing intra prediction based on the second intra prediction mode, a second prediction sample can be obtained. Subsequently, a final prediction sample can be obtained based on the average operation or weighted addition of the first prediction sample and the second prediction sample. In this case, by considering whether the first intra prediction mode is a non-angular / angular prediction mode, whether the second intra prediction mode is a non-angular / angular prediction mode, or at least one of the intra prediction modes of adjacent blocks, weighted values respectively assigned to the first prediction sample and the second prediction sample can be determined.
[0267] The plurality of intra prediction modes may be a combination of a non-angular intra prediction mode and an angular prediction mode, a combination of angular prediction modes, or a combination of non-angular prediction modes.
[0268] FIG. 27 is a diagram showing an angle formed between an angular intra prediction mode and a straight line perpendicular to the x-axis.
[0269] In the example shown in FIG. 27, the angular prediction mode may exist between the lower left diagonal direction and the upper right diagonal direction. When described as the angle formed by the x-axis and the angular prediction mode, the angular prediction mode may exist between 45 degrees (lower left diagonal direction) and -135 degrees (upper right diagonal direction).
[0270] When the current block has a non-square shape, based on the intra prediction mode of the current block, without using the reference sample closest to the prediction sample among the reference samples located on the angular line following the intra prediction angle, the reference sample farthest from the prediction sample is used to derive the prediction sample.
[0271] FIG. 28 is a diagram showing an example of obtaining a prediction sample when the current block is non-square.
[0272] For example, in the example shown in FIG. 28(a), assume that the current block has a non-square shape where the width is greater than the height, and the intra prediction mode of the current block is an angular intra prediction mode with an angle ranging from 0 degrees to 45 degrees. In this case, when deriving the prediction sample A in the vicinity of the right column of the current block, the left reference sample L among the reference samples in the angular mode of the angle, which is far from the prediction sample, may be used instead of the upper reference sample T close to the prediction sample.
[0273] In another example, in the example shown in FIG. 28(b), assume that the current block has a non-square shape where the height is greater than the width, and the intra prediction mode of the current block is an angular intra prediction mode with an angle ranging from -90 degrees to -135 degrees. In the above-described case, when deriving the prediction sample A in the vicinity of the lower row of the current block, the upper reference sample T among the reference samples in the angular mode of the angle, which is far from the prediction sample, may be used instead of the left reference sample L close to the prediction sample.
[0274] To solve the above-described problem, when the current block is non-square, the intra prediction mode of the current block can be replaced with a reverse intra prediction mode. Therefore, for a non-square shaped block, an angular prediction mode having an angle greater than or smaller than the angle of the angular prediction mode shown in FIG. 24 can be used. Such an angular intra prediction mode may be defined as a wide-angle intra prediction mode. The wide-angle intra prediction mode represents an angular intra prediction mode that does not fall within the range of 45 degrees to -135 degrees.
[0275] FIG. 29 is a diagram showing the wide-angle intra prediction mode.
[0276] In the example shown in FIG. 29, the intra prediction modes with indices from -1 to -14 and the intra prediction modes with indices from 67 to 80 represent wide-angle intra prediction modes.
[0277] FIG. 29 shows 14 wide-angle intra prediction modes (-1 to -14) having angles greater than 45 degrees and 14 wide-angle intra prediction modes (67 to 80) having angles less than -135 degrees, but a greater or smaller number of wide-angle intra prediction modes can be defined.
[0278] When using the wide-angle intra prediction mode, the length of the upper reference sample is set to 2W + 1, and the length of the left reference sample is set to 2H + 1.
[0279] When using the wide-angle intra prediction mode, sample A shown in FIG. 28(a) can be predicted using reference sample T, and sample A shown in FIG. 28(b) can be predicted using reference sample L.
[0280] By adding the existing intra prediction modes and N wide-angle intra prediction modes, a total of 67 + N intra prediction modes can be used. For example, Table 3 shows the intra-direction parameters of the intra prediction modes when 20 wide-angle intra prediction modes are defined.
Table 3
[0281] When the current block is non-square in shape and the intra prediction mode of the current block obtained in step S2202 falls within the conversion range, the intra prediction mode of the current block can be converted to a wide-angle intra prediction mode. The conversion range can be determined based on at least one of the size, shape, or ratio of the current block. Here, the ratio can represent the ratio of the width to the height of the current block.
[0282] When the current block is a non-square with a width greater than its height, the conversion range can be set from the upper-right diagonal intra prediction mode index (e.g., 66) to (the index of the upper-right diagonal intra prediction mode - N). Here, N may be determined based on the ratio of the current block. If the intra prediction mode of the current block falls within the conversion range, the intra prediction mode can be converted to a wide-angle intra prediction mode. The conversion can be performed by subtracting a predefined value from the intra prediction mode. The predefined value may be the total number of intra prediction modes other than the wide-angle intra prediction mode (e.g., 67).
[0283] According to the above embodiment, the intra prediction modes between the 66th and the 53rd can be respectively converted to the wide-angle intra prediction modes between the -1st and the -14th.
[0284] When the current block is a non-square with a height greater than its width, the conversion range can be set from the lower-left diagonal intra prediction mode index (e.g., 2) to (the index of the lower-left diagonal intra prediction mode + M). Here, M may be determined based on the ratio of the current block. If the intra prediction mode of the current block falls within the conversion range, the intra prediction mode can be converted to a wide-angle intra prediction mode. The conversion can be performed by adding a predefined value to the intra prediction mode. The predefined value may be the total number of angular intra prediction modes other than the wide-angle intra prediction mode (e.g., 65).
[0285] According to the above embodiment, the intra prediction modes between the 2nd and the 15th can be respectively converted to the wide-angle intra prediction modes between the 67th and the 80th.
[0286] Hereinafter, the intra prediction mode that falls within the conversion range is referred to as a wide-angle intra replacement prediction mode.
[0287] The conversion range may be determined based on the ratio of the current block. For example, Tables 4 and 5 respectively show the conversion ranges when defining 35 intra prediction modes other than the wide-angle intra prediction mode and 67 intra prediction modes. [Table 4] [Table 5]
[0288] As in the examples shown in Tables 4 and 5, the number of wide-angle intra replacement prediction modes within the conversion range may vary depending on the ratio of the current block.
[0289] With the further use of the wide-angle intra prediction mode in addition to the existing intra prediction modes, the resources required for encoding the wide-angle intra prediction mode increase. Therefore, it may reduce the encoding efficiency. Therefore, instead of directly encoding the wide-angle intra prediction mode, encoding is performed for the replacement intra prediction mode related to the wide-angle intra prediction mode to improve the encoding efficiency.
[0290] For example, when encoding the current block using the 67th wide-angle intra prediction mode, the number 2 of the 67th wide-angle replacement intra prediction mode can be encoded as the intra prediction mode of the current block. Note that when encoding the current block using the -1st wide-angle intra prediction mode, the number 66 of the -1st wide-angle replacement intra prediction mode can be encoded as the intra prediction mode of the current block.
[0291] The decoder can perform decoding on the intra prediction mode of the current block and determine whether the decoded intra prediction mode is included in the conversion range. If the decoded intra prediction mode is the wide-angle replacement intra prediction mode, the intra prediction mode can be converted to the wide-angle intra prediction mode.
[0292] Alternatively, when encoding the current block in the wide-angle intra prediction mode, encoding can be directly performed for the wide-angle intra prediction mode.
[0293] The encoding of the intra prediction mode may be realized based on the MPM list. In the following, the method for setting the MPM list will be described in detail. In the embodiments described later, it is assumed that ten wide-angle intra prediction modes (-1 to -10) with an angle greater than 45 degrees and ten wide-angle intra prediction modes (67 to 76) with an angle less than -135 degrees are defined.
[0294] When encoding an adjacent block in the wide-angle intra prediction mode, the MPM can be set based on the wide-angle replacement intra prediction mode corresponding to the wide-angle intra prediction mode. For example, when encoding an adjacent block in the wide-angle intra prediction mode, the variable candIntraPredX (where X is A or B) can be set to the wide-angle replacement intra prediction mode.
[0295] Alternatively, based on the shape of the current block, the method for deriving the MPM can be determined. For example, when the current block is a square shape with the same width and height, candIntraPredX can be set to the wide-angle replacement intra prediction mode. On the other hand, when the current block is a non-square shape, candIntraPredX can be set to the wide-angle intra prediction mode.
[0296] Alternatively, it is possible to determine whether to set candIntraPredX to the wide-angle intra prediction mode based on whether the wide-angle intra prediction mode of an adjacent block can be applied to the current block. For example, if the current block has a non-square shape where the width is greater than the height, the wide-angle intra prediction mode with an index greater than the index of the intra prediction mode in the upper right diagonal direction is directly set as candIntraPredX. For a wide-angle intra prediction mode with an index less than the index of the intra prediction mode in the lower left diagonal direction, the wide-angle replacement intra prediction mode corresponding to the wide-angle intra prediction mode is set as candIntraPredX. On the other hand, if the current block has a non-square shape where the height is greater than the width, the wide-angle intra prediction mode with an index less than the index of the intra prediction mode in the lower left diagonal direction is directly set as candIntraPredX. For a wide-angle intra prediction mode with an index greater than the index of the intra prediction mode in the upper right diagonal direction, the wide-angle replacement intra prediction mode corresponding to the wide-angle intra prediction mode is set as candIntraPredX.
[0297] That is, based on whether the shape of an adjacent block encoded in the wide-angle intra prediction mode is the same as or similar to the shape of the current block, it is possible to determine whether to derive the MPM using the wide-angle intra prediction mode or whether to derive the MPM using the wide-angle replacement intra prediction mode.
[0298] Alternatively, regardless of the shape of the current block, the wide-angle intra prediction mode of an adjacent block can be set as candIntraPredX.
[0299] In short, candIntraPredX can be set to the wide-angle intra prediction mode or the wide-angle replacement intra prediction mode of an adjacent block.
[0300] The MPM can be derived based on candIntraPredA and candIntraPredB. In this case, the MPM can be derived in an intra prediction mode similar to candIntraPredA or candIntraPredB. Based on modulo operation and offset, an intra prediction mode similar to candIntraPredA or candIntraPredB can be derived. In this case, depending on the shape of the current block, the constants and offsets used in the modulo operation can be determined to be different.
[0301] Table 6 shows an example of deriving the MPM based on the shape of the current block.
Table 6
[0302] Assume that candIntraPredA and candIntraPredB are the same and candIntraPredA is in an angular intra prediction mode. If the current block is square-shaped, an intra prediction mode similar to candIntraPredA can be obtained by a modulo operation based on the value obtained by subtracting 1 from the total number of angular intra prediction modes other than the wide-angle intra prediction mode. For example, if the number of angular intra prediction modes other than the wide-angle intra prediction mode is 65, the MPM can be derived by the value obtained by the modulo operation based on candIntraPredA and 64. On the other hand, if the current block is non-square-shaped, an intra prediction mode similar to candIntraPredA can be obtained by a modulo operation based on the value obtained by subtracting 1 from the total number of angular intra prediction modes including the wide-angle intra prediction mode. For example, if the number of wide-angle intra prediction modes is 20, the MPM can be derived by the value obtained by the modulo operation based on candIntrapredA and 84.
[0303] Since the constants used in modulo operations are set to vary depending on the shape of the current block, it is possible to determine whether the wide-angle intra prediction mode can be set to an angular intra prediction mode similar to candIntraPredA. For example, in modulo operations using 64, it may not be possible to set the wide-angle intra prediction mode to an angular intra prediction mode similar to candIntraPredA, but in modulo operations using 84, it is possible to set the wide-angle intra prediction mode to an angular intra prediction mode similar to candIntraPredA.
[0304] Alternatively, when candIntraPredA and candIntraPredB are the same, the MPM can be derived considering the shape of the current block and whether candIntraPredA is a wide-angle intra prediction mode.
[0305] Table 7 shows an example of deriving the MPM based on the shape of the current block.
Table 7
[0306] JPEG0007693872000028.jpg109150
[0307] Assume that candIntraPredA and candIntraPredB are the same.
[0308] When the current block is square-shaped and candIntraPredA is a wide-angle intra prediction mode, the MPM can be set to the default mode. For example, MPM[0], MPM[1], and MPM[2] can be set to the planar mode, the DC mode, and the vertical intra prediction mode, respectively.
[0309] When the current block is square-shaped and candIntraPredA is an angular intra prediction mode other than the wide-angle intra prediction mode, the MPM can be set to an angular intra prediction mode similar to candIntraPredA. For example, MPM[0] can be set to candIntraPredA, and MPM[1] and MPM[2] can be set to angular intra prediction modes similar to candIntraPredA.
[0310] When the current block is non-square-shaped and candIntraPredA is an angular intra prediction mode, the MPM can be set to an angular intra prediction mode similar to candIntraPredA. For example, MPM[0] can be set to candIntraPredA, and MPM[1] and MPM[2] can be set to angular intra prediction modes similar to candIntrapredA.
[0311] An angular intra prediction mode similar to candIntraPredA can be derived using modulo operation and offset. In this case, the constant used in the modulo operation may vary depending on the shape of the current block. Note that based on the shape of the current block, the offset for deriving an angular intra prediction mode similar to candIntraPredA can be set differently. For example, when the current block is a non-square shape with a width larger than the height, an angular intra prediction mode similar to candIntraPredA can be derived by using offset 2. On the other hand, when the current block is a non-square shape with a height larger than the width, an angular intra prediction mode similar to candIntraPredA can be derived by using offset 2 and -8.
[0312] Alternatively, the MPM can be derived by considering whether candIntraPredX is a wide-angle intra prediction mode with the maximum index or a wide-angle intra prediction mode with the minimum index.
[0313] Table 8 shows an example of deriving the MPM by considering the wide-angle intra prediction mode index. [Table 8]
[0314] JPEG0007693872000030.jpg161150
[0315] Assume that candIntraPredA and candIntraPredB are the same. For ease of explanation, a wide-angle intra prediction mode whose index value is less than the index value of the intra prediction mode in the lower left diagonal direction is called a downward wide-angle intra prediction mode, and a wide-angle intra prediction mode whose index value is greater than the index value of the intra prediction mode in the upper right diagonal direction is called a rightward wide-angle intra prediction mode.
[0316] When candIntraPredA is a downward wide-angle intra prediction mode, the MPM can be set to an angular intra prediction mode similar to candIntraPredA. In this case, when candIntraPredA is a downward wide-angle intra prediction mode having the minimum value, the MPM can be set to a downward wide-angle intra prediction mode having a pre-defined index value. Here, the pre-defined index may be an index having the maximum value among the indexes of the downward wide-angle intra prediction modes. For example, when candIntraPredA is -10, MPM[0], MPM[1], and MPM[2] can be set to -10, -1, and -9, respectively.
[0317] When candIntraPredA is a right - wide - angle intra - prediction mode, the MPM can be set to an intra - prediction mode with an angle similar to candIntraPredA. In this case, when candIntraPredA is the right - wide - angle intra - prediction mode with the maximum value, the MPM can be set to a right - wide - angle intra - prediction mode with a predefined index value. Here, the predefined index may be an index having the minimum value among the indexes of the right - wide - angle intra - prediction modes. For example, when candIntraPredA is 77, MPM[0], MPM[1], and MPM[2] can be set to 77, 76, and 67, respectively.
[0318] Alternatively, if the index obtained by subtracting 1 from the index of candIntraPredA is less than the minimum value among the indexes of the intra - prediction modes, or if the index obtained by adding 1 to the index of candIntraPredA is greater than the maximum value, the MPM can be set to the default mode. Here, the default mode may include at least one of the planar mode, the DC mode, the vertical intra - prediction mode, the horizontal intra - prediction mode, and the diagonal intra - prediction mode.
[0319] Alternatively, if the index obtained by subtracting 1 from the index of candIntraPredA is less than the minimum value among the indexes of the intra - prediction modes, or if the index obtained by adding 1 to the index of candIntraPredA is greater than the maximum value, the MPM can be set to an intra - prediction mode opposite to candIntraPredA or an intra - prediction mode similar to the intra - prediction mode opposite to candIntraPredA.
[0320] Alternatively, the MPM candidates can be derived considering the shape of the current block and the shape of the adjacent blocks. For example, the method of deriving the MPM when both the current block and the adjacent block are non - square shapes may be different from the method of deriving the MPM when the current block is square but the adjacent block is non - square.
[0321] At least one of the current block size, the current block shape, the adjacent block size, and the adjacent block shape can be considered to rearrange (or reorder) the MPMs in the MPM list. Here, rearrangement means reassigning the index assigned to each MPM. For example, a small index can be assigned to an MPM that is the same as the intra prediction mode of an adjacent block having the same size or shape as the current block size or shape.
[0322] Assume that MPM[0] and MPM[1] are the intra prediction modes candIntraPredA of the left adjacent block and candIntraPredB of the upper adjacent block, respectively.
[0323] If the current block and the upper adjacent block are non-square shapes where the width is greater than the height, the MPMs can be rearranged so that the intra prediction mode candIntraPredB of the upper adjacent block has a small index. That is, candIntraPredB can be rearranged to MPM[0], and candIntraPredA can be rearranged to MPM[1].
[0324] Alternatively, if the current block and the upper adjacent block are non-square shapes where the height is greater than the width, the MPMs can be rearranged so that the intra prediction mode candIntraPredB of the upper adjacent block has a small index. That is, candIntraPredB can be rearranged to MPM[0], and candIntraPredA can be rearranged to MPM[1].
[0325] Alternatively, when the current block and the upper adjacent block are square-shaped, the MPM can be rearranged such that the upper adjacent block's intra prediction mode candIntraPredB has a small index. That is, candIntraPredB can be rearranged to MPM[0], and candIntraPredA can be rearranged to MPM[1].
[0326] Instead of rearranging the MPM, when candIntraPredX is first assigned to the MPM, at least one of the size of the current block, the shape of the current block, the size of the adjacent block, and the shape of the adjacent block can be considered.
[0327] The MPM can be rearranged based on the size or shape of the current block. For example, if the current block is a non-square shape where the width is larger than the height, the MPM can be rearranged in descending order. On the other hand, if the current block is a non-square shape where the height is larger than the width, the MPM can be rearranged in ascending order.
[0328] By subtracting the original image from the predicted image, a derived residual image can be obtained. In this case, when the residual image is changed to the frequency domain, even if the high-frequency components among the frequency components are removed, the subjective video quality will not be significantly degraded. Therefore, converting the values of the high-frequency components to small values or setting the values of the high-frequency components to 0 has the effect of improving the compression efficiency without causing obvious visual distortion. To reflect the above characteristics, the residual image can be decomposed into two-dimensional frequency components by converting the current block. The above conversion can be performed using a conversion technique such as the Discrete Cosine Transform (DCT) or the Discrete Sine Transform (DST).
[0329] The DCT decomposes (or transforms) a residual image into two-dimensional frequency components using a cosine transform. The DST decomposes (or transforms) a residual image into two-dimensional frequency components using a sine transform. As a conversion result of the residual image, the frequency components may be represented as a basic image. For example, when performing a DCT transform on a block of size N×N, N 2 basic pattern components can be obtained. By the transformation, the size of each basic pattern component included in a block of size N×N can be obtained. According to the transformation technique used, the size of the basic pattern component may be referred to as a DCT coefficient or a DST coefficient.
[0330] The transformation technique DCT is mainly used to perform a transformation on an image in which many non-zero low-frequency components are distributed. The transformation technique DST is mainly used for an image in which many high-frequency components are distributed.
[0331] It is also possible to transform the residual image using a transformation technique other than DCT or DST.
[0332] Hereinafter, the process of converting a residual image into two-dimensional frequency components is referred to as two-dimensional image conversion. Note that the size of the basic pattern component obtained by the conversion may also be referred to as a conversion coefficient. For example, the conversion coefficient may be a DCT coefficient or a DST coefficient. When the main conversion and the secondary conversion described later are applied simultaneously, the conversion coefficient can represent the size of the basic pattern component generated by the result of the secondary conversion.
[0333] The transformation technique can be determined in units of blocks. The transformation technique can be determined based on at least one of the prediction coding mode of the current block, the size of the current block, or the shape of the current block. For example, in the intra prediction mode, when encoding is performed on the current block and the size of the current block is less than N×N, DST can be executed using the transformation technique. On the other hand, when the above conditions cannot be satisfied, the transformation can be executed using the transformation technique DCT.
[0334] In the residual image, it may not be necessary to perform a two-dimensional image transformation on some blocks. Not performing the two-dimensional image transformation may be referred to as transform skip. When applying transform skip, quantization can be applied to the residual values for which the transformation has not been performed.
[0335] After transforming the current block using DCT or DST, the transformed current block can be transformed again. In this case, the transformation based on DCT or DST is defined as the main transformation, and the process of performing a transformation again on the block to which the main transformation is applied is called the secondary transformation.
[0336] The main transformation can be executed using any one of a plurality of transformation kernel candidates. For example, the main transformation can be executed using any one of DCT2, DCT8, or DCT7.
[0337] For the horizontal and vertical directions, different transformation kernels can also be used. Information representing the combination of the horizontal transformation kernel and the vertical transformation kernel can also be transmitted as a signal via the bitstream.
[0338] The execution units of the main transformation and the secondary transformation are different. For example, the main transformation can be executed on an 8×8 block, and the secondary transformation can be executed on a sub-block with a size of 4×4 in the transformed 8×8 block. In this case, the transformation coefficients in the surplus area where the secondary transformation is not performed can also be set to 0.
[0339] Alternatively, the main transformation can be executed on a 4×4 block, and the secondary transformation can be executed on an area with a size of 8×8 including the transformed 4×4 block.
[0340] Information indicating whether to perform the secondary transformation can be transmitted as a signal via the bitstream.
[0341] In a decoder, an inverse transformation of the second transformation (second inverse transformation) can be performed, and an inverse transformation of the main transformation (first inverse transformation) can be performed on the result. As a result of the execution of the second inverse transformation and the first inverse transformation, a residual signal of the current block can be obtained.
[0342] Quantization is used to reduce the energy of a block, and the quantization process includes a process of dividing a transform coefficient by a specific constant. The constant may be derived from a quantization parameter, and the quantization parameter may be defined as a value from 1 to 63.
[0343] When conversion and quantization are performed in an encoder, the decoder can obtain a residual block by inverse quantization and inverse transformation. The decoder can obtain a reconstructed block of the current block by adding a prediction block and the residual block.
[0344] When a reconstructed block of the current block is obtained, information loss generated in the quantization and encoding processes can be reduced by in-loop filtering. The in-loop filter may include at least one of a deblocking filter, a sample adaptive offset filter (SAO), or an adaptive loop filter (ALF). Hereinafter, the reconstructed block before applying the in-loop filter may be referred to as a first reconstructed block, and the reconstructed block after applying the in-loop filter may be referred to as a second reconstructed block.
[0345] A second reconstructed block can be obtained by applying at least one of a deblocking filter, SAO, or ALF to the first reconstructed block. In this case, SAO or ALF can be applied after applying the deblocking filter.
[0346] A deblocking filter is used to reduce image quality degradation (blocking artifact) that occurs at block boundaries when quantization is performed in units of blocks. To apply the deblocking filter, the blocking strength (BS) between the first reconstructed block and the adjacent reconstructed block can be determined.
[0347] FIG. 30 is a flowchart showing a process for determining the blocking strength.
[0348] In the example shown in FIG. 30, P represents the first reconstructed block, and Q represents the adjacent reconstructed block. Here, the adjacent reconstructed block may be adjacent to the left or above the current block.
[0349] In the example shown in FIG. 30, the blocking strength can be determined by considering whether the prediction coding modes of P and Q, whether they include non-zero transform coefficients, whether inter prediction is performed using the same reference image, and whether the difference value of the motion vectors is greater than or equal to a threshold.
[0350] Based on the blocking strength, it can be determined whether the deblocking filter has been applied. For example, if the blocking strength is 0, filtering may not be performed.
[0351] SAO is used to reduce the ringing artifact that occurs when quantization is performed in the frequency domain. SAO can be performed by adding or subtracting a pattern considering the first reconstructed image to an offset determined by adding or subtracting. The method for determining the offset includes an edge offset (EO) or a band offset. EO represents a method for determining the offset of the current sample based on the pattern of surrounding pixels. BO represents a method for applying a common offset to a set of pixels having similar luminance values in a region. Specifically, pixel luminance can be divided into 32 equal intervals, and pixels having similar luminance values can be grouped into one set. For example, four adjacent bands out of 32 bands can be grouped together, and the same offset can be applied to the samples belonging to the four bands.
[0352] ALF is a method for generating a second reconstructed image by applying a filter of a predefined size / shape to the first reconstructed image or the reconstructed image to which a deblocking filter has been applied. The following Equation 21 shows an application example of ALF.
Equation
[0353] One of the predefined filter candidates can be selected in units of an image, a coding tree unit, a coding block, a prediction block, or a transform block. Any one of the sizes or shapes of each filter candidate may be different.
[0354] Figure 31 is a diagram showing predefined filter candidates.
[0355] In the example shown in Figure 31, at least one of 5×5, 7×7, and 9×9 rhombuses can be selected.
[0356] For chromaticity components, only rhombuses with a size of 5×5 can be used.
[0357] For high-resolution videos such as panoramic videos, 360-degree videos, or 4K / 8K UHD (ultra-high definition), in order to perform real-time encoding or low-latency encoding, one image can be divided into a plurality of regions, and encoding / decoding can be performed on the plurality of regions. For this purpose, the image can be divided into tiles (i.e., basic units encoded / decoded in parallel), and these tiles can be processed in parallel.
[0358] The tiles can be restricted to those having a rectangular shape. When performing encoding / decoding on a tile, data of other tiles is not used. Taking the tile as a unit, the probability table of the context-adaptive binary arithmetic coding (CABAC) context can be initialized, and it can be set so that the loop filter is not applied at the tile boundary.
[0359] FIG. 32 shows an example of dividing an image into a plurality of tiles.
[0360] A tile includes at least one coding tree unit, and the tile boundary overlaps with the coding tree unit boundary.
[0361] In the example shown in FIG. 32, the image may be divided into a plurality of tile sets. Information for dividing the image into a plurality of tile sets can be transmitted as a signal via a bitstream.
[0362] According to the division type of the image, the tiles may have the same size in all regions other than the image boundary.
[0363] Alternatively, the image can be divided such that horizontally adjacent tiles have the same height, or the image can be divided such that vertically adjacent tiles have the same width.
[0364] When dividing an image using at least one line, either a vertical or a horizontal line, that intersects the image, each tile belongs to a different column and / or row. In the exemplary embodiments described below, the column to which a tile belongs is referred to as a tile column, and the row to which a tile belongs is referred to as a tile row.
[0365] Via a bitstream, information for determining the shape of dividing an image into tiles can be transmitted in a signal. The information can be encoded and transmitted in a signal by an image parameter set or a sequence parameter set. The information is used to determine the number of tiles in the image and may also include information indicating the number of tile rows and information indicating the number of tile columns. For example, the syntax element num_tile_columns_minus1 indicates a value obtained by subtracting 1 from the number of tile columns, and the syntax element num_tile_rows_minus1 indicates a value obtained by subtracting 1 from the number of tile rows.
[0366] In the example shown in FIG. 32, since the number of tile columns is 4 and the number of tile rows is 3, num_tile_columns_minus1 may be 3, and num_tile_rows_minus1 may be 2.
[0367] When dividing an image into a plurality of tiles, information indicating the size of the tiles can be transmitted in a signal via a bitstream. For example, when dividing an image into a plurality of tile columns, information indicating the width of each tile column is transmitted in a signal via a bitstream. Also, when dividing an image into a plurality of tile rows, information indicating the height of each tile row is transmitted in a signal via a bitstream. For example, for each tile column, the syntax element column_width_minus1 indicating the width of the tile column can be encoded and transmitted in a signal. Also, for each tile row, the syntax element row_height_minus1 indicating the height of the tile row can be encoded and transmitted in a signal.
[0368] column_width_minus1 can indicate a value obtained by subtracting 1 from the width of a tile column. Also, row_height_minus1 can indicate a value obtained by subtracting 1 from the height of a tile row.
[0369] For the last tile column, encoding for column_width_minus1 may be omitted, and for the last tile row, encoding for row_height_minus1 may be omitted. The width of the last tile column and the height of the last row can be derived considering the size of the image.
[0370] The decoder can determine the size of a tile based on column_width_minus1 and row_height_minus1.
[0371] Table 9 shows a syntax table for dividing an image into tiles.
Table 9
[0372] Referring to Table 9, it is possible to transmit, in a signal, a syntax element num_tile_columns_minus1 indicating the number of tile columns and a syntax element num_tile_rows_minus1 indicating the number of tile rows.
[0373] Subsequently, it is possible to transmit, in a signal, a syntax element uniform_spacing_flag indicating whether the image is divided into tiles of equal size. If uniform_spacing_flag is true, tiles in regions other than the image boundary can be divided into tiles of equal size.
[0374] When the uniform_spacing_flag is false, a signal can be used to transmit a syntax element column_width_minus1 indicating the width of each tile column and a syntax element row_height_minus1 indicating the height of each tile row.
[0375] The syntax element loop_filter_across_tiles_enabled_flag indicates whether to allow the application of the loop filter across tile boundaries.
[0376] The tile column with the minimum width among the tile columns may be called the minimum-width tile, and the tile row with the minimum height among the tile rows may be called the minimum-height tile. Information indicating the width of the minimum-width tile and the height of the minimum-height tile can be transmitted as a signal via the bitstream. For example, the syntax element min_column_width_minus1 indicates a value obtained by subtracting 1 from the width of the minimum-width tile, and the syntax element min_row_height_minus1 indicates a value obtained by subtracting 1 from the height of the minimum-height tile.
[0377] For each tile column, information indicating the difference value from the minimum tile width can be transmitted as a signal. For example, the syntax element diff_column_width indicates the width difference value between the current tile column and the minimum tile column. The width difference value may be represented as the difference value of the number of coded tree unit columns. The decoder can derive the width of the current tile by adding the width of the minimum-width tile derived based on min_column_width_minus1 and the width difference value derived based on diff_column_width.
[0378] In addition, for each tile row, information indicating the difference value from the minimum tile height can be transmitted by a signal. For example, the syntax element diff_row_height indicates the height difference value between the current tile row and the minimum tile row. The height difference value may be represented as the difference value of the number of rows of the coding tree unit. The decoder can derive the height of the current tile by adding the height of the minimum height tile derived based on min_row_height_minus1 and the height difference value derived based on diff_row_height.
[0379] Table 10 shows a syntax table including information related to size differences.
Table 10
[0380] An image can be divided so as to have a height different from that of horizontally adjacent tiles, or an image can be divided so as to have a width different from that of vertically adjacent tiles. The above image division method may be called a Flexible Tile division method, and tiles divided by the Flexible Tile division method may be called Flexible Tiles.
[0381] FIG. 33 is a diagram showing an image division mode by Flexible Tile technology.
[0382] The search order of tiles generated by dividing an image may follow a predetermined scanning order. Note that an index can be assigned to each tile according to the predetermined scanning order.
[0383] The scanning order of tiles may be any one of raster scanning, diagonal scanning, vertical scanning, or horizontal scanning. FIGS. 33(a) to 33(d) show examples of assigning an index to each tile based on raster scanning, diagonal scanning, vertical scanning, and horizontal scanning, respectively.
[0384] The next scanning order can be determined based on the current tile size or position. For example, when the height of the current tile is different from the height of the tile adjacent to the right of the current tile (for example, when the height of the right adjacent tile is greater than the height of the current tile), the leftmost tile among the tiles on the vertical line that is the same as the vertical line of the tile adjacent to the bottom of the current tile may be determined as the next tile to be scanned after the current tile.
[0385] The scanning order of tiles can be determined in units of images or sequences.
[0386] Alternatively, by considering the size of the first tile in the image, the scanning order of the tiles can be determined. For example, when the width of the first tile is greater than the height, the scanning order of the tiles can be a horizontal scan. When the height of the first tile is greater than the width, the scanning order of the tiles can be a vertical scan. When the width of the first tile is the same as the height, the scanning order of the tiles can be a raster scan or a diagonal scan.
[0387] Information indicating the total number of tiles can be transmitted as a signal via a bitstream. For example, when applying flexible tile technology, a syntax element number_of_tiles_in_picture_minus2 derived by subtracting 2 from the total number of tiles in the image can be transmitted as a signal. The decoder can recognize the number of tiles included in the current image based on number_of_tiles_in_picture_minus2.
[0388] Table 11 shows a syntax table containing information related to the number of tiles.
Table 11
[0389] To reduce the number of bits required for encoding the size of a tile, information indicating the size of a sub - tile can be encoded and transmitted as a signal. A sub - tile is a basic unit that constitutes a tile, and each tile may be set to include at least one sub - tile. A sub - tile may include one or more encoding tree units.
[0390] For example, the syntax element subtile_width_minus1 indicates a value obtained by subtracting 1 from the width of a sub - tile. The syntax element subtile_height_minus1 indicates a value obtained by subtracting 1 from the height of a sub - tile.
[0391] Information indicating whether a tile other than the first tile has the same size as the previous tile can be encoded and transmitted as a signal. For example, the syntax element use_previous_tile_size_flag indicates whether the size of the current tile is the same as the size of the previous tile. If use_previous_tile_size_flag is true, it indicates that the size of the current tile is the same as the size of the previous tile. If use_previous_tile_size_flag is false, information indicating the size of the current tile can be encoded and transmitted as a signal. For the first tile, the encoding of use_previous_tile_size_flag may be omitted, and the value of the flag may be set to false.
[0392] The information indicating the size of a tile may include the syntax element tile_width_minus1[i] indicating the width of the i - th tile and the syntax element tile_height_minus1[i] indicating the height of the i - th tile.
[0393] The information indicating the size of a tile can indicate the difference value from the size of a child tile. When using the size information of a sub-tile, the encoding / decoding efficiency can be improved by reducing the number of bits required for encoding the size of each tile. For example, based on the following Equation 22, the width tileWidth of the i-th tile can be derived, and based on the following Equation 23, the height tileHeight of the i-th tile can be derived. [Number] [Number]
[0394] Alternatively, the encoding of the size information of the sub-tile may be omitted, and the size of the i-th tile can be directly encoded into the tile size information. The size information of the sub-tile can be selectively encoded. Information indicating whether the size information of the sub-tile is encoded can be transmitted as a signal via a video parameter set, a sequence parameter set, or an image parameter set.
[0395] The information related to the size of a tile may be encoded into something indicating the number of coding tree units and transmitted as a signal. For example, column_width_minus1, min_column_width_minus1, subtile_width_minus1, tile_width_minus1, etc. can indicate the number of coding tree unit columns included in the tile. Note that diff_column_width can indicate the difference value between the number of coding tree unit columns included in the minimum-width tile and the number of coding tree unit columns included in the current tile.
[0396] Note that row_height_minus1, min_row_height_minus1, subtile_height_minus1, tile_height_minus1, etc. can indicate the number of coded tree unit rows included in a tile. Note that diff_row_height can indicate the difference value between the number of coded tree unit rows included in the minimum height tile and the number of coded tree unit rows included in the current tile.
[0397] The decoder can determine the size of a tile based on the number of coded tree unit columns and / or the number of coded tree unit rows derived based on syntax elements, and the size of the coded tree unit. For example, the width of the i-th tile may be set to (tile_width_minus1[i]+1) * (width of the coded tree unit), and the height of the i-th tile may be set to (tile_height_minus1[i]+1) * (height of the coded tree unit).
[0398] At the same time, information indicating the size of the coded tree unit can be transmitted as a signal via a sequence parameter set or a picture parameter set.
[0399] In Table 11, the syntax element use_previous_tile_size_flag is used to explain whether the size of the current tile is the same as the size of the previous tile. As another example, information indicating whether the width of the current tile is the same as the width of the previous tile or information indicating whether the height of the current tile is the same as the height of the previous tile can be coded and transmitted as a signal.
[0400] Table 12 shows a syntax table including information indicating whether the width of the current tile is the same as the width of the previous tile.
Table 12
[0401] The syntax element use_previous_tile_width_flag indicates whether the width of the current tile is the same as the width of the previous tile. If use_previous_tile_width_flag is true, the width of the current tile can be set to be equal to the width of the previous tile. In this case, the encoding of the information indicating the width of the current tile may be omitted, and the width of the current tile can be derived from the width of the previous tile.
[0402] If use_previous_tile_width_flag is false, the information indicating the width of the current tile can be signaled. For example, tile_width_minus1[i] can indicate the value obtained by subtracting 1 from the width of the i-th tile.
[0403] The syntax element use_previous_tile_width_flag can be encoded and signaled only when it is determined that the size of the current tile is different from the size of the previous tile (for example, when the value of use_previous_tile_size_flag is 0).
[0404] tile_width_minus1[i] may have the value obtained by subtracting 1 from the number of columns of the coded tree unit sequence included in the i-th tile. The decoder can derive the number of columns of the coded tree units belonging to the i-th tile by adding 1 to tile_width_minus1[i], and can calculate the width of the tile by multiplying the derived value by the width of the coded tree unit.
[0405] Table 13 shows a syntax table that further includes information indicating whether the height of the current tile is the same as the height of the previous tile.
Table 13
[0406] JPEG0007693872000039.jpg95150
[0407] The syntax element use_previous_tile_height_flag indicates whether the height of the current tile is the same as the height of the previous tile. When the use_previous_tile_height_flag is true, the height of the current tile can be set to be equal to the height of the previous tile. In this case, the encoding of the information indicating the height of the current tile may be omitted, and the height of the current tile can be derived from the height of the previous tile.
[0408] When the use_previous_tile_height_flag is false, the information indicating the height of the current tile can be signaled. For example, tile_height_minus1[i] can indicate a value obtained by subtracting 1 from the height of the i-th tile.
[0409] The syntax element use_previous_tile_height_flag can be encoded and signaled only when it is determined that the size of the current tile is different from the size of the previous tile (for example, when the value of use_previous_tile_size_flag is 0). Note that the syntax element use_previous_tile_height_flag is signaled only when use_previous_tile_width_flag is false.
[0410] Table 12 shows an example of using use_previous_tile_width_flag, and Table 13 shows an example of using use_previous_tile_width_flag and use_previous_tile_height_flag. Although not shown in the above tables, the encoding of use_previous_tile_width_flag may be omitted, and only use_previous_tile_height_flag can be used.
[0411] Based on at least one of the tile scanning order, the width and height of the first tile, and the width and height of the previous tile, it is possible to determine which of use_previous_tile_height_flag and use_previous_tile_size_flag to use. For example, when the tile scanning order is vertical, use_previous_tile_height_flag can be used, and when the tile scanning order is horizontal, use_previous_tile_width_flag can be used. Alternatively, if the first tile or the previous tile is a non-square shape with a width larger than the height, use_previous_tile_width_flag can be used. If the first tile or the previous tile is a non-square shape with a height larger than the width, use_previous_tile_height_flag can be used.
[0412] When transmitting a signal indicating the number of tiles included in an image, for the last tile, encoding of information related to the tile size may be omitted.
[0413] Table 14 shows an example of omitting the encoding of tile size information for the last tile. [Table 14]
[0414] When specifying the size of tiles other than the last tile, the surplus area in the image can be made the last tile.
[0415] For each coding tree unit, an identifier (hereinafter referred to as tile ID, TileID) for recognizing the tile to which the coding tree unit belongs can be assigned.
[0416] FIG. 34 is a diagram showing an example of assigning tile IDs to each coding tree unit.
[0417] The same tile ID can be assigned to the encoding tree units belonging to the same tile. Specifically, the Nth TileID can be assigned to the encoding tree unit belonging to tile N.
[0418] To determine the tile ID assigned to each encoding tree unit, variables x and y indicating the position of the encoding tree unit in the image can be determined. Here, x represents the value obtained by dividing the x-axis coordinate of the position (x0, y0) of the top-left sample of the encoding tree unit by the width of the encoding tree unit, and y represents the value obtained by dividing the y-axis coordinate of the position (x0, y0) of the top-left sample of the encoding tree unit by the height of the encoding tree unit. Specifically, x and y can be derived by the following equations 24 and 25.
Equation
Equation
[0419] The operation of assigning tile IDs to each encoding tree unit can be executed by the process described below.
[0420] i) Initialization of tile ID The tile ID of each encoding tree unit may be initialized to the value obtained by subtracting 1 from the number of tiles in the image.
Table 15
[0421] ii) Derivation of tile ID
Table 16
[0422] In the above embodiments, it has been described that a flag indicating whether to allow applying a loop filter at the boundary of a tile is transmitted as a signal via an image parameter set. However, if it is set not to use the loop filter at any boundary of all tiles, problems such as a decrease in subjective image quality and a decrease in coding efficiency may occur.
[0423] Therefore, information indicating whether each tile allows applying a loop filter can be encoded and transmitted as a signal.
[0424] FIG. 35 is a diagram showing an example of selectively determining whether to apply a loop filter to each tile.
[0425] In the example shown in FIG. 35, for each tile, it is possible to determine whether to allow applying a loop filter (e.g., a deblocking filter, SAO, and / or ALF) at a horizontal or vertical boundary.
[0426] Table 17 shows an example of encoding information indicating whether to allow applying a loop filter to each tile. [Table 17]
[0427] In the example of Table 17, the syntax element loop_filter_across_tiles_flag[i] indicates whether to allow applying a loop filter to the i-th tile. When the value of loop_filter_across_tile_flag[i] is 1, it indicates that the loop filter can be used at the horizontal and vertical boundaries of the tile with tile ID i. When the value of loop_filter_across_tile_flag[i] is 0, it indicates that the loop filter is not used at the horizontal and vertical boundaries of the tile with tile ID i.
[0428] Information indicating whether to allow the application of the loop filter in each of the horizontal and vertical directions can be encoded.
[0429] Table 18 shows examples of encoding information indicating whether to allow the application of the loop filter in the horizontal and vertical directions, respectively. [Table 18]
[0430] In the example of Table 18, the syntax element loop_filter_hor_across_tiles_flag[i] indicates whether to allow the application of the loop filter at the position crossing the i-th tile in the horizontal direction. The syntax element loop_filter_ver_across_tiles_flag[i] indicates whether to allow the application of the loop filter at the position crossing the i-th tile in the vertical direction.
[0431] If the value of loop_filter_hor_across_tile_flag[i] is 1, it indicates that the loop filter can be used at the horizontal boundary of the tile with tile ID i. If the value of loop_filter_hor_across_tile_flag[i] is 0, it indicates that the loop filter is not used at the vertical boundary of the tile with tile ID i.
[0432] If the value of loop_filter_ver_across_tile_flag[i] is 1, it indicates that the loop filter can be used at the vertical boundary of the tile with tile ID i. If the value of loop_filter_ver_across_tile_flag[i] is 0, it indicates that the loop filter is not used at the vertical boundary of the tile with tile ID i.
[0433] Information indicating whether a tile group including a plurality of tiles permits the application of a loop filter can be encoded and transmitted as a signal. Based on this information, it is possible to determine whether the plurality of tiles included in the tile group permit the application of the loop filter.
[0434] To determine a tile group, at least one of the number of tiles belonging to the tile group, the size of the tile group, and the image division information can be transmitted as a signal via a bit stream. Alternatively, a region with a pre-defined size in an encoder and a decoder can be used as the tile group.
[0435] Encoding of information indicating whether to permit the application of the loop filter may be omitted, and based on at least one of the number of coded tree units included in a tile, the width of the tile, and the height of the tile, it is possible to determine whether to permit the application of the loop filter. For example, when the width of a tile is less than a reference value, applying the loop filter in the horizontal direction is permitted, and when the height of the tile is less than the reference value, the loop filter can be applied in the vertical direction.
[0436] When using a loop filter at the tile boundary, reconstructed data outside the tile can be generated based on the data included in the tile. In this case, by performing padding or interpolation on the data included in the tile, a reconstructed video outside the tile can be obtained. Subsequently, the loop filter can be applied by using the reconstructed data outside the tile.
[0437] Examples described with an emphasis on the decoding process or the encoding process and those used in the decoding process or the encoding process are also included within the scope of the present invention. Those obtained by changing a plurality of examples described in a predetermined order in an order different from the described order are also included within the scope of the present invention.
[0438] Examples have been described based on a series of steps or flowcharts, which do not limit the chronological order of the invention. Also, if necessary, they may be executed simultaneously or in other orders. In the above examples, the components (e.g., units, modules, etc.) constituting the block diagrams may each be implemented as a hardware device or software. Also, a plurality of components may be combined and executed as a single hardware device or software. The above examples may be executed in the form of program instructions. The program instructions may be executed by various computer components and recorded on a computer-readable storage medium. The computer-readable storage medium may include a program instruction, a data file, a data structure, etc. alone or in combination. Examples of computer-readable storage media include magnetic media such as hard disks, flexible disks, and magnetic tapes, optical recording media such as CD-ROMs, DVDs, magneto-optical media such as floptical disks, and hardware devices specifically arranged in a manner to store and execute program instructions such as ROMs, RAMs, flash memories, etc. The hardware device may operate as one or more software modules and be configured to execute the processing according to the present invention, and vice versa.
Industrial Applicability
[0439] The present invention is applicable to electronic devices that perform encoding / decoding on videos.
Claims
1. 1. A video decoding method, comprising: generating a merge candidate list for the current block; designating one of a plurality of merging candidates included in the merging candidate list; deriving a first affine seed vector (sv0) and a second affine seed vector (sv1) of the current block based on a first affine seed vector (nv0) and a second affine seed vector (nv1) of a designated merging candidate; deriving an affine vector of a sub-block in the current block by using the first affine seed vector (sv0) and the second affine seed vector (sv1) of the current block, the sub-block being an area having a size smaller than the size of the current block; performing motion compensated prediction on the sub-block based on the affine vector; the first affine seed vector (nv0) and the second affine seed vector (nv1) of the merging candidate are derived based on motion information of a neighboring block adjacent to the current block; If the neighboring block is included in a coding tree unit different from a coding tree unit of the current block, the first affine seed vector (nv0) and the second affine seed vector (nv1) of the merging candidate are derived based on motion vectors of a lower left sub-block and a lower right sub-block of the neighboring block; the lower-left sub-block includes a lower-left affine reference sample control point (xn4, yn4) located at the lower-left corner of the adjacent block, and the lower-right sub-block is adjacent to a lower-right affine reference sample control point (xn5, yn5) located to the right of the lower-right sample of the lower-right sub-block.
2. the first affine seed vector and the second affine seed vector of the merging candidate are derived from values obtained by performing a shift operation on a difference value of a motion vector between the lower left sub-block and the lower right sub-block; The shift operation shifts the difference value by a scale factor, the scale factor being derived from a value obtained by adding a horizontal distance between a lower left reference sample and a lower right reference sample and an offset. The video decoding method of claim 1.
3. the first affine seed vector and the second affine seed vector of the merging candidate are derived from values obtained by performing a shift operation on a difference value of a motion vector between the lower left sub-block and the lower right sub-block; The shift operation shifts the difference value by a scale factor, and the scale factor is derived from a distance between the lower left reference sample and an adjacent sample to the right of the lower right reference sample. The video decoding method of claim 1.
4. 1. A video encoding method comprising the steps of: generating a merge candidate list for the current block; designating one of a plurality of merging candidates included in the merging candidate list; deriving a first affine seed vector (sv0) and a second affine seed vector (sv1) of the current block based on a first affine seed vector (nv0) and a second affine seed vector (nv1) of a designated merging candidate; deriving an affine vector of a sub-block in the current block by using the first affine seed vector (sv0) and the second affine seed vector (sv1) of the current block, the sub-block being an area having a size smaller than the size of the current block; performing motion compensated prediction on the sub-block based on the affine vector; the first affine seed vector (nv0) and the second affine seed vector (nv1) of the merging candidate are derived based on motion information of a neighboring block adjacent to the current block; If the neighboring block is included in a coding tree unit different from a coding tree unit of the current block, the first affine seed vector (nv0) and the second affine seed vector (nv1) of the merging candidate are derived based on motion vectors of a lower left sub-block and a lower right sub-block of the neighboring block; the lower-left sub-block includes a lower-left affine reference sample control point (xn4, yn4) located at the lower-left corner of the adjacent block, and the lower-right sub-block is adjacent to a lower-right affine reference sample control point (xn5, yn5) located to the right of the lower-right sample of the lower-right sub-block.
5. the first affine seed vector and the second affine seed vector of the merging candidate are derived from values obtained by performing a shift operation on a difference value of a motion vector between the lower left sub-block and the lower right sub-block; The shift operation shifts the difference value by a scale factor, and the scale factor is derived from a value obtained by adding a horizontal distance between the lower-left reference sample and the lower-right reference sample and an offset, or the scale factor is derived from a distance between the lower-left reference sample and an adjacent sample adjacent to the right of the lower-right reference sample.
5. The video encoding method of claim 4.
6. A video decoder configured to perform the video decoding method according to any one of claims 1 to 3.
7. A video encoder configured to perform the video encoding method according to any one of claims 4 to 5.
Citation Information
Patent Citations
Prediction image generation device, moving image decoding device, and moving image encoding device
WO2017130696A1
Affine motion prediction for video coding
WO2017200771A1
Affine motion vector derivation device, prediction image generation device, moving image decoding device, and moving image coding device
WO2018061563A1
Motion vector prediction for affine motion models in video coding
WO2018067823A1