Video signal encoding / decoding method and apparatus for said method
By processing the residual coefficients with non-zero flags, absolute value information, and parity flags, the problem of increased data volume in high-definition video services is solved, achieving more efficient video signal encoding/decoding.
Patent Information
- Application Number
- CN202310145089.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-02-26
- Filing Date
- 2019-09-20
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2039-09-20
AI Technical Summary
Existing video encoding technologies face the problem of a significant increase in data volume in high-definition video services. HEVC's compression performance has gradually revealed its limitations, and more efficient encoding/decoding methods are needed to handle residual coefficients.
By parsing the non-zero flag indicating whether the residual coefficient is non-zero, its absolute value is determined. Encoding/decoding is performed using the comparison flag between the residual coefficient and the threshold and the parity flag. Only when the residual coefficient is greater than the first value is the parity flag and residual value information of the residual coefficient further parsed and adjusted.
It achieves efficient encoding/decoding of residual coefficients, thereby improving the encoding efficiency of video signals.
Smart Images

Figure CN116320397B_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese patent application No. 201980058803.3, entitled "Video Signal Encoding / Decoding Method and Apparatus for the Method", which entered the Chinese national phase on PCT international patent application PCT / KR2019 / 012291 filed on September 20, 2019. Technical Field
[0002] This invention relates to a video signal encoding / decoding method and an apparatus for the method. Background Technology
[0003] With the trend of increasingly larger display panels, there is a growing need for higher-quality video services. The biggest problem with high-definition video services is the significant increase in data volume. To address this issue, research is actively underway to improve video compression rates. As a representative example, in 2009, the Moving Picture Experts Group (MPEG) and the Video Coding Experts Group (VCEG) of the International Telecommunication Union-Telecommunication (ITU-T) established the Joint Collaborative Team on Video Coding (JCT-VC). JCT-VC proposed the video compression standard HEVC (High Efficiency Video Coding), which was approved on January 25, 2013, and its compression performance is approximately twice that of H.264 / AVC. However, with the rapid development of high-definition video services, the limitations of HEVC have gradually become apparent. Summary of the Invention
[0004] Technical problems to be solved
[0005] The purpose of this invention is to provide a method for encoding / decoding residual coefficients during the encoding / decoding of video signals, and an apparatus for using the method.
[0006] The purpose of this invention is to provide a method for encoding / decoding residual coefficients by comparing the magnitude of the residual coefficients with a threshold when encoding / decoding video signals, and an apparatus for using the method.
[0007] The purpose of this invention is to provide a method and an apparatus for encoding / decoding residual coefficients by using a flag indicating whether the residual coefficients are even or odd when encoding / decoding video signals.
[0008] The technical problems to be solved by the present invention are not limited to those mentioned above, and those skilled in the art to which the present invention pertains will clearly understand other technical problems not mentioned through the following description.
[0009] Technical solution
[0010] The video signal decoding method of the present invention includes the following steps: parsing a non-zero flag indicating whether the residual coefficient is non-zero from the bitstream; when the non-zero flag indicates that the residual coefficient is not non-zero, parsing absolute value information from the bitstream, the absolute value information being used to determine the absolute value of the residual coefficient; and determining the absolute value of the residual coefficient based on the absolute value information. The absolute value information may include a residual coefficient comparison flag indicating whether the residual coefficient is greater than a first value; only when the residual coefficient is greater than the first value is a parity flag further parsed from the bitstream.
[0011] The video signal encoding method of the present invention may include the following steps: encoding a non-zero flag indicating whether the residual coefficient is non-zero; and encoding absolute value information when the residual coefficient is not non-zero, the absolute value information being used to determine the absolute value of the residual coefficient. The absolute value information may include a residual coefficient comparison flag indicating whether the residual coefficient is greater than a first value, and only when the residual coefficient is greater than the first value is a parity flag of the residual coefficient further encoded.
[0012] In the video signal decoding / encoding method of the present invention, the parity flag can indicate whether the value of the residual coefficient is even or odd.
[0013] In the video signal decoding / encoding method of the present invention, when the residual coefficient is greater than the first value, the first adjustment residual coefficient comparison flag can be further analyzed. The first adjustment residual coefficient comparison flag indicates whether the adjustment residual coefficient derived by shifting the residual coefficient to the right by 1 is greater than the second value.
[0014] In the video signal decoding / encoding method of the present invention, when the adjusted residual coefficient is below the second value, the residual coefficient can be determined as 2N or 2N+1 according to the value of the parity flag.
[0015] In the video signal decoding / encoding method of the present invention, when the adjustment residual coefficient is greater than the second value, a second adjustment residual coefficient comparison flag can be further analyzed. The second adjustment residual coefficient comparison flag indicates whether the adjustment residual coefficient is greater than a third value.
[0016] In the video signal decoding / encoding method of the present invention, when the adjustment residual coefficient is greater than the second value, the residual value information can be further analyzed. The residual value information is a value obtained by subtracting the second value from the adjustment residual coefficient.
[0017] The features briefly outlined above are merely exemplary embodiments of the invention as described in the detailed description to follow, and do not limit the scope of the invention.
[0018] Invention Effects
[0019] According to the present invention, residual coefficients can be encoded / decoded efficiently.
[0020] According to the present invention, residual coefficients can be efficiently encoded / decoded by using a flag that compares the magnitude of the residual coefficients with a threshold.
[0021] According to the present invention, residual coefficients can be efficiently encoded / decoded by using a flag indicating whether the residual coefficients are even or odd.
[0022] The effects that can be obtained in this invention are not limited to those described above, and other effects not mentioned will be clearly understood by those skilled in the art through the following description. Attached Figure Description
[0023] Figure 1 This is a block diagram of a video encoder according to an embodiment of the present invention.
[0024] Figure 2 This is a block diagram of a video decoder according to an embodiment of the present invention.
[0025] Figure 3 This is a diagram illustrating the basic coding tree unit of an embodiment of the present invention.
[0026] Figure 4 This is a diagram illustrating the various partitioning types of coded blocks.
[0027] Figure 5 This is a diagram illustrating an example of the division of a coding tree unit.
[0028] Figure 6 This is a diagram showing an example of a block smaller than a coding tree unit of a preset size appearing at the image boundary.
[0029] Figure 7 This is a diagram illustrating an example of performing quadtree partitioning on boundary blocks with atypical boundaries.
[0030] Figure 8 This is a diagram illustrating an example of performing a quadtree partition on a block adjacent to an image boundary.
[0031] Figure 9 This is a diagram showing the partitioning pattern of blocks adjacent to the image boundary.
[0032] Figure 10 This is a diagram showing the encoding patterns of blocks adjacent to the image boundaries.
[0033] Figure 11 This is a flowchart of the inter-frame prediction method according to an embodiment of the present invention.
[0034] Figure 12 This is a flowchart of the process of exporting motion information of the current block in merge mode.
[0035] Figure 13 This is a diagram showing the candidate blocks used to derive the merge candidates.
[0036] Figure 14 This is a diagram showing the location of the reference sample.
[0037] Figure 15 This is a diagram showing the candidate blocks used to derive the merge candidates.
[0038] Figure 16 This is a diagram illustrating an example of changing the position of a reference sample.
[0039] Figure 17 This is a diagram illustrating an example of changing the position of a reference sample.
[0040] Figure 18 This is a diagram illustrating an example of updating the list of motion information between frames.
[0041] Figure 19 This is a diagram illustrating an example of an updated inter-frame merge candidate list.
[0042] Figure 20 This is a diagram showing an example of how the index of a stored inter-frame merge candidate is updated.
[0043] Figure 21 This is a diagram showing the location of a representative sub-block.
[0044] Figure 22 An example of generating a list of inter-frame motion information for different inter-frame prediction modes is shown.
[0045] Figure 23 This is a diagram illustrating an example of adding inter-frame merge candidates included in the long-term motion information list to the merge candidate list.
[0046] Figure 24 This is a diagram illustrating an example of performing redundancy checks only on some of the merge candidates.
[0047] Figure 25This is a diagram illustrating an example of omitting redundancy checks for a specific merge candidate.
[0048] Figure 26 This is a diagram illustrating an example where candidate blocks included in the same side-by-side merge region as the current block are set to be unusable as merge candidates.
[0049] Figure 27 This is a diagram showing a list of time-domain motion information.
[0050] Figure 28 This is a diagram illustrating an example of merging the inter-frame motion information list and the temporal motion information list.
[0051] Figure 29 This is a flowchart of the intra-frame prediction method according to an embodiment of the present invention.
[0052] Figure 30 This is a diagram showing the reference samples included in each reference sample line.
[0053] Figure 31 This is a diagram illustrating the intra-frame prediction mode.
[0054] Figure 32 and Figure 33 This is a diagram illustrating an example of a one-dimensional arrangement of reference samples in a row.
[0055] Figure 34 This is a diagram showing the angle formed between the prediction pattern within the angular frame and a straight line parallel to the x-axis.
[0056] Figure 35 This is a diagram showing an example of obtaining a predicted sample when the current block is not square.
[0057] Figure 36 This is a diagram illustrating the wide-angle intra-frame prediction mode.
[0058] Figure 37 This is a diagram illustrating the application of PDPC.
[0059] Figure 38 and Figure 39 This is a diagram showing the sub-block to which a second transformation will be performed.
[0060] Figure 40 This is a diagram used to illustrate an example of determining the transformation type of the current block.
[0061] Figure 41 This is a flowchart illustrating a method for encoding residual coefficients.
[0062] Figure 42 and Figure 43 It is a graph showing the arrangement of residual coefficients according to different scanning orders.
[0063] Figure 44 An example of encoding the position of the last non-zero coefficient is shown.
[0064] Figure 45 This is a flowchart of the process of encoding the absolute value of the residual coefficients.
[0065] Figure 46 This is a flowchart of the process of encoding the absolute value of the residual coefficients.
[0066] Figure 47 This is a flowchart of the process of encoding the absolute value of the residual coefficients.
[0067] Figure 48 This is a flowchart illustrating the process of determining block strength.
[0068] Figure 49 Predefined filter candidates are shown. Detailed Implementation
[0069] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0070] Video encoding and decoding are performed on a block-by-block basis. For example, encoding / decoding processes such as transform, quantization, prediction, in-loop filtering, or reconstruction can be performed on encoded blocks, transform blocks, or prediction blocks.
[0071] Hereinafter, the block to be encoded / decoded will be referred to as the "current block". For example, depending on the current encoding / decoding process step, the current block can represent an encoded block, a transform block, or a prediction block.
[0072] Additionally, as used herein, the term "unit" refers to a basic unit used to perform a specific encoding / decoding process, and "block" can be understood as representing a sample array of a predetermined size. Unless otherwise stated, "block" and "unit" are used interchangeably. For example, in the embodiments described later, encoding block and encoding unit can be understood to have the same meaning.
[0073] Figure 1 This is a block diagram of a video encoder according to an embodiment of the present invention.
[0074] Reference Figure 1 The video encoding device 100 may include an image segmentation unit 110, a prediction unit 120 and 125, a transformation unit 130, a quantization unit 135, a rearrangement unit 160, an entropy coding unit 165, an inverse quantization unit 140, an inverse transformation unit 145, a filter unit 150, and a memory 155.
[0075] Figure 1The components shown are illustrated individually to illustrate the distinct functionalities of the video encoding device and do not imply that each component is composed of separate hardware or a single software component. That is, for ease of explanation, the components are arranged such that at least two components are combined into one, or one component is divided into multiple components, thereby performing functions. Such embodiments of integrated components and embodiments of separated components are also within the scope of this invention, provided they do not depart from its spirit.
[0076] Furthermore, some structural elements are not essential structural elements for performing the essential functions of this invention, but rather optional structural elements used only to improve performance. This invention can be implemented by including only the components necessary for realizing the essence of the invention, excluding the structural elements used only to improve performance, and structures including only the essential structural elements, excluding the optional structural elements used only to improve performance, are also within the scope of this invention.
[0077] The image partitioning unit 110 can divide the input image into at least one processing unit. In this case, the processing unit can be a prediction unit (PU), a transformation unit (TU), or a coding unit (CU). The image partitioning unit 110 divides an image into a combination of multiple coding units, prediction units, and transformation units, and can select a combination of coding units, prediction units, and transformation units to encode the image based on a predetermined criterion (e.g., a cost function).
[0078] For example, an image can be divided into multiple coding units. To segment an image into coding units, a recursive tree structure such as a quadtree structure can be used. A video or the largest coding unit can be used as the root, and the coding unit can be divided into additional coding units with a number of child nodes equivalent to the number of coding units in the division. Coding units that are no longer divided according to certain constraints become leaf nodes. That is, when it is assumed that a coding unit can only be divided into squares, a coding unit can be divided into a maximum of four other coding units.
[0079] In the embodiments of the present invention, the encoding unit may mean a unit that performs encoding, or it may mean a unit that performs decoding.
[0080] A prediction unit within a coding unit can be divided into at least one shape of the same size, such as a square or a rectangle, or a prediction unit within a coding unit can be divided into shapes and / or sizes different from those of another prediction unit.
[0081] When the prediction unit for intra-frame prediction based on the coding unit is not the smallest coding unit, intra-frame prediction can be performed without splitting into multiple prediction units N×N.
[0082] Prediction units 120 and 125 may include an inter-frame prediction unit 120 that performs inter-frame prediction and an intra-frame prediction unit 125 that performs intra-frame prediction. It can be determined whether inter-frame prediction or intra-frame prediction is used for the prediction unit, and specific information (e.g., intra-frame prediction mode, motion vector, reference image, etc.) is determined based on each prediction method. In this case, the processing unit performing the prediction may be different from the processing unit that determines the prediction method and specific content. For example, the prediction unit may determine the prediction method and prediction mode, and the transformation unit may perform the prediction. The residual value (residual block) between the generated prediction block and the original block can be input to the transformation unit 130. Furthermore, the prediction mode information, motion vector information, etc., used for prediction, along with the residual value, can be encoded in the entropy coding unit 165 and transmitted to the decoder. When using a specific coding mode, the original block can also be directly encoded and transmitted to the decoder without generating a prediction block through the prediction units 120 and 125.
[0083] The inter-frame prediction unit 120 can predict prediction units based on information from at least one image preceding or following the current image. In some cases, it can also predict prediction units based on information from a portion of the encoded region within the current image. The inter-frame prediction unit 120 may include a reference image interpolation unit, a motion prediction unit, and a motion compensation unit.
[0084] The reference image interpolation unit receives reference image information from memory 155 and can generate pixel information of integer pixels or fractional pixels from the reference image. For luminance pixels, in order to generate fractional pixel information in 1 / 4 pixel units, an 8th-order DCT-based interpolation filter with different filter coefficients can be used. For chrominance signals, in order to generate fractional pixel information in 1 / 8 pixel units, a 4th-order DCT-based interpolation filter with different filter coefficients can be used.
[0085] The motion prediction unit can perform motion prediction based on a reference image interpolated by the reference image interpolation unit. Various methods can be used to calculate motion vectors, such as the Full Search-based Block Matching Algorithm (FBMA), the Three-Step Search (TSS), and the New Three-Step Search Algorithm (NTS). Motion vectors can have values in units of 1 / 2 pixel or 1 / 4 pixel based on the interpolated pixels. Different motion prediction methods can be used in the motion prediction unit to predict the current prediction unit. These methods include skipping, merging, Advanced Motion Vector Prediction (AMVP), and Intra Block Copying.
[0086] The intra-prediction unit 125 can generate prediction units based on reference pixel information surrounding the current block, which serves as pixel information within the current image. When the neighboring block of the current prediction unit is a block that has already undergone inter-frame prediction, and the reference pixel is a pixel that has undergone inter-frame prediction, the reference pixel included in the block that has undergone inter-frame prediction can be used as the reference pixel information for the surrounding block that has undergone intra-frame prediction. That is, when a reference pixel is unavailable, at least one of the available reference pixels can be used to replace the unavailable reference pixel information.
[0087] In intra-frame prediction, the prediction mode can have an angular prediction mode that uses reference pixel information in the prediction direction and a non-angular mode that does not use direction information when performing prediction. The mode used to predict luminance information and the mode used to predict chrominance information can be different. To predict chrominance information, either the intra-frame prediction mode information used for predicting luminance information or the predicted luminance signal information can be applied.
[0088] When performing intra-frame prediction, if the size of the prediction unit is the same as the size of the transform unit, intra-frame prediction can be performed based on pixels to the left, upper left, and upper right of the prediction unit. However, when performing intra-frame prediction, if the size of the prediction unit is different from the size of the transform unit, intra-frame prediction can be performed using reference pixels based on the transform unit. Furthermore, intra-frame prediction using an N×N partition only for the smallest coding unit can be applied.
[0089] Intra-prediction methods can generate prediction blocks after applying an Adaptive Intra Smoothing (AIS) filter to a reference pixel based on the prediction mode. The type of AIS filter used for the reference pixel may vary. To perform intra-prediction, the intra-prediction mode of the current prediction unit can be predicted from the intra-prediction modes of prediction units existing in the vicinity of the current prediction unit. When using mode information predicted from surrounding prediction units to predict the prediction mode of the current prediction unit, if the intra-prediction modes of the current prediction unit and those of the surrounding prediction units are the same, predetermined flag information can be used to convey information indicating that the prediction modes of the current prediction unit and those of the surrounding prediction units are the same. If the prediction modes of the current prediction unit and those of the surrounding prediction units are different, entropy coding can be performed to encode the prediction mode information of the current block.
[0090] Furthermore, a residual block can be generated that includes residual value information, which is the difference between the original block of the prediction unit and the prediction unit that performs prediction based on the prediction unit generated in the prediction units 120 and 125. The generated residual block can be input to the transformation unit 130.
[0091] The transform unit 130 can transform the original block and the residual block, which includes the residual coefficient information of the prediction units generated by the prediction units 120 and 125, using transformation methods such as Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), or KL Transform (KLT). Whether to apply DCT, DST, or KLT to transform the residual block can be determined based on the intra-frame prediction mode information of the prediction units used to generate the residual block.
[0092] The quantization unit 135 can quantize the values that have been transformed into the frequency domain in the transformation unit 130. The quantization coefficients can be changed according to the importance of the block or video. The values calculated in the quantization unit 135 can be provided to the inverse quantization unit 140 and the rearrangement unit 160.
[0093] The rearrangement unit 160 can rearrange the coefficient values of the quantized residual values.
[0094] The rearrangement unit 160 can transform 2D block shape coefficients into 1D vector form using a coefficient scanning method. For example, the rearrangement unit 160 can use a zig-zag scan method to scan the DC coefficients and even the coefficients in the high-frequency domain, and transform them into 1D vector form. Depending on the size of the transform unit and the intra-frame prediction mode, instead of zig-zag scanning, vertical scanning along the column direction and horizontal scanning along the row direction can also be used to scan the 2D block shape coefficients. That is, the choice between zig-zag scanning, vertical scanning, and horizontal scanning can be determined based on the size of the transform unit and the intra-frame prediction mode.
[0095] The entropy coding unit 165 can perform entropy coding based on the value calculated by the rearrangement unit 160. For example, entropy coding can use various coding methods such as Exponential Golomb code, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC).
[0096] The entropy coding unit 165 can encode various information such as residual coefficient information, block type information, prediction mode information, partitioning unit information, prediction unit information and transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information of the coding units originating from the rearrangement unit 160 and the prediction units 120 and 125.
[0097] The coefficient values of the coding units input from the rearrangement unit 160 can be entropy encoded in the entropy coding unit 165.
[0098] The inverse quantization unit 140 and the inverse transform unit 145 perform inverse quantization on the multiple values quantized in the quantization unit 135, and perform inverse transform on the values transformed in the transform unit 130. The residual values generated in the inverse quantization unit 140 and the inverse transform unit 145 can be merged with the prediction units predicted by the motion prediction unit, motion compensation unit and intra-frame prediction unit included in the prediction units 120 and 125 to generate a reconstructed block.
[0099] The filter unit 150 may include at least one of a deblocking filter, an offset correction unit, and an adaptive loop filter (ALF).
[0100] Deblocking filters remove block distortion generated in the reconstructed image due to the boundaries between blocks. To determine whether to perform deblocking, the number of pixels in the columns or rows included in the block can be used to decide whether to apply a deblocking filter to the current block. When applying a deblocking filter to a block, a strong or weak filter can be applied depending on the desired deblocking intensity. Furthermore, during the use of deblocking filters, horizontal and vertical filtering can be processed simultaneously when performing vertical and horizontal filtering.
[0101] The offset correction unit can correct the offset between the video being deblocked and the original video on a pixel-by-pixel basis. To perform offset correction on a specific image, the following methods can be used: after dividing the pixels included in the video into a predetermined number of regions, determine the region to be offset and apply the offset to the corresponding region, or apply the offset by taking into account the edge information of each pixel.
[0102] Adaptive Loop Filtering (ALF) can be performed based on a comparison between the filtered reconstructed image and the original video. After dividing the pixels in the video into predetermined groups, filtering can be performed differently for each group by determining a filter to be used for the corresponding group. Information related to whether adaptive loop filtering is applied, along with the luminance signal, can be transmitted per coding unit (CU). The shape and filter coefficients of the adaptive loop filter to be applied can vary depending on the block. Furthermore, it is possible to apply the same type (fixed type) of adaptive loop filter regardless of the characteristics of the block to which it is applied.
[0103] The memory 155 can store the reconstructed blocks or images calculated by the filter unit 150, and can provide the stored reconstructed blocks or images to the prediction units 120 and 125 when performing inter-frame prediction.
[0104] Figure 2 This is a block diagram of a video decoder according to an embodiment of the present invention.
[0105] Reference Figure 2 The video decoder 200 may include an entropy decoding unit 210, a rearrangement unit 215, an inverse quantization unit 220, an inverse transform unit 225, a prediction unit 230, a prediction unit 235, a filter unit 240, and a memory 245.
[0106] When a video stream is input into a video encoder, the input stream can be decoded in the reverse order of the video encoder steps.
[0107] The entropy decoding unit 210 can perform entropy decoding in the reverse order of the entropy encoding steps performed in the entropy encoding unit of the video encoder. For example, corresponding to the methods performed in the video encoder, various methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) can be applied.
[0108] The entropy decoding unit 210 can decode information related to intra-frame prediction and inter-frame prediction performed by the encoder.
[0109] The rearrangement unit 215 can perform rearrangement based on a method used in the encoding unit to rearrange the bitstream that has been entropily decoded by the entropy decoding unit 210. Rearrangement can be performed by reconstructing multiple coefficients represented in 1D vector form into 2D block-shaped coefficients. The rearrangement unit 215 receives information related to the coefficient scan performed in the encoding unit and can perform rearrangement by performing a reverse scan based on the scan order performed in the corresponding encoding unit.
[0110] The inverse quantization unit 220 can perform inverse quantization based on the quantization parameters provided by the encoder and the coefficient values of the rearranged blocks.
[0111] Based on the quantization result performed by the video encoder, the inverse transform unit 225 can perform inverse transforms, namely inverse DCT, inverse DST, and inverse KLT, on the transforms performed by the transform unit. The inverse transform can be performed based on the transmission unit determined in the video encoder. In the inverse transform unit 225 of the video decoder, transform methods (e.g., DCT, DST, KLT) can be selectively performed based on multiple factors such as the prediction method, the size of the current block, and the prediction direction.
[0112] Prediction units 230 and 235 can generate prediction blocks based on information related to prediction block generation provided by entropy decoding unit 210 and previously decoded block or image information provided by memory 245.
[0113] As described above, when intra-prediction is performed in the same manner as in the video encoder, if the size of the prediction unit is the same as the size of the transform unit, intra-prediction is performed on the prediction unit based on the pixels to its left, the pixels to its upper left, and the pixels above it. If the size of the prediction unit is different from the size of the transform unit, intra-prediction can be performed using reference pixels based on the transform unit. Furthermore, intra-prediction using only N×N partitioning for the smallest coding unit can also be applied.
[0114] Prediction units 230 and 235 may include a prediction unit determination unit, an inter-frame prediction unit, and an intra-frame prediction unit. The prediction unit determination unit receives various information input from the entropy decoding unit 210, such as prediction unit information, prediction mode information of the intra-frame prediction method, and motion prediction-related information of the inter-frame prediction method. It classifies prediction units according to the current coding unit and determines whether the prediction unit is performing inter-frame prediction or intra-frame prediction. The inter-frame prediction unit 230 can use the information required for inter-frame prediction of the current prediction unit provided by the video encoder and perform inter-frame prediction on the current prediction unit based on information included in at least one of the preceding or subsequent images of the current image to which the current prediction unit belongs. Alternatively, inter-frame prediction can also be performed based on information from a portion of the reconstructed region within the current image to which the current prediction unit belongs.
[0115] In order to perform inter-frame prediction, it is possible to determine, based on the coding unit, which of the following modes of motion prediction method is used for the prediction units included in the corresponding coding unit: Skip Mode, Merge Mode, Advanced Motion Vector Prediction Mode (AMVP Mode), or Intra-block Copy Mode.
[0116] The intra-prediction unit 235 can generate prediction blocks based on pixel information within the current image. When the prediction unit is one that has already performed intra-prediction, intra-prediction can be performed based on the intra-prediction mode information of the prediction unit provided by the video encoder. The intra-prediction unit 235 may include an adaptive intra-smoothing (AIS) filter, a reference pixel interpolation unit, and a DC filter. The adaptive intra-smoothing filter is the part that performs filtering on the reference pixels of the current block, and whether to apply the filter can be determined according to the prediction mode of the current prediction unit. Adaptive intra-smoothing filtering can be performed on the reference pixels of the current block using the prediction mode of the prediction unit provided by the video encoder and the adaptive intra-smoothing filter information. If the prediction mode of the current block is a mode that does not perform adaptive intra-smoothing filtering, then the adaptive intra-smoothing filter may not be applied.
[0117] For the reference pixel interpolation unit, if the prediction mode of the prediction unit is a prediction unit that performs intra-frame prediction based on pixel values interpolated from reference pixels, then reference pixels with integer or fractional pixel units can be generated by interpolating the reference pixels. If the prediction mode of the current prediction unit is a prediction mode that generates prediction blocks without interpolating reference pixels, then interpolation of reference pixels is not required. If the prediction mode of the current block is DC mode, then the DC filter can generate prediction blocks by filtering.
[0118] The reconstructed blocks or images can be provided to the filter unit 240. The filter unit 240 may include a deblocking filter, an offset correction unit, and an ALF.
[0119] Information related to whether to apply a deblocking filter to a corresponding block or image can be received from the video encoder, as well as information regarding whether to apply strong or weak filtering when applying the deblocking filter. Information related to the deblocking filter provided by the video encoder can be received from the video decoder's deblocking filter, and deblocking filtering can be performed on the corresponding block at the video decoder.
[0120] The offset correction unit can perform offset correction on the reconstructed video based on the type and amount of offset correction used during video encoding.
[0121] The ALF can be applied to the coding unit based on information provided by the encoder, such as whether the ALF is applied and ALF coefficient information. This ALF information can be provided by including it in a specific parameter set.
[0122] The memory 245 stores the reconstructed image or block, such that the image or block can be used as a reference image or reference block, and the reconstructed image can be provided to the output unit.
[0123] Figure 3 This is a diagram illustrating the basic coding tree unit of an embodiment of the present invention.
[0124] The largest coding block can be defined as the coding tree block. An image can be divided into multiple coding tree units (CTUs). The coding tree unit is the largest coding unit and can also be called the largest coding unit (LCU). Figure 3 An example of dividing an image into multiple coding tree units is shown.
[0125] The size of a coding tree unit can be defined at the image level or the sequence level. Therefore, information representing the size of a coding tree unit can be transmitted via signals using either an image parameter set or a sequence parameter set.
[0126] For example, the size of the coding tree unit for the entire image within the sequence can be set to 128×128. Alternatively, either 128×128 or 256×256 at the image level can be determined as the size of the coding tree unit. For example, the size of the coding tree unit in the first image can be set to 128×128, and the size of the coding tree unit in the second image can be set to 256×256.
[0127] Coded blocks can be generated by dividing the coding tree into units. A coded block represents the basic unit used for encoding / decoding processing. For example, prediction or transformation can be performed on different coded blocks, or prediction coding modes can be determined on different coded blocks. The prediction coding mode represents the method for generating the predicted image. For example, prediction coding modes can include intra-prediction, inter-prediction, current picture referencing (CPR, or intra-block copy (IBC)), or combined prediction. For a coded block, at least one of the prediction coding modes—intra-prediction, inter-prediction, current picture referencing, or combined prediction—can be used to generate the prediction block associated with that coded block.
[0128] Information indicating the predictive coding mode of the current block can be transmitted via signals in the bitstream. For example, this information could be a 1-bit flag indicating whether the predictive coding mode is intra-frame or inter-frame. Current image reference or combined prediction can be used only if the predictive coding mode of the current block is determined to be inter-frame.
[0129] The current image reference is used to set the current image as the reference image and obtain the prediction block of the current block from the encoded / decoded region within the current image. Here, the current image means the image that includes the current block. Information indicating whether the current image reference is applied to the current block can be sent via signals in the bitstream. For example, this information could be a 1-bit flag. When the flag is true, the prediction coding mode of the current block can be determined as the current image reference; when the flag is false, the prediction mode of the current block can be determined as inter-frame prediction.
[0130] Alternatively, the predictive coding mode for the current block can be determined based on a reference image index. For example, when the reference image index points to the current image, the predictive coding mode for the current block can be determined as current image reference. When the reference image index points to another image instead of the current image, the predictive coding mode for the current block can be determined as inter-frame prediction. That is, current image reference is a prediction method that uses information from already encoded / decoded regions within the current image, and inter-frame prediction is a prediction method that uses information from other encoded / decoded images.
[0131] Combinatorial prediction represents a coding mode composed of two or more of intra-frame prediction, inter-frame prediction, and current image reference. For example, when applying combinatorial prediction, a first prediction block can be generated based on one of intra-frame prediction, inter-frame prediction, or the current image reference, and a second prediction block can be generated based on another. If a first and a second prediction block are generated, a final prediction block can be generated by averaging or weighted summing the first and second prediction blocks. Information indicating whether combinatorial prediction is applied can be transmitted via a signal in the bitstream. This information can be a 1-bit flag.
[0132] Figure 4 This is a diagram illustrating the various partitioning types of coded blocks.
[0133] A coded block can be divided into multiple coded blocks based on quadtree partitioning, binary tree partitioning, or ternary tree partitioning. Furthermore, the divided coded blocks can be further divided into multiple coded blocks based on quadtree partitioning, binary tree partitioning, or ternary tree partitioning.
[0134] Quadtree partitioning refers to a partitioning technique that divides the current block into four blocks. As a result of quadtree partitioning, the current block can be divided into four square partitions (see reference). Figure 4 ("SPLIT_QT(partition_quadtree)") in part (a).
[0135] Binary tree partitioning refers to a partitioning technique that divides the current block into two blocks. The process of partitioning the current block along a vertical direction (i.e., using a vertical line that crosses the current block) is called vertical binary tree partitioning, and the process of partitioning the current block along a horizontal direction (i.e., using a horizontal line that crosses the current block) is called horizontal binary tree partitioning. After binary tree partitioning, the current block can be divided into two non-square partitions. Figure 4 In part (b), "SPLIT_BT_VER(Partition_Binary_Vertical)" represents the result of a vertical binary tree partition, and Figure 4 In part (c), “SPLIT_BT_HOR(Partition_Binary_Level)” represents the result of a horizontal binary tree partition.
[0136] Ternary tree partitioning refers to a partitioning technique that divides the current block into three blocks. The process of dividing the current block into three blocks along a vertical direction (i.e., using two vertical lines crossing the current block) is called vertical ternary tree partitioning, and the process of dividing the current block into three blocks along a horizontal direction (i.e., using two horizontal lines crossing the current block) is called horizontal ternary tree partitioning. After ternary tree partitioning, the current block can be divided into three non-square partitions. In this case, the width / height of the partition located at the center of the current block can be twice the width / height of the other partitions. Figure 4In part (d), "SPLIT_TT_VER(Partition_Ternary_Vertical)" represents the result of a ternary tree partition in the vertical direction, and Figure 4 In part (e), “SPLIT_TT_HOR(Partition_Ternary_Horizontal)” represents the result of horizontal ternary tree partitioning.
[0137] The number of times a coding tree unit is divided can be defined as the partitioning depth. The maximum partitioning depth of a coding tree unit can be determined at the sequence or image level. Therefore, the maximum partitioning depth of a coding tree unit can vary depending on different sequences or images.
[0138] Alternatively, the maximum partitioning depth can be determined individually for each of the multiple partitioning techniques. For example, the maximum partitioning depth allowed for quadtree partitioning can be different from the maximum partitioning depth allowed for binary tree partitioning and / or ternary tree partitioning.
[0139] The encoder can transmit information via signals from the bitstream representing at least one of the partition shape or partition depth of the current block. The decoder can determine the partition shape and partition depth of the coding tree unit based on the information parsed from the bitstream.
[0140] Figure 5 This is a diagram illustrating an example of the division of a coding tree unit.
[0141] The process of dividing coding blocks using partitioning techniques such as quadtree partitioning, binary tree partitioning, and / or ternary tree partitioning is called multitree partitioning.
[0142] The coded blocks generated by applying a multi-way tree partitioning to the coded block can be called multiple downstream coded blocks. When the partitioning depth of the coded block is k, the partitioning depth of the multiple downstream coded blocks is set to k+1.
[0143] On the other hand, for multiple coding blocks with a partitioning depth of k+1, the coding block with a partitioning depth of k can be called the upstream coding block.
[0144] The partition type of the current coding block can be determined based on at least one of the partition shape of the upstream coding block or the partition type of the adjacent coding blocks. The adjacent coding blocks are adjacent to the current coding block and can include at least one of the current coding block's upper adjacent block, left adjacent block, or adjacent block to its upper left corner. The partition type can include at least one of whether to partition into a quadtree, whether to partition into a binary tree, the binary tree partition direction, whether to partition into a ternary tree, or the ternary tree partition direction.
[0145] To determine the shape of the coded block partition, information indicating whether the coded block has been partitioned can be sent via a signal in the bitstream. This information is a 1-bit flag "split_cu_flag," and when the flag is true, it indicates that the coded block has been partitioned using a multi-way tree partitioning technique.
[0146] When "split_cu_flag" is true, information indicating whether the coded block has been partitioned by a quadtree can be sent via a signal in the bitstream. This information is a 1-bit flag "split_qt_flag". When this flag is true, the coded block can be divided into 4 blocks.
[0147] For example, in Figure 5 The example shown illustrates how the coding tree unit is partitioned by a quadtree to generate four coding blocks with a partition depth of 1. Furthermore, the example illustrates applying quadtree partitioning again to the first and fourth coding blocks generated as a result of the quadtree partitioning. Ultimately, four coding blocks with a partition depth of 2 can be generated.
[0148] Furthermore, a coded block with a partition depth of 3 can be generated by applying a quadtree partition to the coded block with a partition depth of 2 again.
[0149] When a quadtree partitioning is not applied to the coded block, it can be determined whether to perform a binary tree partitioning or a ternary tree partitioning on the coded block by considering at least one of the following: the size of the coded block, whether the coded block is located at an image boundary, the maximum partitioning depth, or the partitioning shape of adjacent blocks. When it is determined whether to perform a binary tree partitioning or a ternary tree partitioning on the coded block, information indicating the partitioning direction can be transmitted via a signal in the bitstream. This information can be a 1-bit flag "mtt_split_cu_vertical_flag". The partitioning direction (vertical or horizontal) can be determined based on this flag. Additionally, information indicating whether a binary tree partitioning or a ternary tree partitioning is applied to the coded block can be transmitted via a signal in the bitstream. This information can be a 1-bit flag "mtt_split_cu_binary_flag". The binary tree partitioning or ternary tree partitioning can be determined based on this flag.
[0150] For example, in Figure 5 The example shown illustrates the application of a vertical binary tree partitioning to a coded block with a partitioning depth of 1, the application of a vertical ternary tree partitioning to the left coded block in the resulting coded block, and the application of a vertical binary tree partitioning to the right coded block.
[0151] When an image is divided into coding tree units, there may be blocks smaller than a predefined size in the region adjacent to the right or bottom boundary of the image. Assuming these blocks are coding tree units, coding tree units smaller than the predefined size can be generated at the right or bottom boundary of the image. In this case, the size of the coding tree unit can be determined based on information transmitted via signals through a sequence parameter set or an image parameter set.
[0152] Figure 6 This is a diagram showing an example of a block smaller than a coding tree unit of a preset size appearing at the image boundary.
[0153] like Figure 6 As shown in the example, when a 1292×1080 image is divided into 128×128 coding tree units, there will be blocks smaller than 128×128 at the right and bottom boundaries of the image. In the embodiments described later, the blocks of coding tree units smaller than the predefined size generated at the image boundaries are referred to as boundary blocks of atypical boundaries.
[0154] For boundary blocks with atypical boundaries, only predefined partitioning methods may be allowed. These predefined partitioning methods may include at least one of quadtree partitioning, ternary tree partitioning, or binary tree partitioning.
[0155] For example, for boundary blocks with atypical boundaries, only quadtree partitioning may be allowed. In this case, quadtree partitioning can be repeated until the image boundary block reaches the minimum quadtree partition size. The minimum quadtree partition size can be predefined in the encoder and decoder. Alternatively, information indicating the minimum quadtree partition size can be sent via signals in the bitstream.
[0156] Figure 7 This is a diagram illustrating an example of performing a quadtree partitioning on a boundary block with atypical boundaries. For ease of illustration, it is assumed that the minimum size of the quadtree partition is 4×4.
[0157] The partitioning of boundary blocks for atypical boundaries can be performed based on square blocks. These square blocks can be derived based on the larger of the width or height of the boundary block for the atypical boundary. For example, a power of 2 greater than and closest to the reference value can be considered the length of one side of the square block. Figure 7 The 12×20 block shown is considered to belong to the 32×32 block, and the 32×32 partition result can be applied to the 12×20 block.
[0158] When a 12×20 block is partitioned into a quadtree, the block can be divided into blocks of size 12×16 and blocks of size 12×4. If these blocks are then repartitioned into quadtrees, the 12×16 block is divided into two 8×8 blocks and two 4×8 blocks, and the 12×4 block is divided into two 8×4 blocks and two 4×4 blocks.
[0159] Dividing the 4×8 block located at the image boundary into a quadtree again results in two 4×4 blocks. Similarly, dividing the 8×4 block located at the image boundary into a quadtree again results in two 4×4 blocks.
[0160] Alternatively, binary tree partitioning can be performed when at least one of the block width or height is the same as or less than the minimum quadtree partition size. The minimum quadtree partition size can represent either the minimum quadtree partition width or the minimum quadtree partition height. For example, when the minimum quadtree partition size is 4×4, both the minimum quadtree partition width and the minimum quadtree partition height can be 4.
[0161] At this point, when the width of the blocks is the same as or less than the minimum quadtree partition size, a vertical binary tree partition can be performed; when the height of the blocks is the same as or less than the minimum quadtree partition height, a horizontal binary tree partition can be performed.
[0162] On the other hand, quadtree partitioning can be performed when the width or height of a block is greater than the minimum quadtree partitioning size. For example, when the upper right and lower left positions of a block are off-center from the image, and the width or height of the block is greater than the minimum quadtree partitioning size, quadtree partitioning can be applied to the corresponding block.
[0163] Figure 8 This is a diagram illustrating an example of performing a quadtree partition on blocks that are simultaneously adjacent to the right and bottom boundaries of the image. For ease of illustration, it is assumed that the minimum size of the quadtree partition is 4×4.
[0164] In such Figure 8 In the example shown, when a quadtree partition is performed on a 32×32 block containing a 12×20 block, four 16×16 blocks are generated. The two 16×16 blocks containing texture data can be partitioned again using a quadtree. As a result, the x-axis and y-axis coordinates deviate from the image boundaries, generating an 8×8 block containing 4×4 texture data. Since the width and height of the 8×8 block are greater than the minimum quadtree partition size, a quadtree partition can be performed on this block.
[0165] When a 12×20 block is partitioned into a quadtree, the block can be divided into blocks of size 12×16 and blocks of size 12×4. If these blocks are then repartitioned into quadtrees, the 12×16 block is divided into two 8×8 blocks and two 4×8 blocks, and the 12×4 block is divided into two 8×4 blocks and two 4×4 blocks.
[0166] Since the width of the 4×8 block located at the image boundary is the same as the minimum quadtree partition size, a binary tree partition can be performed on the 4×8 block. Specifically, a vertical binary tree partition can be performed based on a square block containing the 4×8 block (i.e., 8×8).
[0167] Furthermore, since the width of the 8×4 block located at the image boundary is the same as the minimum quadtree partition size, binary tree partitioning can be performed on the 8×4 block. Specifically, horizontal binary tree partitioning can be performed based on a square block containing 4×8 blocks (i.e., 8×8).
[0168] The partitioning results show that blocks of size 8×4, 4×4, and 4×8 can be set adjacent to the image boundary.
[0169] Alternatively, a binary tree partition can be performed if at least one of the block's width or height is less than or equal to a threshold; otherwise, a quadtree partition can be performed. The threshold can be derived based on the minimum quadtree partition size. For example, when the minimum quadtree partition size is minQTsize, the threshold can be set to "minQTsize<<1". Alternatively, information for determining the threshold can be sent via a signal in the bitstream.
[0170] Figure 9 This is a diagram illustrating the partitioning pattern of blocks adjacent to the image boundaries. For ease of illustration, it is assumed that the minimum quadtree partition size is 4×4. The threshold can be set to 8.
[0171] First, a quadtree partitioning can be performed on the 12×20 block. As a result, the block can be divided into 12×16 blocks and 12×4 blocks. Since the width and height of the 12×16 block are both greater than a threshold, quadtree partitioning can be applied to this block. Thus, the block can be divided into two 8×8 blocks and two 4×8 blocks.
[0172] The width of a 12×4 block is greater than a threshold. Therefore, a quadtree partitioning can be applied to the 12×4 block. As a result, the block can be divided into 8×4 blocks and 4×4 blocks.
[0173] Subsequently, since the width and height of the 4×8 block and the 8×4 block located at the image boundary are both less than or equal to the threshold, a binary tree partition can be applied to these blocks.
[0174] Contrary to the example above, a binary tree partition can be performed when at least one of the width or height of the block is greater than the threshold; otherwise, a quadtree partition can be performed.
[0175] Alternatively, for boundary blocks with atypical boundaries, quadtree partitioning or binary tree partitioning can be applied only. For example, quadtree partitioning can be repeated until the block located at the image boundary has the minimum size, or binary tree partitioning can be repeated until the block located at the image boundary has the minimum size.
[0176] Boundary blocks with atypical boundaries can be designated as coding units, and a skip mode can be applied to these blocks in a fixed manner, or all transform coefficients can be set to 0. Therefore, the Coded Block Flag (CBF), which indicates whether a boundary block with atypical boundaries has non-zero transform coefficients, can be set to 0. Coding units encoded in skip mode or with transform coefficients set to 0 can be called boundary zero coding units.
[0177] Alternatively, it can be determined whether to set the coding unit as a boundary zero coding unit by comparing at least one of the width or height of the coding block generated by dividing the boundary block of atypical boundaries with a threshold. For example, coding blocks whose width or height is less than the threshold can be encoded in a skip mode, or the transform coefficients of the coding block can be set to 0.
[0178] Figure 10 This is a graph showing the encoding patterns of blocks adjacent to the image boundaries. Assume a threshold of 8.
[0179] A coding block whose width or height is less than a threshold can be set as a boundary zero coding unit in a coding block generated by dividing atypical boundary blocks.
[0180] For example, in such Figure 10 In the example shown, a 4×16 coded block, an 8×4 coded block, and a 4×4 coded block can be set as boundary zero coded units. Therefore, the block can be encoded in a skip mode or the transform coefficients of the block can be set to 0.
[0181] For coded blocks whose width and height are equal to or greater than a threshold, a skip mode can be selectively applied or the transform coefficients can be set to 0. To this end, a flag indicating whether to apply a skip mode to the coded block or a flag indicating whether to set the transform coefficients to 0 can be encoded and transmitted as a signal.
[0182] Alternatively, only coding units generated by binary tree partitioning may be allowed to be set as boundary zero coding units. Alternatively, only coding units generated by quadtree partitioning may be allowed to be set as boundary zero coding units.
[0183] Inter-frame prediction refers to using information from prior images to predict the predictive coding mode of the current block. For example, a block in a prior image that is at the same position as the current block (hereinafter referred to as a collocated block) can be set as the predicted block for the current block. Hereinafter, the predicted block generated based on the block at the same position as the current block will be called a collocated prediction block.
[0184] On the other hand, if an object existing in a prior image has moved to a different location in the current image, the object's motion can be used to effectively predict the current block. For example, if the direction and size of the object's movement can be known by comparing the prior and current images, the object's motion information can be considered to generate a predicted block (or predicted image) for the current block. Hereinafter, the predicted block generated using motion information can be referred to as a motion prediction block.
[0185] Residual blocks can be generated by subtracting prediction blocks from the current block. In this case, when there is motion of the object, motion prediction blocks can be used instead of co-position prediction blocks, thereby reducing the energy of the residual blocks and improving their compression performance.
[0186] As mentioned above, the process of generating prediction blocks using motion information can be called motion-compensated prediction. In most inter-frame predictions, prediction blocks can be generated based on motion-compensated prediction.
[0187] Motion information may include at least one of motion vectors, reference image indices, prediction directions, or bidirectional weighted indexes. Motion vectors represent the direction and magnitude of an object's movement. Reference image indices specify the reference image for the current block among a list of reference images. Prediction directions refer to any one of unidirectional L0 prediction, unidirectional L1 prediction, or bidirectional prediction (L0 and L1 prediction). Motion information in either the L0 or L1 direction can be used based on the prediction direction of the current block. Bidirectional weighted indexes specify the weights applied to the L0 prediction block and the weights applied to the L1 prediction block.
[0188] Figure 11 This is a flowchart of the inter-frame prediction method according to an embodiment of the present invention.
[0189] Reference Figure 11 The inter-frame prediction method includes: determining the inter-frame prediction mode of the current block (S1101); obtaining motion information of the current block according to the determined inter-frame prediction mode (S1102); and performing motion compensation prediction of the current block based on the obtained motion information (S1103).
[0190] Inter-frame prediction modes represent various techniques used to determine the motion information of the current block, and can include inter-frame prediction modes using translational motion information and inter-frame prediction modes using affine motion information. For example, inter-frame prediction modes using translational motion information can include merging mode and advanced motion vector prediction mode, while inter-frame prediction modes using affine motion information can include affine merging mode and affine motion vector prediction mode. Based on the inter-frame prediction mode, the motion information of the current block can be determined based on neighboring blocks adjacent to the current block or information parsed from the bitstream.
[0191] The following section will describe in detail the inter-frame prediction method that uses translational motion information.
[0192] Motion information for the current block can be derived from the motion information of other blocks. These other blocks can be those that are prioritized for inter-frame prediction encoding / decoding compared to the current block. Setting the motion information of the current block to be the same as that of other blocks is defined as a merging mode. Furthermore, setting the motion vectors of other blocks to the predicted values of the motion vectors of the current block is defined as a motion vector prediction mode.
[0193] Figure 12 This is a flowchart of the process of exporting motion information of the current block in merge mode.
[0194] Merging candidates for the current block can be exported (S1201). Merging candidates for the current block can be derived from blocks that were encoded / decoded using inter-frame prediction before the current block.
[0195] Figure 13 This is a diagram showing the candidate blocks used to derive the merge candidates.
[0196] Candidate blocks can include at least one of the following: neighboring blocks containing samples adjacent to the current block, or non-neighboring blocks containing samples not adjacent to the current block. Hereinafter, the samples used to determine candidate blocks will be designated as reference samples. Furthermore, reference samples adjacent to the current block will be referred to as neighboring reference samples, and reference samples not adjacent to the current block will be referred to as non-neighboring reference samples.
[0197] Adjacent reference samples can be included in the adjacent column of the leftmost column of the current block or the adjacent row of the topmost row of the current block. For example, if the coordinates of the top-left sample of the current block are (0, 0), then at least one of the following blocks—a block including a reference sample at position (-1, H-1), a block including a reference sample at position (W-1, -1), a block including a reference sample at position (W, -1), a block including a reference sample at position (-1, H), or a block including a reference sample at position (-1, -1)—can be used as candidate blocks. Referring to the accompanying drawings, adjacent blocks with indices 0 to 4 can be used as candidate blocks.
[0198] A non-adjacent reference sample refers to a sample whose x-axis distance or y-axis distance to the reference sample adjacent to the current block has a predefined value. For example, a block containing a reference sample whose x-axis distance to the left reference sample is a predefined value, a block containing a non-adjacent sample whose y-axis distance to the upper reference sample is a predefined value, or a block containing non-adjacent samples whose x-axis and y-axis distances to the upper-left reference sample are both predefined values can be used as a candidate block. The predefined value can be an integer such as 4, 8, 12, 16, etc. Referring to the accompanying drawings, at least one of the blocks with indices from 5 to 26 can be used as a candidate block.
[0199] Samples that are not on the same vertical, horizontal, or diagonal line as adjacent reference samples can be set as non-adjacent reference samples.
[0200] Figure 14 This is a diagram showing the location of the reference sample.
[0201] like Figure 14 In the example shown, the x-coordinate of the upper non-adjacent reference sample can be set differently from the x-coordinate of the upper adjacent reference sample. For example, when the position of the upper adjacent reference sample is (W-1, -1), the position of the upper non-adjacent reference sample that is N away from the upper adjacent reference sample along the y-axis can be set to ((W / 2)-1, -1-N), and the position of the upper non-adjacent reference sample that is 2N away from the upper adjacent reference sample along the y-axis can be set to (0, -1-2N). That is, the position of the non-adjacent reference sample can be determined based on the position of the adjacent reference sample and the distance between the adjacent reference samples.
[0202] In the following text, a candidate block containing an adjacent reference sample is called a neighboring block, and a block containing a non-adjacent reference sample is called a non-adjacent block.
[0203] When the distance between the current block and a candidate block is greater than or equal to a threshold, the candidate block can be set as unusable as a merging candidate. The threshold can be determined based on the size of the coding tree unit. For example, the threshold can be set to the height of the coding tree unit (ctu_height), or the height of the coding tree unit plus or minus an offset value (e.g., ctu_height ± N). The offset value N is a predefined value in the encoder and decoder, and can be set to 4, 8, 16, 32, or ctu_height.
[0204] If the difference between the y-axis coordinate of the current block and the y-axis coordinate of the samples included in the candidate block is greater than a threshold, the candidate block can be determined as unsuitable for merging.
[0205] Alternatively, candidate blocks that do not belong to the same coding tree unit as the current block can be set as unsuitable for merging. For example, when the reference sample exceeds the upper boundary of the coding tree unit to which the current block belongs, candidate blocks that include the reference sample can be set as unsuitable for merging.
[0206] If the upper boundary of the current block is adjacent to the upper boundary of a coding tree unit, multiple candidate blocks will be determined as unsuitable for merging, which will reduce the encoding / decoding efficiency of the current block. To resolve this issue, candidate blocks can be configured such that the number of candidate blocks above the current block is greater than the number of candidate blocks to the left of the current block.
[0207] Figure 15 This is a diagram showing the candidate blocks used to derive the merge candidates.
[0208] As in Figure 15 In the example shown, the block above the current block (N blocks above it) and the block to the left of the current block (M blocks to its left) can be set as candidate blocks. In this case, by setting M to be greater than N, the number of left candidate blocks can be set to be greater than the number of top candidate blocks.
[0209] For example, the difference between the y-axis coordinate of the reference sample within the current block and the y-axis coordinate of the block above which can be used as a candidate block can be set to no more than N times the height of the current block. Additionally, the difference between the x-axis coordinate of the reference sample within the current block and the x-axis coordinate of the block to the left of which can be used as a candidate block can be set to no more than M times the width of the current block.
[0210] For example, in Figure 15 The example shown illustrates setting blocks belonging to the two blocks above the current block and blocks belonging to the five blocks to the left of the current block as candidate blocks.
[0211] As another example, when a candidate block does not belong to the same coding tree unit as the current block, a merge candidate can be derived by using a block that belongs to the same coding tree unit as the current block, or a block that contains a reference sample adjacent to the boundary of the coding tree unit, instead of the candidate block.
[0212] Figure 16 This is a diagram illustrating an example of changing the position of a reference sample.
[0213] When a reference sample is included in a coding tree unit that is different from the current block, and the reference sample is not adjacent to the boundary of the coding tree unit, a candidate block reference sample can be determined by using a reference sample adjacent to the boundary of the coding tree unit instead of the reference sample.
[0214] For example, in Figure 16 (a) and Figure 16 In the example shown in (b), when the upper boundary of the current block touches the upper boundary of the coding tree unit, the reference sample above the current block belongs to a coding tree unit different from the current block. A reference sample belonging to a coding tree unit different from the current block that is not adjacent to the upper boundary of the coding tree unit can be replaced with a sample adjacent to the upper boundary of the coding tree unit.
[0215] For example, such as Figure 16 In the example shown in (a), the reference sample at position 6 is replaced with the sample at position 6' located at the upper boundary of the coding tree unit, as follows: Figure 16 As shown in example (b), the reference sample at position 15 is replaced with the sample at position 15', which is located at the upper boundary of the coding tree unit. In this case, the y-coordinate of the replacement sample can be changed to that of an adjacent position in the coding tree unit, and the x-coordinate of the replacement sample can be set to be the same as that of the reference sample. For example, the sample at position 6' can have the same x-coordinate as the sample at position 6, and the sample at position 15' can have the same x-coordinate as the sample at position 15.
[0216] Alternatively, the x-coordinate of the replacement sample can be set by adding or subtracting the offset value from the x-coordinate of the reference sample. For example, when the x-coordinates of adjacent and non-adjacent reference samples above the current block are the same, the x-coordinate of the replacement sample can be set by adding or subtracting the offset value from the x-coordinate of the reference sample. This is to prevent the replacement sample used to replace a non-adjacent reference sample from being in the same position as other non-adjacent or adjacent reference samples.
[0217] Figure 17 This is a diagram illustrating an example of changing the position of a reference sample.
[0218] When replacing a reference sample that is included in a different coding tree unit than the current block and is not adjacent to the boundary of the coding tree unit with a sample located at the boundary of the coding tree unit, the x-coordinate of the replacement sample can be set by adding or subtracting the offset value to the x-coordinate of the reference sample.
[0219] For example, in Figure 17 In the example shown, the reference sample at position 6 and the reference sample at position 15 can be replaced with the sample at position 6' and the sample at position 15', respectively, having the same y-coordinate as the row adjacent to the upper boundary of the coding tree unit. In this case, the x-coordinate of the sample at position 6' can be set to a value where the difference between its x-coordinate and the x-coordinate of the reference sample at position 6 is W / 2, and the x-coordinate of the sample at position 15' can be set to a value where the difference between its x-coordinate and the x-coordinate of the reference sample at position 15 is W-1.
[0220] Unlike Figure 16 and Figure 17 In the example shown, the y-coordinate of the row above the top row of the current block or the y-coordinate of the upper boundary of the coding tree unit can also be set to the y-coordinate of the replacement sample.
[0221] Although not illustrated, the sample replacing the reference sample can also be determined based on the left boundary of the coding tree unit. For example, when the reference sample is not included in the same coding tree unit as the current block and is not adjacent to the left boundary of the coding tree unit, the reference sample can be replaced with a sample adjacent to the left boundary of the coding tree unit. In this case, the replacement sample can have the same y-coordinate as the reference sample, or it can have a y-coordinate obtained by adding or subtracting an offset value to the y-coordinate of the reference sample.
[0222] Then, the block containing the replacement sample can be set as a candidate block, and the merge candidates for the current block can be derived based on the candidate blocks.
[0223] Merge candidates can also be derived from temporally adjacent blocks included in images different from the current block. For example, merge candidates can be derived from co-located blocks included in co-located images.
[0224] The motion information of the merged candidate can be set to be the same as that of the candidate block. For example, at least one of the motion vector, reference image index, prediction direction, or bidirectional weighted index of the candidate block can be set as the motion information of the merged candidate.
[0225] A list of merge candidates, including merge candidates, can be generated (S1202). The merge candidates can be classified into adjacent merge candidates derived from adjacent blocks adjacent to the current block, and non-adjacent merge candidates derived from non-adjacent blocks.
[0226] The indices of multiple merge candidates within the merge candidate list can be assigned in a predetermined order. For example, the index assigned to an adjacent merge candidate can have a smaller value than the index assigned to a non-adjacent merge candidate. Alternatively, based on Figure 13 or Figure 15 The index shown for each block can be assigned to each merge candidate.
[0227] When the merge candidate list includes multiple merge candidates, at least one of the multiple merge candidates can be selected (S1203). At this time, information indicating whether the motion information of the current block is derived from adjacent merge candidates can be sent via a signal in the code stream. The information can be a 1-bit flag. For example, the syntax element isAdjancentMergeFlag indicating whether the motion information of the current block is derived from adjacent merge candidates can be sent via a signal in the code stream. When the value of the syntax element isAdjancentMergeFlag is 1, the motion information of the current block can be derived based on adjacent merge candidates. On the other hand, when the value of the syntax element isAdjancentMergeFlag is 0, the motion information of the current block can be derived based on non-adjacent merge candidates.
[0228] Table 1 shows the syntax table including the syntax element isAdjancentMergeFlag.
[0229] Table 1
[0230]
[0231] Information specifying any one of multiple merge candidates can be sent via signals in the bitstream. For example, information indicating the index of any merge candidate included in the merge candidate list can be sent via signals in the bitstream.
[0232] When isAdjacentMergeflag is 1, the syntax element merge_idx can be signaled to determine which of the adjacent merge candidates is being merged. The maximum value of the syntax element merge_idx can be set to a value that is 1 greater than the difference between the number of adjacent merge candidates.
[0233] When isAdjacentMergeflag is 0, the syntax element NA_merge_idx can be signaled to determine any of the non-adjacent merge candidates. The syntax element NA_merge_idx indicates the value obtained by subtracting the index of the non-adjacent merge candidate from the number of adjacent merge candidates. The decoder can select a non-adjacent merge candidate by adding the number of adjacent merge candidates to the index determined by NA_merge_idx.
[0234] When the number of merge candidates included in the merge candidate list is less than the maximum value, merge candidates included in the inter-frame motion information list can be added to the merge candidate list. The inter-frame motion information list can include merge candidates derived based on blocks encoded / decoded before the current block.
[0235] The inter-frame motion information list includes merging candidates derived from blocks encoded / decoded based on inter-frame prediction within the current image. For example, the motion information of the merging candidates included in the inter-frame motion information list can be set to be the same as the motion information of the blocks encoded / decoded based on inter-frame prediction. The motion information may include at least one of motion vectors, reference image indexes, prediction directions, or bidirectional weighted indexes.
[0236] For ease of explanation, the merging candidates included in the inter-frame motion information list are referred to as inter-frame merging candidates.
[0237] The maximum number of merge candidates that can be included in the inter-frame motion information list can be predefined in the encoder and decoder. For example, the maximum number of merge candidates that can be included in the inter-frame motion information list can be 1, 2, 3, 4, 5, 6, 7, 8 or greater (e.g., 16).
[0238] Alternatively, information representing the maximum number of merge candidates in the inter-frame motion information list can be transmitted via a signal in the bitstream. This information can be transmitted at the sequence level, image level, or strip level.
[0239] Alternatively, the maximum number of merged candidates for the inter-frame motion information list can be determined based on the image size, the strip size, or the size of the coding tree unit.
[0240] The inter-frame motion information list can be initialized at the image, strip, tile, brick, coding tree unit, or coding tree unit line (row or column) level. For example, the inter-frame motion information list is also initialized during strip initialization, and it may not include any merge candidates.
[0241] Alternatively, information indicating whether to initialize the inter-frame motion information list can be sent via a signal in the bitstream. This information can be sent at the strip, tile, brick, or block level. A pre-configured inter-frame motion information list can be used before the information indicates initialization of the inter-frame motion information list.
[0242] Alternatively, information related to inter-frame merge candidates can be signaled via an image parameter set or a strip header. Even if the strip is initialized, the inter-frame motion information list can include initial inter-frame merge candidates. Thus, inter-frame merge candidates can be used for the first block encoded / decoded within a strip.
[0243] The blocks are encoded / decoded according to the encoding / decoding order, and multiple blocks encoded / decoded based on inter-frame prediction can be set as inter-frame merging candidates in sequence according to the encoding / decoding order.
[0244] Figure 18 This is a diagram illustrating an example of updating the list of motion information between frames.
[0245] When performing inter-frame prediction on the current block (S1801), inter-frame merging candidates can be derived based on the current block (S1802). The motion information of the inter-frame merging candidates can be set to be the same as the motion information of the current block.
[0246] When the inter-frame motion information list is empty (S1803), inter-frame merge candidates derived from the current block can be added to the inter-frame motion information list (S1804).
[0247] When the inter-frame merge candidate is already included in the inter-frame motion information list (S1803), a redundancy check can be performed on the motion information of the current block (or the inter-frame merge candidate derived from the current block) (S1805). The redundancy check is used to determine whether the motion information of the inter-frame merge candidates stored in the inter-frame motion information list is the same as the motion information of the current block. Redundancy checks can be performed on all inter-frame merge candidates stored in the inter-frame motion information list. Alternatively, redundancy checks can be performed on inter-frame merge candidates whose index is above or below a threshold among the inter-frame merge candidates stored in the inter-frame motion information list.
[0248] If inter-frame prediction merging candidates with the same motion information as the current block are not included, inter-frame merging candidates derived from the current block can be added to the inter-frame motion information list (S1808). Whether inter-frame merging candidates are the same can be determined based on whether the motion information (e.g., motion vectors and / or reference image indexes, etc.) of the inter-frame merging candidates are the same.
[0249] At this point, when the maximum number of inter-frame merge candidates has been stored in the inter-frame motion information list (S1806), the earliest inter-frame merge candidate is deleted (S1807), and inter-frame merge candidates derived based on the current block can be added to the inter-frame motion information list (S1808).
[0250] Multiple inter-frame merge candidates can be identified based on their indices. When adding an inter-frame merge candidate derived from the current block to the inter-frame motion information list, the candidate is assigned the lowest index (e.g., 0), and the indices of already stored inter-frame merge candidates can be incremented by 1. In this case, when the maximum number of inter-frame merge candidates is stored in the inter-frame motion information list, the candidate with the highest index is removed.
[0251] Alternatively, when adding inter-frame merge candidates derived from the current block to the inter-frame motion information list, the largest index can be assigned to the inter-frame merge candidate. For example, when the number of inter-frame prediction merge candidates already stored in the inter-frame motion information list is less than the maximum value, an index with the same value as the number of stored inter-frame prediction merge candidates can be assigned to the inter-frame merge candidate. Alternatively, when the number of inter-frame merge candidates already stored in the inter-frame motion information list is equal to the maximum value, an index that is 1 less than the maximum value can be assigned to the inter-frame merge candidate. Furthermore, the inter-frame merge candidate with the smallest index is removed, and the indices of the remaining stored inter-frame merge candidates are each reduced by 1.
[0252] Figure 19 This is a diagram illustrating an example of an updated inter-frame merge candidate list.
[0253] Assume that inter-frame merge candidates derived from the current block are added to the inter-frame merge candidate list, and the largest index is assigned to the inter-frame merge candidate. Also, assume that the inter-frame merge candidate list already stores the maximum number of inter-frame merge candidates.
[0254] When adding the inter-frame merge candidate HmvpCand[n+1] exported from the current block to the inter-frame merge candidate list HmvpCandList, the inter-frame merge candidate HmvpCand[0] with the smallest index is removed from the stored inter-frame merge candidates, and the indices of the remaining inter-frame merge candidates are decreased by 1 respectively. Alternatively, the index of the inter-frame merge candidate HmvpCand[n+1] exported from the current block can be set to the maximum value (in... Figure 19 In the example shown, n).
[0255] If an inter-frame merge candidate that is the same as the inter-frame merge candidate derived from the current block is already stored (S1805), the inter-frame merge candidate derived from the current block may not be added to the inter-frame motion information list (S1809).
[0256] Alternatively, as inter-frame merge candidates derived from the current block are added to the inter-frame motion information list, previously stored inter-frame merge candidates that are identical to those candidates can also be removed. In this case, the indexes of the previously stored inter-frame merge candidates will be updated.
[0257] Figure 20 This is a diagram showing an example of how the index of a stored inter-frame merge candidate is updated.
[0258] When the index of a stored inter-frame merge candidate that is the same as the inter-frame merge candidate mvCand derived based on the current block is hIdx, deleting the stored inter-frame merge candidate can reduce the index of each inter-frame merge candidate with an index greater than hIdx by 1. For example, in Figure 20 The example shown illustrates removing HmvpCand[2], which is identical to mvCand, from the inter-frame motion information list HvmpCandList, and reducing the indices of HmvpCand[3] to HmvpCand[n] by 1.
[0259] Furthermore, inter-frame merge candidate mvCands derived based on the current block can be added to the end of the inter-frame motion information list.
[0260] Alternatively, the index of a stored inter-frame merge candidate that is assigned to the same inter-frame merge candidate derived based on the current block can be updated. For example, the index of a stored inter-frame merge candidate can be changed to the minimum or maximum value.
[0261] Motion information of blocks included in a predetermined region can be set to not be added to the inter-frame motion information list. For example, inter-frame merge candidates derived based on motion information of blocks included in a parallel merge region can be excluded from the inter-frame motion information list. Since the encoding / decoding order of the blocks included in the parallel merge region is not specified, it is inappropriate to use the motion information of any of the blocks when performing inter-frame prediction on other blocks. Therefore, inter-frame merge candidates derived based on blocks included in the parallel merge region can be excluded from the inter-frame motion information list.
[0262] When performing motion compensation prediction using sub-block units, inter-frame merging candidates can be derived from the motion information of representative sub-blocks within the current block. For example, when using sub-block merging candidates for the current block, inter-frame merging candidates can be derived from the motion information of representative sub-blocks within the sub-block.
[0263] The motion vector of a sub-block can be derived in the following order. First, any of the merge candidates included in the merge candidate list of the current block can be selected, and the initial shift vector (shVector) can be derived based on the motion vector of the selected merge candidate. Then, by adding the initial shift vector to the positions (xSb, ySb) of the reference samples (e.g., the top-left sample or the middle sample) of each sub-block within the coded block, a shifted sub-block with reference sample positions (xColSb, yColSb) can be derived. Equation 1 below shows the equation used to derive the shifted sub-block.
[0264] Equation 1
[0265] (xColSb, yColSb) = (xSb+shVecror[0]>>4, ySb+shVector[I]>>4)
[0266] Next, the motion vector of the co-position block corresponding to the center position of the sub-block including (xColSb, yColSb) is set as the motion vector of the sub-block including (xSb, ySb).
[0267] A representative sub-block can mean a sub-block that includes the top-left sample or the center sample of the current block.
[0268] Figure 21 This is a diagram showing the location of a representative sub-block.
[0269] Figure 21 (a) shows an example of setting the child block located to the upper left of the current block as the representative child block. Figure 21 (b) shows an example of setting the sub-block located at the center of the current block as the representative sub-block. When performing motion compensation prediction on a sub-block basis, inter-frame merge candidates for the current block can be derived based on the motion vectors of sub-blocks that include the top-left sample of the current block or sub-blocks that include the center sample of the current block.
[0270] Based on the inter-frame prediction mode of the current block, it can also be determined whether the current block should be used as an inter-frame merging candidate. For example, blocks encoded / decoded based on an affine motion model can be set as non-inter-frame merging candidates. Thus, even if the current block is encoded / decoded using inter-frame prediction, the inter-frame prediction motion information list will not be updated based on the current block if the current block's inter-frame prediction mode is affine prediction mode.
[0271] Alternatively, inter-frame merge candidates can be derived from at least one sub-block vector within the sub-blocks included in the block being encoded / decoded based on an affine motion model. For example, an inter-frame merge candidate can be derived using a sub-block located to the upper left, center, or upper right of the current block. Alternatively, the average of the sub-block vectors of multiple sub-blocks can be used as the motion vector for the inter-frame merge candidate.
[0272] Alternatively, inter-frame merge candidates can be derived based on the average of the affine seed vectors of the blocks encoded / decoded using an affine motion model. For example, the average of at least one of the first, second, or third affine seed vectors of the current block can be set as the motion vector of the inter-frame merge candidate.
[0273] Alternatively, the inter-frame motion information list can be configured for different inter-frame prediction modes. For example, at least one of the following can be defined: an inter-frame motion information list for blocks encoded / decoded via intra-block copying, an inter-frame motion information list for blocks encoded / decoded based on a translational motion model, or an inter-frame motion information list for blocks encoded / decoded based on an affine motion model. Any one of the multiple inter-frame motion information lists can be selected depending on the inter-frame prediction mode of the current block.
[0274] Figure 22 An example of generating a list of inter-frame motion information for different inter-frame prediction modes is shown.
[0275] When encoding / decoding a block based on a non-affine motion model, inter-frame merge candidate `mvCand` derived from the block can be added to the inter-frame non-affine motion information list `HmvpCandList`. Conversely, when encoding / decoding a block based on an affine motion model, inter-frame merge candidate `mvAfCand` derived from the block can be added to the inter-frame affine motion information list `HmvpAfCandList`.
[0276] The affine seed vector of a block can be stored in an inter-frame merge candidate derived from the block encoded / decoded based on the affine motion model. Thus, the inter-frame merge candidate can be used as a merge candidate for deriving the affine seed vector of the current block.
[0277] In addition to the described list of inter-frame motion information, another list of inter-frame motion information can be defined. Besides the described list of inter-frame motion information (hereinafter referred to as the first inter-frame motion information list), a long-term motion information list (hereinafter referred to as the second inter-frame motion information list) can also be defined. The long-term motion information list includes long-term merging candidates.
[0278] When both the first and second inter-frame motion information lists are empty, inter-frame merge candidates can be added to the second inter-frame motion information list first. Only after the maximum number of available inter-frame merge candidates in the second inter-frame motion information list has been reached can inter-frame merge candidates be added to the first inter-frame motion information list.
[0279] Alternatively, an inter-frame merge candidate can be added to both the second inter-frame motion information list and the first inter-frame motion information list.
[0280] In this case, the already configured second inter-frame motion information list may no longer be updated. Alternatively, the second inter-frame motion information list may be updated when the decoded region is above a predetermined ratio of the stripes. Alternatively, the second inter-frame motion information list may be updated every N coding tree unit rows.
[0281] On the other hand, the first inter-frame motion information list can be updated whenever a block is generated using inter-frame prediction for encoding / decoding. However, inter-frame merge candidates added to the second inter-frame motion information list can also be set not to be used to update the first inter-frame motion information list.
[0282] Information for selecting either a first inter-frame motion information list or a second inter-frame motion information list can be transmitted via a signal in the bitstream. When the number of merge candidates included in the merge candidate list is less than the maximum value, the merge candidates included in the inter-frame motion information list indicated by the information can be added to the merge candidate list.
[0283] Alternatively, the list of inter-frame motion information can be selected based on the size and shape of the current block, the inter-frame prediction mode, whether bidirectional prediction is enabled or disabled, whether motion vectors are refined or not, or whether triangulation is enabled or not.
[0284] Alternatively, if the number of merge candidates included in the merge candidate list is still less than the maximum number of merges even after adding the inter-frame merge candidates included in the first inter-frame motion information list, then the inter-frame merge candidates included in the second inter-frame motion information list can be added to the merge candidate list.
[0285] Figure 23 This is a diagram illustrating an example of adding inter-frame merge candidates included in the long-term motion information list to the merge candidate list.
[0286] If the number of merge candidates in the merge candidate list is less than the maximum number, inter-frame merge candidates included in the first inter-frame motion information list HmvpCandList can be added to the merge candidate list. Even if the number of merge candidates in the merge candidate list is still less than the maximum number after adding inter-frame merge candidates included in the first inter-frame motion information list, then inter-frame merge candidates included in the long-term motion information list HmvpLTCandList can be added to the merge candidate list.
[0287] Table 2 illustrates the process of adding inter-frame merging candidates, which are included in the long-term motion information list, to the merge candidate list.
[0288] Table 2
[0289]
[0290] Inter-frame merge candidates can be configured to include additional information besides motion information. For example, the size, shape, or partitioning information of storage blocks can be added to the inter-frame merge candidates. When constructing the merge candidate list for the current block, only inter-frame merge candidates with the same or similar size, shape, or partitioning information as the current block are used in the inter-frame merge candidate list, or inter-frame merge candidates with the same or similar size, shape, or partitioning information as the current block are preferentially added to the merge candidate list.
[0291] Alternatively, inter-frame motion information lists can be generated for different block sizes, shapes, or partitioning information. Multiple inter-frame motion information lists corresponding to the shape, size, or partitioning information of the current block can be used to generate a merge candidate list for the current block.
[0292] Alternatively, inter-frame motion information lists can be generated based on different resolutions of the motion vectors. For example, when the motion vector of the current block has a resolution of 1 / 4 pixel, inter-frame merge candidates derived from the current block can be added to the inter-frame region 1 / 4 pixel motion information list. When the motion vector of the current block has a resolution of 1 integer pixel, inter-frame merge candidates derived from the current block can be added to the inter-frame region integer pixel motion information list. When the motion vector of the current block has a resolution of 4 integer pixels, inter-frame merge candidates derived from the current block can be added to the inter-frame region 4 integer pixel motion information list. One of multiple inter-frame motion information lists can be selected based on the motion vector resolution of the encoded / decoded object block.
[0293] When applying the merge offset vector coding method to the current block, inter-frame merge candidates derived from the current block can be added to the inter-frame region merge offset motion information list HmvpHMVDCandList instead of the inter-frame motion information list HmvpCandList. In this case, the inter-frame merge candidates can include the motion vector offset information of the current block. HmvpHMVDCandList can be used to derive the offset of the block to which the merge offset vector coding method is applied.
[0294] If the number of merge candidates included in the current block's merge candidate list is not the maximum value, inter-frame merge candidates included in the inter-frame motion information list can be added to the merge candidate list. The addition process is performed in ascending or descending order of the index. For example, the inter-frame merge candidate with the largest index can be added to the merge candidate list.
[0295] When adding inter-frame merge candidates included in the inter-frame motion information list to the merge candidate list, a redundancy check can be performed between the inter-frame merge candidates and multiple merge candidates already stored in the merge candidate list.
[0296] For example, Table 3 shows the process of adding inter-frame merge candidates to the merge candidate list.
[0297] Table 3
[0298]
[0299] Redundancy checks can also be performed only on some of the inter-frame merge candidates included in the inter-frame motion information list. For example, redundancy checks can be performed only on inter-frame merge candidates whose index is above or below a threshold.
[0300] Alternatively, redundancy checks can be performed only on some of the merge candidates already stored in the merge candidate list. For example, redundancy checks can be performed only on merge candidates with an index above or below a threshold, or on merge candidates derived from a block at a specific location. A specific location may include at least one of the current block's left neighbor, top neighbor, top-right neighbor, or bottom-left neighbor.
[0301] Figure 24 This is a diagram illustrating an example of performing redundancy checks only on some of the merge candidates.
[0302] When adding an inter-frame merge candidate HmvpCand[j] to the merge candidate list, a redundancy check can be performed between the inter-frame merge candidate and the two merge candidates with the largest indices, mergeCandList[NumMerge-2] and mergeCandList[NumMerge-1]. Here, NumMerge represents the number of available spatial and temporal merge candidates.
[0303] If a merge candidate that is identical to the first inter-frame merge candidate is found, the redundancy check of the merge candidate that is identical to the first inter-frame merge candidate can be omitted when performing a redundancy check on the second inter-frame merge candidate.
[0304] Figure 25 This is a diagram illustrating an example of omitting redundancy checks for a specific merge candidate.
[0305] When adding the inter-frame merge candidate HmvpCand[i] at index i to the merge candidate list, a redundancy check can be performed between the inter-frame merge candidate and the merge candidates already stored in the merge candidate list. In this case, if a merge candidate mergeCandList[j] identical to the inter-frame merge candidate HmvpCand[i] is found, the inter-frame merge candidate HmvpCand[i] will not be added to the merge candidate list, and a redundancy check can be performed between the inter-frame merge candidate HmvpCand[i-1] at index i-1 and the merge candidate. In this case, the redundancy check between the inter-frame merge candidate HmvpCand[i-1] and the merge candidate mergeCandList[j] can be omitted.
[0306] For example, in Figure 25 In the example shown, HmvpCand[i] is determined to be the same as mergeCandList[2]. Therefore, HmvpCand[i] is not added to the merge candidate list, and a redundancy check can be performed on HmvpCand[i-1]. In this case, the redundancy check between HvmpCand[i-1] and mergeCandList[2] can be omitted.
[0307] When the number of merge candidates in the current block's merge candidate list is less than the maximum value, in addition to inter-frame merge candidates, at least one of pairwise merge candidates or zero merge candidates may be included. Pairwise merge candidates are those whose motion vectors are averaged from two or more merge candidates, while zero merge candidates are those whose motion vectors are 0.
[0308] Merge candidates for the current block can be added in the following order.
[0309] Spatial merge candidate - Temporal merge candidate - Inter-frame merge candidate - (Inter-frame affine merge candidate) - Pairwise merge candidate - Zero merge candidate
[0310] Spatial merge candidates refer to merge candidates derived from at least one of adjacent or non-adjacent blocks, while temporal merge candidates refer to merge candidates derived from a prior reference image. The inter-frame affine merge candidate column represents inter-frame merge candidates derived from blocks encoded / decoded using an affine motion model.
[0311] Inter-frame motion information lists can also be used in advanced motion vector prediction mode. For example, if the number of motion vector prediction candidates included in the current block's motion vector prediction candidate list is less than the maximum value, inter-frame merging candidates included in the inter-frame motion information list can be set as motion vector prediction candidates for the current block. Specifically, the motion vectors of the inter-frame merging candidates are set as motion vector prediction candidates.
[0312] If any one of the motion vector prediction candidates included in the motion vector prediction candidate list for the current block is selected, the selected candidate is set as the motion vector prediction value for the current block. After decoding the motion vector residual value for the current block, the motion vector for the current block can be obtained by adding the motion vector prediction value and the motion vector residual value.
[0313] The candidate list for motion vector prediction of the current block can be constructed in the following order.
[0314] Spatial motion vector prediction candidate - Temporal motion vector prediction candidate - Inter-frame decoding region merging candidate - (Inter-frame decoding region affine merging candidate) - Zero motion vector prediction candidate
[0315] Spatial motion vector prediction candidates are those derived from at least one of neighboring or non-neighboring blocks, while temporal motion vector prediction candidates are those derived from a prior reference image. The inter-frame affine merging candidate column represents inter-frame motion vector prediction candidates derived from blocks encoded / decoded using an affine motion model. Zero motion vector prediction candidates represent candidates with a motion vector value of 0.
[0316] A merge processing region larger than the encoded block can be specified. Encoded blocks included in the merge processing region can be encoded / decoded in parallel, not sequentially. "Not encoded / decoded in sequence" means that the encoding / decoding order is not specified. Therefore, the encoding / decoding process of blocks included in the merge processing region can be processed independently. Alternatively, blocks included in the merge processing region can share merge candidates. These merge candidates can be derived based on the merge processing region.
[0317] Based on the aforementioned characteristics, the merged processing region can also be referred to as the parallel processing region, the shared merge region (SMR), or the merge estimation region (MER).
[0318] Merge candidates for the current block can be derived based on the encoded blocks. However, when the current block is included in a parallel merge region that is larger than the current block, candidate blocks included in the same parallel merge region as the current block can be set as unusable as merge candidates.
[0319] Figure 26 This is a diagram illustrating an example where candidate blocks included in the same side-by-side merge region as the current block are set to be unusable as merge candidates.
[0320] exist Figure 26In the example shown in (a), when encoding / decoding CU5, blocks that include reference samples adjacent to CU5 can be set as candidate blocks. In this case, candidate blocks X3 and X4, which are included in the same parallel merging region as CU5, can be set as merging candidates that are not usable as CU5. On the other hand, candidate blocks X0, X1, and X2, which are not included in the same parallel merging region as CU5, can be set as merging candidates.
[0321] exist Figure 26 In the example shown in (b), when encoding / decoding CU8, blocks that include reference samples adjacent to CU8 can be set as candidate blocks. In this case, candidate blocks X6, X7, and X8, which are included in the same parallel merging region as CU8, can be set as non-merging candidates. On the other hand, candidate blocks X5 and X9, which are not included in the same parallel merging region as CU5, can be set as merging candidates.
[0322] The parallel merging region can be square or non-square. Information for determining the parallel merging region can be transmitted via signals in the bitstream. This information may include at least one of information indicating the shape of the parallel merging region and information indicating the size of the parallel merging region. When the parallel merging region is non-square, at least one of information indicating the size of the parallel merging region, information indicating the width and / or height of the parallel merging region, or information indicating the width-to-height ratio of the parallel merging region can be transmitted via signals in the bitstream.
[0323] The size of the parallel merging region can be determined based on at least one of the information transmitted by signaling through the bitstream, the image resolution, the strip size, or the tile size.
[0324] When performing motion compensation prediction on blocks included in a parallel merging region, inter-frame merging candidates derived from the motion information of blocks whose motion compensation prediction has already been performed can be added to the inter-frame motion.
[0325] However, when inter-frame merge candidates derived from blocks included in a parallel merge region are added to the inter-frame motion information list, the inter-frame merge candidates derived from those blocks may be used during the encoding / decoding of other blocks within the parallel merge region that are actually slower than those blocks. That is, although inter-block dependencies should be eliminated when encoding / decoding blocks included in a parallel merge region, there may be cases where motion information from other blocks included in the parallel merge region is used to perform motion prediction compensation. To address this issue, even after encoding / decoding of blocks included in a parallel merge region is completed, the motion information of the encoded / decoded blocks may not be added to the inter-frame motion information list.
[0326] Alternatively, if motion compensation prediction is performed on blocks included in the parallel merging region, inter-frame merging candidates derived from said blocks can be added to the inter-frame motion information list in a predefined order. This predefined order can be determined based on the scan order of coded blocks within the parallel merging region or coded tree unit. The scan order can be at least one of raster scan, horizontal scan, vertical scan, or zigzag scan. Alternatively, the predefined order can be determined based on the number of blocks with motion information for each block or the number of blocks with the same motion information.
[0327] Alternatively, inter-frame merge candidates containing unidirectional motion information can be added to the inter-frame region merge row table before inter-frame merge candidates containing bidirectional motion information. Conversely, inter-frame merge candidates containing bidirectional motion information can be added to the inter-frame merge candidate list before inter-frame merge candidates containing unidirectional motion information.
[0328] Alternatively, inter-frame merging candidates can be added to the inter-frame motion information list in order of high or low usage frequency within the parallel merging region or coding tree unit.
[0329] If the current block is included in a parallel merge region and the number of merge candidates included in the current block's merge candidate list is less than the maximum number, inter-frame merge candidates included in the inter-frame motion information list can be added to the merge candidate list. In this case, it can be configured not to add inter-frame merge candidates derived from blocks included in the same parallel merge region as the current block to the current block's merge candidate list.
[0330] Alternatively, when including the current block in a parallel merging region, it can be configured not to use inter-frame merge candidates included in the inter-frame motion information list. That is, even if the number of merge candidates included in the merge candidate list of the current block is less than the maximum number, inter-frame merge candidates included in the inter-frame motion information list may not be added to the merge candidate list.
[0331] Inter-frame motion information lists can be configured for parallel merging regions or coding tree units. These lists temporarily store motion information for blocks within the parallel merging region. To distinguish between general inter-frame motion information lists and those for parallel merging regions or coding tree units, the list for parallel merging regions or coding tree units is referred to as a temporal motion information list. Furthermore, inter-frame merging candidates stored in the temporal motion information list are called temporal merging candidates.
[0332] Figure 27 This is a diagram showing a list of time-domain motion information.
[0333] A temporal motion information list for coding tree units or parallel merging regions can be configured. When motion compensation prediction has already been performed on the current block included in a coding tree unit or parallel merging region, the motion information of that block may not be added to the inter-frame prediction motion information list HmvpCandList. Instead, temporal merging candidates derived from that block can be added to the temporal motion information list HmvpMERCandList. That is, temporal merging candidates added to the temporal motion information list may not be added to the inter-frame motion information list. Therefore, the inter-frame motion information list may not include inter-frame merging candidates derived from the motion information of blocks included in the coding tree unit or parallel merging region.
[0334] The maximum number of merging candidates that can be included in the temporal motion information list can be set to the same as the maximum number of merging candidates that can be included in the inter-frame motion information list. Alternatively, the maximum number of merging candidates that can be included in the temporal motion information list can be determined based on the size of the coding tree unit or the parallel merging region.
[0335] It can be configured so that the current block included in the coding tree unit or parallel merging region does not use the temporal motion information list for the coding tree unit or the parallel merging region. That is, when the number of merging candidates included in the merging candidate list of the current block is less than the maximum value, the inter-frame merging candidates included in the inter-frame motion information list are added to the merging candidate list, and the temporal merging candidates included in the temporal motion information list are not added to the merging candidate list. Thus, motion information from other blocks included in the same coding tree unit or parallel merging region as the current block can be excluded from motion compensation prediction for the current block.
[0336] Once the encoding / decoding of all blocks included in the coding tree unit or parallel merging region is complete, the inter-frame motion information list and the temporal motion information list can be merged.
[0337] Figure 28 This is a diagram illustrating an example of merging the inter-frame motion information list and the temporal motion information list.
[0338] When the encoding / decoding of all blocks included in the coding tree unit or parallel merging region is completed, such as Figure 28 As shown in the example, the inter-frame motion information list can be updated using temporal merge candidates included in the temporal motion information list.
[0339] At this point, the temporal merge candidates included in the temporal motion information list can be added to the inter-frame motion information list in the order they were inserted into the temporal motion information list (i.e., in ascending or descending order of the index values).
[0340] As another example, temporal merge candidates included in the temporal motion information list can be added to the inter-frame motion information list in a predefined order.
[0341] The predefined order can be determined based on the scanning order of coding blocks within a parallel merging region or coding tree unit. The scanning order can be at least one of raster scanning, horizontal scanning, vertical scanning, or zigzag scanning. Alternatively, the predefined order can be determined based on the number of blocks with motion information for each block or the number of blocks with the same motion information.
[0342] Alternatively, temporal merge candidates including unidirectional motion information can be added to the inter-frame merge candidate list before those including bidirectional motion information. Conversely, temporal merge candidates including bidirectional motion information can be added to the inter-frame merge candidate list before those including unidirectional motion information.
[0343] Alternatively, temporal merging candidates can be added to the inter-frame motion information list based on the order of high or low usage frequency within the parallel merging region or coding tree unit.
[0344] When adding a temporal merge candidate from the temporal motion information list to the inter-frame motion information list, redundancy checks can be performed on the temporal merge candidate. For example, if an inter-frame merge candidate identical to one included in the temporal motion information list is already stored in the inter-frame motion information list, it is not necessary to add the temporal merge candidate to the inter-frame motion information list. In this case, redundancy checks can be performed on some inter-frame merge candidates included in the inter-frame motion information list. For example, redundancy checks can be performed on inter-frame prediction merge candidates with indices greater than or equal to a threshold. For example, if a temporal merge candidate is the same as an inter-frame merge candidate with an index greater than or equal to a predefined value, it is not necessary to add the temporal merge candidate to the inter-frame motion information list.
[0345] Intra-frame prediction uses reconstructed samples that have already been encoded / decoded from the surrounding blocks to predict the current block. In this case, intra-frame prediction of the current block can use reconstructed samples before the application of the in-loop filter.
[0346] Intra-prediction techniques include matrix-based intra-prediction and general intra-prediction that takes into account the directionality with surrounding reconstructed samples. Information indicating the intra-prediction technique for the current block can be signaled via the bitstream. This information may be a 1-bit flag. Alternatively, the intra-prediction technique for the current block can be determined based on at least one of the intra-prediction techniques of the current block's position, size, shape, or neighboring blocks. For example, when the current block crosses an image boundary, the current block is set not to apply matrix-based intra-prediction.
[0347] Matrix-based intra-frame prediction is a method that obtains the predicted block for the current block by performing matrix multiplication between the matrices stored in the encoder and decoder and the reconstructed samples surrounding the current block. Information specifying any one of the stored matrices can be sent via a signal in the bitstream. The decoder can then determine the matrix used for intra-frame prediction of the current block based on this information and the size of the current block.
[0348] General intra-frame prediction is a method for obtaining the prediction block associated with the current block based on non-angular intra-frame prediction mode or angular intra-frame prediction mode. The process of performing intra-frame prediction based on general intra-frame prediction is described in more detail below with reference to the accompanying drawings.
[0349] Figure 29 This is a flowchart of the intra-frame prediction method according to an embodiment of the present invention.
[0350] The reference sample line for the current block can be determined (S2901). The reference sample line refers to the set of reference samples included in the Kth row above and / or to the left of the current block. The reference samples can be derived from the reconstructed samples that have been encoded / decoded around the current block.
[0351] Index information of the reference sample lines identifying the current block among multiple reference sample lines can be transmitted via signaling in the bitstream. The multiple reference sample lines may include at least one of the first, second, third, or fourth rows / columns above and / or to the left of the current block. Table 4 shows the index assigned to each reference sample line. In Table 4, it is assumed that the first, second, and fourth rows / columns are used as reference sample line candidates.
[0352] Table 4
[0353] index Reference sample line 0 First reference sample line 1 Second reference sample line 2 Fourth reference sample line
[0354] The reference sample line for the current block can also be determined based on at least one of the following: the position, size, shape, or predicted coding patterns of adjacent blocks. For example, when the current block is adjacent to the boundary of an image, tile, strip, or coding tree unit, the first reference sample line can be determined as the reference sample line for the current block.
[0355] The reference sample line can include an upper reference sample located above the current block and a left reference sample located to the left of the current block. The upper and left reference samples can be derived from the reconstructed samples surrounding the current block. The reconstructed samples can be in a state prior to the application of the in-loop filter.
[0356] Figure 30 This is a diagram showing the reference samples included in each reference sample line.
[0357] Based on the intra-prediction mode of the current block, a prediction sample can be obtained using at least one of the reference samples belonging to the reference sample line.
[0358] Next, the intra-prediction mode of the current block can be determined (S2902). For the intra-prediction mode of the current block, at least one of a non-angular intra-prediction mode or an angular intra-prediction mode can be determined as the intra-prediction mode of the current block. Non-angular intra-prediction modes include planar and DC (diagonal) modes, and angular intra-prediction modes include 33 or 65 modes from the lower left diagonal to the upper right diagonal.
[0359] Figure 31 This is a diagram illustrating the intra-frame prediction mode.
[0360] Figure 31 (a) shows 35 intra-frame prediction modes, and Figure 31 (b) shows 67 intra-frame prediction modes.
[0361] Can be defined with Figure 31 The examples shown are for more or fewer intra-frame prediction modes.
[0362] The Most Probable Mode (MPM) can be set based on the intra-prediction modes of neighboring blocks adjacent to the current block. Neighboring blocks can include the left-side neighboring block to the left of the current block and the top-side neighboring block above the current block. When the coordinates of the top-left sample of the current block are (0, 0), the left-side neighboring block can include samples at positions (-1, 0), (-1, H-1), or (-1, (H-1) / 2), where H represents the height of the current block. The top-side neighboring block can include samples at positions (0, -1), (W-1, -1), or ((W-1) / 2, -1), where W represents the width of the current block.
[0363] When encoding adjacent blocks using general intra-prediction, the MPM can be derived based on the intra-prediction modes of the adjacent blocks. Specifically, the intra-prediction mode of the left adjacent block can be set to the variable candIntraPredModeA, and the intra-prediction mode of the upper adjacent block can be set to the variable candIntraPredModeB.
[0364] In this case, when a neighboring block is unavailable (e.g., when a neighboring block has not yet been encoded / decoded or when the neighboring block's position deviates from the image boundary), the variable `candIntraPredModeX` (where X is A or B), derived from the neighboring block's intra-prediction mode, can be set to the default mode if the neighboring block is encoded using matrix-based intra-prediction, inter-prediction, or is included in a different coding tree unit than the current block. The default mode can include at least one of planar mode, DC mode, vertical mode, or horizontal mode.
[0365] Alternatively, when encoding adjacent blocks using matrix-based intra-prediction, the intra-prediction mode corresponding to the index value used to specify any one of the matrices can be set to candIntraPredModeX. For this purpose, a lookup table indicating the mapping between the index values used to specify the matrices and the intra-prediction modes can be pre-stored in the encoder and decoder.
[0366] The MPM can be derived based on the variables candIntraPredModeA and candIntraPredModeB. The number of MPMs included in the MPM list can be predefined in the encoder and decoder. For example, the number of MPMs can be 3, 4, 5, or 6. Alternatively, information indicating the number of MPMs can be sent via signals in the bitstream. Alternatively, the number of MPMs can be determined based on at least one of the predictive coding modes of neighboring blocks, the size of the current block, or its shape.
[0367] In the embodiments described later, it is assumed that there are 3 MPMs, which will be referred to as MPM[0], MPM[1], and MPM[2]. When there are more than 3 MPMs, the MPMs may include the 3 MPMs described in the embodiments described later.
[0368] When candIntraPredA and candIntraPredB are the same and candIntraPredA is in planar mode or DC mode, MPM[0] and MPM[1] can be set to planar mode and DC mode respectively. MPM[2] can be set to vertical intra-prediction mode, horizontal intra-prediction mode or diagonal intra-prediction mode. Diagonal intra-prediction mode can be lower left diagonal intra-prediction mode, upper left intra-prediction mode or upper right intra-prediction mode.
[0369] When candIntraPredA and candIntraPredB are the same and candIntraPredA is in intra-prediction mode, MPM[0] can be set to be the same as candIntraPredA. MPM[1] and MPM[2] can be set to intra-prediction modes similar to candIntraPredA. Intra-prediction modes similar to candIntraPredA can be intra-prediction modes with an index difference of ±1 or ±2 from candIntraPredA. Intra-prediction modes similar to candIntraPredA can be derived using modulo operations (%) and offsets.
[0370] When candIntraPredA and candIntraPredB are different, MPM[0] can be set to be the same as candIntraPredA, and MPM[1] can be set to be the same as candIntraPredB. In this case, when both candIntraPredA and candIntraPredB are non-angular intra-prediction modes, MPM[2] can be set to a vertical intra-prediction mode, a horizontal intra-prediction mode, or a diagonal intra-prediction mode. Alternatively, when at least one of candIntraPredA and candIntraPredB is an angular intra-prediction mode, MPM[2] can be set to an intra-prediction mode derived by adding or subtracting an offset value from the larger of the plane, DC, or candIntraPredA or candIntraPredB values. The offset value can be 1 or 2.
[0371] An MPM list comprising multiple MPMs can be generated, and information indicating whether an MPM with the same intra-prediction mode as the current block is included in the MPM list can be signaled via the bitstream. This information is a 1-bit flag, referred to as the MPM flag. When the MPM flag indicates that an MPM with the same mode as the current block is included in the MPM list, index information identifying one of the MPMs can be signaled via the bitstream. The MPM specified by the index information can be set as the intra-prediction mode of the current block. When the MPM flag indicates that an MPM with the same mode as the current block is not included in the MPM list, residual mode information indicating any of the remaining intra-prediction modes other than the MPM can be signaled via the bitstream. The residual mode information represents the index value corresponding to the intra-prediction mode of the current block when the index is reallocated to the remaining intra-prediction modes other than the MPM. The decoder can sort the MPMs in ascending order and determine the intra-prediction mode of the current block by comparing the residual mode information with the MPMs. For example, when the residual mode information is the same as or smaller than the MPM, the intra-prediction mode of the current block can be derived by adding 1 to the residual mode information.
[0372] Instead of setting the default mode to MPM, information indicating whether the intra-prediction mode of the current block is the default mode can be signaled via the bitstream. This information is a 1-bit flag, which may be called the default mode flag. The default mode flag is signaled only if the MPM flag indicates that the same MPM as the current block is included in the MPM list. As mentioned above, the default mode can include at least one of planar, DC, vertical, or horizontal modes. For example, when planar is set as the default mode, the default mode flag can indicate whether the intra-prediction mode of the current block is planar. When the default mode flag indicates that the intra-prediction mode of the current block is not the default mode, one of the MPMs indicated by index information can be set as the intra-prediction mode of the current block.
[0373] When multiple intra-prediction modes are set as the default mode, index information indicating any of the default modes can be further sent using a signal. The intra-prediction mode of the current block can be set to the default mode indicated by the index information.
[0374] When the index of the reference sample line of the current block is not 0, the default mode is set not to be used. Therefore, when the index of the reference sample line is not 0, the default mode flag is not sent by signal, and the value of the default mode flag can be set to a predefined value (i.e., false).
[0375] After determining the intra-prediction mode of the current block, the prediction samples of the current block can be obtained based on the determined intra-prediction mode (S2903).
[0376] When DC mode is selected, predicted samples related to the current block can be generated based on the average value of reference samples. Specifically, the value of the overall sample within the predicted block can be generated based on the average value of reference samples. The average value can be derived using at least one of the upper reference sample located above the current block and the left reference sample located to the left of the current block.
[0377] The number or range of reference samples used to derive the average value will vary depending on the shape of the current block. For example, when the current block is a non-square block with a width greater than its height, the average value can be calculated using only the top reference sample. On the other hand, when the current block is a non-square block with a width less than its height, the average value can be calculated using only the left reference sample. That is, when the width and height of the current block are different, the average value can be calculated using only the reference sample adjacent to the longer side. Alternatively, the choice between using only the top reference sample or only the left reference sample can be determined based on the width-to-height ratio of the current block.
[0378] When the planar mode is selected, prediction samples can be obtained using horizontal and vertical prediction samples. Specifically, the horizontal prediction sample is obtained based on the left and right reference samples located on the same horizontal line as the prediction sample, and the vertical prediction sample is obtained based on the upper and lower reference samples located on the same vertical line as the prediction sample. The right reference sample can be generated by copying the reference sample adjacent to the upper right corner of the current block, and the lower reference sample can be generated by copying the reference sample adjacent to the lower left corner of the current block. The horizontal prediction sample can be obtained based on a weighted sum of the left and right reference samples, and the vertical prediction sample can be obtained based on a weighted sum of the upper and lower reference samples. In this case, the weighting value assigned to each reference sample can be determined based on the position of the prediction sample. The prediction sample can also be obtained based on the average or weighted sum of the horizontal and vertical prediction samples. When performing a weighted sum operation, the weighting value assigned to the horizontal and vertical prediction samples can be determined based on the position of the prediction sample.
[0379] When an angle prediction mode is selected, parameters representing the prediction direction (or prediction angle) of the selected angle prediction mode can be determined. Table 5 below shows the intrapredAng parameter for each intrapredangling prediction mode.
[0380] Table 5
[0381]
[0382]
[0383] Table 5 shows the intra-prediction parameters for each intra-prediction mode with an index of any one of 2 to 34 when 35 intra-prediction modes are defined. When more than 33 angular intra-prediction modes are defined, Table 5 further breaks down the intra-prediction parameters for setting each angular intra-prediction mode.
[0384] After arranging the upper and left reference samples of the current block into a column, the predicted sample can be obtained based on the value of the intra-frame orientation parameter. In this case, when the value of the intra-frame orientation parameter is negative, the left and upper reference samples can be arranged into a column.
[0385] Figure 32 and Figure 33 This is a diagram illustrating an example of a one-dimensional arrangement of reference samples in a row.
[0386] Figure 32 An example of a vertically oriented one-dimensional array of reference samples is shown, and Figure 33An example of a horizontally oriented one-dimensional array of reference samples is shown. This will be described under the assumption of defining 35 intra-frame prediction modes. Figure 32 and 33 Examples of implementations.
[0387] When the intra-prediction mode index is any one of 11 to 18, a one-dimensional horizontal arrangement of the upper reference sample can be applied, rotating counterclockwise. When the intra-prediction mode index is any one of 19 to 25, a one-dimensional vertical arrangement of the left reference sample can be applied, rotating clockwise. When arranging the reference samples in a column, the intra-prediction mode angle can be considered.
[0388] Reference sample determination parameters can be determined based on intra-frame orientation parameters. These parameters may include a reference sample index for specifying the reference sample and weighting parameters for determining the weights applied to the reference sample.
[0389] The reference sample index iIdx and the weight parameter ifact can be obtained through Equations 2 and 3 below, respectively.
[0390] Equation 2
[0391] iIdx=(y+1)*P ang / 32
[0392] Equation 3
[0393] i fact =[(y+1)*P ang ]&31
[0394] In equations 2 and 3, Pang represents the intra-frame orientation parameter. The reference sample specified by the reference sample index iIdx is equivalent to an integer pixel (Integer pel).
[0395] To derive predicted samples, more than one reference sample can be specified. Specifically, considering the slope of the prediction pattern, the position of the reference sample used when deriving predicted samples can be specified. For example, using the reference sample index iIdx, the reference sample used when deriving predicted samples can be specified.
[0396] In this scenario, when the slope of the intra-prediction mode is not represented by a single reference sample, a prediction sample can be generated by interpolating multiple reference samples. For example, when the slope of the intra-prediction mode is the value between the slope between the prediction sample and the first reference sample and the slope between the prediction sample and the second reference sample, the prediction sample can be obtained by interpolating the first and second reference samples. That is, when the angular line following the intra-prediction angle does not pass through a reference sample located at an integer pixel, the prediction sample can be obtained by interpolating the reference samples that are adjacent to the left, right, or top and bottom of the position through which the angular line passes.
[0397] Equation 4 below shows an example of obtaining a predicted sample based on a reference sample.
[0398] Equation 4
[0399] P(x, y)=((32-i) fact ) / 32)*Ref_1D(x+iIdx+1)+(i fact / 32)*Ref_1D(x+iIdx+2)
[0400] In Equation 4, P represents the predicted sample, and Ref_1D represents any one of the reference samples in a one-dimensional arrangement. In this case, the position of the reference sample can be determined based on the position (x, y) of the predicted sample and the index iIdx of the reference sample.
[0401] When the slope of the intra-prediction mode can be represented by a reference sample, the weighting parameter ifact can be set to 0. Therefore, Equation 4 can be simplified to Equation 5 below.
[0402] Equation 5
[0403] P(x, y) = Ref_1D(x + iIdx + 1)
[0404] Intra-prediction can also be performed on the current block based on multiple intra-prediction modes. For example, intra-prediction modes can be derived for different prediction samples, and prediction samples can be derived based on the intra-prediction modes assigned to each prediction sample.
[0405] Alternatively, intra-prediction modes can be derived for different regions, and intra-prediction can be performed on each region based on the intra-prediction modes assigned to each region. Each region may include at least one sample. The size or shape of the region can be adaptively determined based on at least one of the size, shape, or intra-prediction mode of the current block. Alternatively, in the encoder and decoder, it is possible to predefine at least one of the size or shape of the region, independent of the size or shape of the current block.
[0406] Alternatively, intra-prediction can be performed based on multiple intra-prediction methods, and the final prediction sample can be derived based on the average or weighted sum of multiple prediction samples obtained through multiple intra-prediction operations. For example, a first prediction sample can be obtained by performing intra-prediction based on a first intra-prediction mode, and a second prediction sample can be obtained by performing intra-prediction based on a second intra-prediction mode. Then, the final prediction sample can be obtained based on the average or weighted sum of the first and second prediction samples. In this case, the weighting values assigned to the first and second prediction samples can be determined by considering at least one of whether the first intra-prediction mode is a non-angle / angle prediction mode, whether the second intra-prediction mode is a non-angle / angle prediction mode, or the intra-prediction modes of adjacent blocks.
[0407] Multiple intra-frame prediction modes can be a combination of non-angle intra-frame prediction modes and angle prediction modes, a combination of angle prediction modes, or a combination of non-angle prediction modes.
[0408] Figure 34 This is a diagram showing the angle formed between the prediction pattern within the angular frame and a straight line parallel to the x-axis.
[0409] like Figure 34 In the example shown, the angle prediction pattern can exist between the lower left diagonal and the upper right diagonal. When described as the angle formed by the x-axis and the angle prediction pattern, the angle prediction pattern can exist between 45 degrees (lower left diagonal) and -135 degrees (upper right diagonal).
[0410] When the current block is not square, the following will happen: depending on the intra-prediction mode of the current block, the prediction sample is derived by using the reference sample that is farther away from the prediction sample from the reference samples located on the corner that follows the intra-prediction angle, rather than the reference sample that is closer to the prediction sample.
[0411] Figure 35 This is a diagram showing an example of obtaining a predicted sample when the current block is not square.
[0412] For example, as in Figure 35 In the example shown in (a), it is assumed that the current block is a non-square shape with a width greater than its height, and the intra-frame prediction mode of the current block is an angular intra-frame prediction mode with an angle between 0 and 45 degrees. In this case, when deriving the prediction sample A near the right column of the current block, a left reference sample L, which is far from the prediction sample in the angular mode located at the angle, is used instead of the upper reference sample T, which is close to the prediction sample.
[0413] As another example, such as in Figure 35In the example shown in (b), it is assumed that the current block is a non-square shape with a height greater than its width, and the intra-frame prediction mode of the current block is an angular intra-frame prediction mode with an angle between -90 degrees and -135 degrees. In the above case, when deriving the prediction sample A near the lower row of the current block, a situation occurs where an upper reference sample T, which is far from the prediction sample in the angular mode located at the angle, is used instead of a left reference sample L that is close to the prediction sample.
[0414] To address the above issue, when the current block is not square, the intra-prediction mode of the current block can be replaced with an intra-prediction mode in the opposite direction. Therefore, for non-square blocks, a mode with a higher frequency of prediction can be used. Figure 31 The angle prediction modes shown are for angles larger or smaller than the indicated angles. This type of intra-frame prediction mode can be defined as a wide-angle intra-frame prediction mode. A wide-angle intra-frame prediction mode refers to an intra-frame prediction mode that does not fall within the range of 45 degrees to -135 degrees.
[0415] Figure 36 This is a diagram illustrating the wide-angle intra-frame prediction mode.
[0416] exist Figure 36 In the example shown, the intra-prediction modes with indices -1 to -14 and the intra-prediction modes with indices 67 to 80 represent wide-angle intra-prediction modes.
[0417] Despite Figure 36 The diagram shows 14 wide-angle intra-prediction modes (-1 to -14) with angles greater than 45 degrees and 14 wide-angle intra-prediction modes (67 to 80) with angles less than -135 degrees, but more or fewer wide-angle intra-prediction modes can be defined.
[0418] When using the wide-angle intra-frame prediction mode, the length of the upper reference sample is set to 2W+1, and the length of the left reference sample is set to 2H+1.
[0419] When using the wide-angle intra-frame prediction mode, a reference sample T can be used for prediction. Figure 35 The sample A shown in (a) can be used to predict the reference sample L. Figure 35 Sample A is shown in (b).
[0420] By adding the existing intra-prediction modes to N wide-angle intra-prediction modes, a total of 67+N intra-prediction modes can be used. For example, Table 6 shows the intra-direction parameters of the intra-prediction modes when 20 wide-angle intra-prediction modes are defined.
[0421] Table 6
[0422]
[0423]
[0424] The intra-prediction parameter can be set differently based on at least one of the size, shape, or reference sample line of the current block. For example, the intra-prediction parameter for a given intra-prediction mode can differ depending on whether the current block is square or non-square. For instance, the intra-prediction parameter intraPredAngle of intra-prediction mode 15 can have a larger value when the current block is square than when the current block is non-square.
[0425] Alternatively, the intraPredAngle parameter of intraPred mode 75 can have a larger value when the index of the reference sample line of the current block is 1 or higher than when the index of the reference sample line of the current block is 0.
[0426] When the current block is non-square and the intra-prediction mode of the current block obtained in step S2902 falls within the transformation range, the intra-prediction mode of the current block can be transformed into a wide-angle intra-prediction mode. The transformation range can be determined based on at least one of the size, shape, or ratio of the current block. The ratio can represent the ratio between the width and height of the current block.
[0427] When the current block is a non-square with a width greater than its height, the transformation range can be set from the intra-prediction mode index in the upper-right diagonal direction (e.g., 66) to (the index of the intra-prediction mode in the upper-right diagonal direction - N). Here, N can be determined based on the ratio of the current block. When the intra-prediction mode of the current block falls within the transformation range, the intra-prediction mode can be transformed into a wide-angle intra-prediction mode. The transformation can be performed by subtracting a predefined value from the intra-prediction mode; the predefined value can be the total number of intra-prediction modes other than the wide-angle intra-prediction mode (e.g., 67).
[0428] According to the embodiment, the intra-frame prediction modes between the 66th and 53rd frames can be transformed into wide-angle intra-frame prediction modes between the -1st and -14th frames, respectively.
[0429] When the current block is a non-square with a height greater than its width, the transformation range can be set from the intra-prediction mode index (e.g., 2) along the lower left diagonal to (the index of the intra-prediction mode along the lower left diagonal + M). Here, M can be determined based on the ratio of the current block. When the intra-prediction mode of the current block falls within the transformation range, the intra-prediction mode can be transformed into a wide-angle intra-prediction mode. The transformation can be performed by adding a predefined value to the intra-prediction mode; the predefined value can be the total number of angular intra-prediction modes other than the wide-angle intra-prediction mode (e.g., 65).
[0430] According to the embodiment, the intra-frame prediction modes between the 2nd and 15th frames are transformed into wide-angle intra-frame prediction modes between the 67th and 80th frames, respectively.
[0431] Hereinafter, the intra-frame prediction modes that fall within the transform range will be referred to as wide-angle intra-frame replacement prediction modes.
[0432] The transform range can be determined based on the ratio of the current block. For example, Tables 7 and 8 show the transform range when 35 intra-prediction modes and 67 intra-prediction modes, excluding the wide-angle intra-prediction mode, are defined, respectively.
[0433] Table 7
[0434] condition Replace Intra-Prediction Mode W / H = 2 Pattern 2, 3, 4 W / H>2 Patterns 2, 3, 4, 5, 6 W / H = 1 none H / W = 1 / 2 Patterns 32, 33, 34 H / W<1 / 2 Patterns 30, 31, 32, 33, 34
[0435] Table 8
[0436] condition Replace Intra-Prediction Mode W / H = 2 Patterns 2, 3, 4, 5, 6, 7 W / H>2 Patterns 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 W / H = 1 none H / W = 1 / 2 Patterns 61, 62, 63, 64, 65, 66 H / W<1 / 2 Patterns 57, 58, 59, 60, 61, 62, 63, 64, 65, 66
[0437] As shown in the examples in Tables 7 and 8, the number of wide-angle intra-frame replacement prediction modes falling within the transform range can vary depending on the ratio of the current block.
[0438] With the use of a wide-angle intra prediction mode in addition to the existing intra prediction modes, the resources required for encoding the wide-angle intra prediction mode increase, potentially reducing coding efficiency. Therefore, instead of directly encoding the wide-angle intra prediction mode, encoding an alternative intra prediction mode associated with it can improve coding efficiency.
[0439] For example, when encoding the current block using the 67th wide-angle intra-prediction mode, the 67th wide-angle replacement intra-prediction mode (number 2) can be encoded as the intra-prediction mode for the current block. Similarly, when encoding the current block using the -1st wide-angle intra-prediction mode, the -1st wide-angle replacement intra-prediction mode (number 66) can be encoded as the intra-prediction mode for the current block.
[0440] The decoder can decode the intra-prediction mode of the current block and determine whether the decoded intra-prediction mode is included in the transform range. When the decoded intra-prediction mode is a wide-angle replacement intra-prediction mode, it can transform the intra-prediction mode into a wide-angle intra-prediction mode.
[0441] Alternatively, when encoding the current block in wide-angle intra-prediction mode, wide-angle intra-prediction mode can also be encoded directly.
[0442] Encoding of intra-prediction modes can be implemented based on the MPM list. Specifically, when encoding adjacent blocks in wide-angle intra-prediction mode, the MPM can be set based on the wide-angle replacement intra-prediction mode corresponding to the wide-angle intra-prediction mode. For example, when adjacent blocks are encoded in wide-angle intra-prediction mode, the variable candIntraPredX (where X is A or B) can be set to the wide-angle replacement intra-prediction mode.
[0443] If a prediction block is generated from the results of intra-frame prediction, the prediction samples can be updated based on the position of each prediction sample included in the prediction block. The update method described above can be called an intra-frame weighted prediction method based on sample position (or position-dependent prediction combination (PDPC)).
[0444] Whether to use PDPC can be determined by considering the intra-prediction mode of the current block, the reference sample line of the current block, the size of the current block, or the color component. For example, PDPC can be used if the intra-prediction mode of the current block is at least one of planar mode, DC mode, vertical mode, horizontal mode, a mode with an index value smaller than the vertical mode, or a mode with an index value larger than the horizontal mode. Alternatively, PDPC can be used only if at least one of the width or height of the current block is greater than 4. Alternatively, PDPC can be used only if the index of the reference image line of the current block is 0. Alternatively, PDPC can be used only if the index of the reference image line of the current block is greater than or equal to a predefined value. Alternatively, PDPC can be used only for the luma component. Alternatively, whether to use PDPC can be determined based on whether two or more of the listed conditions are met.
[0445] As another example, information indicating whether PDPC is applied can be sent via signaling through the bitstream.
[0446] If a prediction sample is obtained through intra-frame prediction, a reference sample for correcting the prediction sample can be determined based on the position of the obtained prediction sample. For ease of explanation, in the following embodiments, the reference sample used to correct the prediction sample will be referred to as the PDPC reference sample. Furthermore, the prediction sample obtained through intra-frame prediction will be referred to as the first prediction sample, and the prediction sample obtained by correcting the first prediction sample will be referred to as the second prediction sample.
[0447] Figure 37 This is a diagram illustrating the application of PDPC.
[0448] The first prediction sample can be corrected using at least one PDPC reference sample. The PDPC reference sample may include at least one of the following: a reference sample adjacent to the top-left corner of the current block, an upper reference sample above the current block, or a left reference sample to the left of the current block.
[0449] At least one of the reference samples belonging to the reference sample line of the current block can be set as a PDPC reference sample. Alternatively, regardless of the reference sample line of the current block, at least one of the reference samples belonging to the reference sample line with index 0 can be set as a PDPC reference sample. For example, even if the first prediction sample is obtained using reference samples included in the reference sample line with index 1 or index 2, the second prediction sample can also be obtained using reference samples included in the reference sample line with index 0.
[0450] The number or location of PDPC reference samples used to correct the first prediction sample can be determined by considering at least one of the intra-prediction mode of the current block, the size of the current block, the shape of the current block, or the location of the first prediction sample.
[0451] For example, when the intra-prediction mode of the current block is planar mode or DC mode, a second prediction sample can be obtained using the upper reference sample and the left reference sample. In this case, the upper reference sample can be a reference sample perpendicular to the first prediction sample (e.g., a reference sample with the same x-coordinate), and the left reference sample can be a reference sample horizontal to the first prediction sample (e.g., a reference sample with the same y-coordinate).
[0452] When the intra-prediction mode of the current block is horizontal intra-prediction mode, a second prediction sample can be obtained using the upper reference sample. In this case, the upper reference sample can be a reference sample perpendicular to the first prediction sample.
[0453] When the intra-prediction mode of the current block is vertical intra-prediction mode, a second prediction sample can be obtained using the left reference sample. In this case, the left reference sample can be a reference sample at the same level as the first prediction sample.
[0454] When the intra-prediction mode of the current block is either the bottom-left diagonal intra-prediction mode or the top-right diagonal intra-prediction mode, a second prediction sample can be obtained based on the top-left reference sample, the top reference sample, and the left reference sample. The top-left reference sample can be the reference sample adjacent to the top-left corner of the current block (e.g., the reference sample at position (-1, -1)). The top reference sample can be the reference sample located in the top-right diagonal direction of the first prediction sample, and the left reference sample can be the reference sample located in the bottom-left diagonal direction of the first prediction sample.
[0455] In summary, when the position of the first predicted sample is (x, y), R(-1, -1) can be set as the top-left reference sample, and R(x+y+1, -1) or R(x, -1) can be set as the top reference sample. Alternatively, R(-1, x+y+1) or R(-1, y) can be set as the left-side reference sample.
[0456] As another example, the position of the left or top reference sample can be determined by considering at least one of the current block shape or whether a wide-angle intra-frame mode is applied.
[0457] Specifically, when the intra-prediction mode of the current block is wide-angle intra-prediction mode, a reference sample that is offset by a certain value from the reference sample located diagonally opposite to the first prediction sample can be set as a PDPC reference sample. For example, the upper reference sample R(x+y+k+1, -1) and the left reference sample R(-1, x+y-k+1) can be set as PDPC reference samples.
[0458] At this point, the offset value k can be determined based on the wide-angle intra-frame prediction mode. Equations 6 and 7 show examples of deriving the offset value based on the wide-angle intra-frame prediction mode.
[0459] Equation 6
[0460] k = CurrIntraMode-66
[0461] if (CurrIntraMtode > 66)
[0462] Equation 7
[0463] k = -CurrIntraMode
[0464] if (CurrIntraMode < 0)
[0465] The second predicted sample can be determined based on a weighted sum of the first predicted sample and the PDPC reference sample. For example, the second predicted sample can be obtained based on the following Equation 8.
[0466] Equation 8
[0467] pred(x, y) = (xL*R) L +wT*R T -wTL*R TL +(64-wL-wT+wTL)*pred(x,y)+32)>>6
[0468] In Equation 8, RL represents the left reference sample, RT represents the top reference sample, and RTL represents the top-left reference sample. pred(x, y) represents the predicted sample at position (x, y). wL represents the weighting value assigned to the left reference sample, wT represents the weighting value assigned to the top reference sample, and wTL represents the weighting value assigned to the top-left reference sample. The weighting value assigned to the first predicted sample can be derived by subtracting the weighting value assigned to the reference sample from the maximum value. For ease of explanation, the weighting value assigned to the PDPC reference sample is referred to as the PDPC weighting value.
[0469] The weighting values assigned to each reference sample can be determined based on at least one of the intra-prediction mode of the current block or the position of the first predicted sample.
[0470] For example, at least one of wL, wT, or wTL may be directly or inversely proportional to at least one of the x-axis or y-axis coordinate values of the predicted sample. Alternatively, at least one of wL, wT, or wTL may be directly or inversely proportional to at least one of the width or height of the current block.
[0471] When the intra-prediction mode of the current block is DC mode, the PDPC weighting value can be determined as shown in Equation 9.
[0472] Equation 9
[0473] wT=32>>((y<<1)>>shift)
[0474] wL=32>>((x<<1)>>shift)
[0475] wTL = (wL >> 4) + (wT >> 4)
[0476] In Equation 9, x and y represent the positions of the first predicted sample.
[0477] In Equation 9, the variable `shift` used in the shift operation can be derived based on the width or height of the current block. For example, the variable `shift` can be derived based on the following Equations 10 or 11.
[0478] Equation 10
[0479] shift=(log2(width)-2+log2(height)-2+2)>>2
[0480] Equation 11
[0481] shift=((Log1(nTbW)+Log2(nTbH)-2)>>2)
[0482] Alternatively, the intra-frame orientation parameter of the current block can be considered to derive the variable shift.
[0483] The number or type of parameters used to derive the variable `shift` can be determined based on the intra-prediction mode of the current block. For example, when the intra-prediction mode of the current block is planar, DC, vertical, or horizontal, as shown in Equation 10 or Equation 11, the width and height of the current block can be used to derive the variable `shift`. When the intra-prediction mode of the current block has an index larger than that of the vertical intra-prediction mode, the height and intra-direction parameters of the current block can be used to derive the variable `shift`. When the intra-prediction mode of the current block has an index smaller than that of the horizontal intra-prediction mode, the width and intra-direction parameters of the current block can be used to derive the variable `shift`.
[0484] When the intra-prediction mode of the current block is planar mode, the value of wTL can be set to 0. wL and wT can be derived based on the following Equation 12.
[0485] Equation 12
[0486] wT[y]=32>>((y<<1)>>nScale)
[0487] wL[x]=32>>((x<<1)>>nScale)
[0488] When the intra-prediction mode of the current block is horizontal intra-prediction mode, wT can be set to 0, and wTL and wL can be set to the same value. On the other hand, when the intra-prediction mode of the current block is vertical intra-prediction mode, wL can be set to 0, and wTL and wT can be set to the same value.
[0489] When the intra-prediction mode of the current block is an intra-prediction mode pointing to the upper right and with an index value greater than that of the vertical direction, the PDPC weighting value can be derived as shown in Equation 13 below.
[0490] Equation 13
[0491] wT=16>>((y<<1)>>shift)
[0492] wL=16>>((x<<1)>>shift)
[0493] wTL=0
[0494] On the other hand, when the intra-prediction mode of the current block is an intra-prediction mode pointing to the lower left with an index value smaller than that of the horizontal direction, the PDPC weighting value can be derived as shown in Equation 14 below.
[0495] Equation 14
[0496] wT16>>((y<<1)>>shift)
[0497] wL=16>>((x<<1)>>shift)
[0498] wTL=0
[0499] As in the above embodiments, the PDPC weighting value can be determined based on the predicted sample's position x and y.
[0500] As another example, the weighting values assigned to each PDPC reference sample can also be determined on a sub-block basis. Predicted samples included in a sub-block can share the same PDPC weighting values.
[0501] The size of the sub-block, which serves as the basic unit for determining the weighting value, can be predefined in the encoder and decoder. For example, the weighting value can be determined for individual sub-blocks of size 2×2 or 4×4.
[0502] Alternatively, the size, shape, or number of sub-blocks can be determined based on the size or shape of the current block. For example, the coded block can be divided into 4 sub-blocks regardless of its size. Alternatively, the coded block can be divided into 4 or 16 sub-blocks depending on its size.
[0503] Alternatively, the size, shape, or number of sub-blocks can be determined based on the intra-prediction mode of the current block. For example, when the intra-prediction mode of the current block is horizontal, N columns (or N rows) can be set as a sub-block; conversely, when the intra-prediction mode of the current block is vertical, N rows (or N columns) can be set as a sub-block.
[0504] Equations 15 to 17 show examples of determining the PDPC weighting value for a 2×2 size sub-block. Equation 15 shows an example when the intra-prediction mode of the current block is DC mode.
[0505] Equation 15
[0506] wT=32>>(((y<<log2K)>>log2K)<<1)>>shift)
[0507] wL=32>>(((x<<log2K))>>log2K))<<1)>>shift)
[0508] wTL = (wL >> 4) + (wT >> 4)
[0509] In Equation 15, K can be a value determined based on the size of the sub-block or the intra-prediction mode.
[0510] Equation 16 shows an example of an intra-prediction mode in which the intra-prediction mode of the current block is an intra-prediction mode pointing to the upper right and with an index value greater than that of the vertical direction.
[0511] Equation 16
[0512] wT=16>>(((y<<log2K))>>log2K))<<1)>>shift)
[0513] wL16>>(((x<<log2K))>>log2K))<<1)>>shift)
[0514] wTL=0
[0515] Equation 17 shows an example of an intra-prediction mode where the intra-prediction mode of the current block is an intra-prediction mode pointing to the lower left with an index value smaller than that of the horizontal direction.
[0516] Equation 17
[0517] wT=16>>(((y<<log2K))>>log2K))<<1)>>shift)
[0518] wL=16>>(((x<<log2K))>>log2K))<<1)>>shift)
[0519] wTL=0
[0520] In equations 15 to 17, x and y represent the positions of the reference sample within the sub-block. The reference sample can be any one of the following: the sample located at the upper left of the sub-block, the sample located at the center of the sub-block, or the sample located at the lower right of the sub-block.
[0521] Equations 18 to 20 show examples of determining the PDPC weighting value for a 4×4 size sub-block. Equation 18 shows an example when the intra-prediction mode of the current block is DC mode.
[0522] Equation 18
[0523] wT=32>>(((y<<2)>>2)<<1)>>shift)
[0524] wL=32>>(((x<<2)>>2)<<1)>>shift)
[0525] wTL = (wL >> 4) + (wT >> 4)
[0526] Equation 19 shows an example of an intra-prediction mode in which the intra-prediction mode of the current block is an intra-prediction mode pointing to the upper right and with an index value greater than that of the vertical direction.
[0527] Equation 19
[0528] wT=16>>(((y<<2)>>2)<<1)>>shift)
[0529] wL=16>>(((x<<2)>>2)<<1)>>shift)
[0530] wTL=0
[0531] Equation 20 shows an example of an intra-prediction mode where the intra-prediction mode of the current block is an intra-prediction mode pointing to the lower left with an index value smaller than that of the horizontal direction.
[0532] Equation 20
[0533] wT=16>>(((y<<2)>>2)<<1)>>shift)
[0534] wL=16>>(((x<<2)>>2)<<1)>>shift)
[0535] wTL=0
[0536] The above embodiments illustrate the use of the position of predicted samples included in the first predicted sample or sub-block to determine the PDPC weighting value. The shape of the current block can also be further considered in determining the PDPC weighting value.
[0537] For example, in DC mode, the method for deriving the PDPC weighting value will differ depending on whether the current block is a non-square with a width greater than its height or a non-square with a height greater than its width.
[0538] Equation 21 shows an example of deriving the PDPC weighted value when the current block is a non-square with a width greater than its height, and Equation 22 shows an example of deriving the PDPC weighted value when the current block is a non-square with a height greater than its width.
[0539] Equation 21
[0540] wT=32>>((y<<1)>>shift)
[0541] wL = 32 >> (x >> shift)
[0542] wTL = (wL >> 4) + (wT >> 4)
[0543] Equation 22
[0544] wT >> (y >> shift)
[0545] wL=32>>((x<<1)>>shift)
[0546] wTL = (xL >> 4) + (wT >> 4)
[0547] When the current block is not square, the wide-angle intra-frame prediction mode can be used to predict the current block. In this way, when applying the wide-angle intra-frame prediction mode, PDPC can also be applied to update the first prediction sample.
[0548] When applying wide-angle intra-frame prediction to the current block, the shape of the coded block can be considered to determine the PDPC weighting value.
[0549] For example, when the current block is a non-square with a width greater than its height, depending on the position of the first predicted sample, the upper reference sample located to the upper right of the first predicted sample may be closer to the first predicted sample than the left reference sample located to the lower left of the first predicted sample. Therefore, when correcting the first predicted sample, the weighting value applied to the upper reference sample can be set to a larger value than the weighting value applied to the left reference sample.
[0550] On the other hand, when the current block is a non-square with a height greater than its width, depending on the position of the first predicted sample, a left reference sample located to the lower left of the first predicted sample may be closer to the first predicted sample than an upper reference sample located to the upper right of the first predicted sample. Therefore, when correcting the first predicted sample, the weighting value applied to the left reference sample can be set to a value that is larger than the weighting value applied to the upper reference sample.
[0551] Equation 23 shows an example of deriving the PDPC weighting value when the intra-prediction mode of the current block is a wide-angle intra-prediction mode with an index greater than 66.
[0552] Equation 23
[0553] wT = 16 >> (y >> shift)
[0554] wL=16>>((x<<1)>>shift)
[0555] wTL=0
[0556] Equation 24 shows an example of deriving the PDPC weights when the intra-prediction mode of the current block is a wide-angle intra-prediction mode with an index less than 0.
[0557] Equation 24
[0558] wT=16>>((y<<1)>>shift)
[0559] wL = 16 >> (x >> shift)
[0560] wTL=0
[0561] The PDPC weighting value can also be determined based on the ratio of the current block. The ratio of the current block shows the ratio of the width to the height of the current block and can be defined as shown in Equation 25 below.
[0562] Equation 25
[0563] whRatio = CUwidth / CUheight
[0564] The method for variably determining the derived PDPC weighting value can be based on the intra-prediction mode of the current block.
[0565] For example, Equations 26 and 27 illustrate examples of deriving PDPC weights when the intra-prediction mode of the current block is DC mode. Specifically, Equation 26 is an example for the case where the current block is a non-square with a width greater than its height, and Equation 27 is an example for the case where the current block is a non-square with a height greater than its width.
[0566] Equation 26
[0567] wT=32>>((y<<1)>>shift)
[0568] wL=32>>(((x<<1)>>whRatio)>>shift)
[0569] wTL = (wL >> 4) + (wT >> 4)
[0570] Equation 27
[0571] wT=32>>(((y<<1)>>1 / whRatio)>>shift)
[0572] wL=32>>((x<<1)>>shift)
[0573] wTL = (wL >> 4) + (wT >> 4)
[0574] Equation 28 shows an example of deriving the PDPC weighting value when the intra-prediction mode of the current block is a wide-angle intra-prediction mode with an index greater than 66.
[0575] Equation 28
[0576] wT=16>>(((y<<1)>>1 / whRatio)>>shift)
[0577] wL=16>>((x<<1)>>shift)
[0578] wTL=0
[0579] Equation 29 shows an example of deriving the PDPC weights when the intra-prediction mode of the current block is a wide-angle intra-prediction mode with an index less than 0.
[0580] Equation 29
[0581] wT=16>>((y<<1)>>shift)
[0582] wL=16>>(((x<<1)>>whRatio)>>shift
[0583] wTL=0
[0584] The residual image can be derived by subtracting the predicted image from the original image. In this case, when the residual image is transformed into the frequency domain, even if high-frequency components are removed, the subjective image quality of the video is not significantly degraded. Therefore, reducing the value of high-frequency components or setting the value of high-frequency components to 0 can improve compression efficiency without causing significant visual distortion. Reflecting these characteristics, the current block can be transformed to decompose the residual image into 2D frequency components. This transformation can be performed using transformation methods such as Discrete Cosine Transform (DCT) or Discrete Sine Transform (DST).
[0585] DCT uses a cosine transform to decompose (or transform) the residual image into 2D frequency components, while DST uses a sine transform to do the same. As the result of the transformation of the residual image, the frequency components can be represented as fundamental images. For example, performing a DCT transform on an N×N block yields N² fundamental pattern components. The magnitudes of each fundamental pattern component within the N×N block can be obtained through the transform. Depending on the transform technique used, the magnitudes of the fundamental pattern components can be referred to as DCT coefficients or DST coefficients.
[0586] The Direct Transformation Technique (DCT) is primarily used to transform images with a high proportion of low-frequency non-zero components. The Direct Transformation Technique (DST) is primarily used for images with a high proportion of high-frequency components.
[0587] Transformation techniques other than DCT or DST can also be used to transform residual images.
[0588] In the following text, the process of transforming the residual image into two-dimensional frequency components is referred to as two-dimensional image transformation. Furthermore, the magnitudes of the fundamental pattern components obtained from the transformation are called transformation coefficients. For example, transformation coefficients can refer to DCT coefficients or DST coefficients. When the primary transformation and secondary transformation, which will be described later, are applied simultaneously, the transformation coefficients can represent the magnitudes of the fundamental pattern components generated by the result of the secondary transformation.
[0589] Transform techniques can be determined on a block-by-block basis. A transform technique can be determined based on at least one of the predictive coding mode of the current block, the size of the current block, or the shape of the current block. For example, when the current block is coded in intra-predictive mode and the size of the current block is less than N×N, the transform technique DST can be used to perform the transform. On the other hand, when the aforementioned conditions cannot be met, the transform technique DCT can be used to perform the transform.
[0590] In the residual image, a portion of the block may not undergo 2D image transformation. This omission of 2D image transformation is called transform skipping. When transform skipping is applied, quantization can be applied to the residual values for which no transformation was performed.
[0591] After transforming the current block using DCT or DST, the transformed current block can be transformed again. In this case, the DCT- or DST-based transformation can be defined as the primary transformation, and the process of transforming the block again using the primary transformation can be called the secondary transformation.
[0592] The main transform can be performed using any of a number of transform kernel candidates. For example, the main transform can be performed using any of DCT2, DCT8, or DCT7.
[0593] Different transform kernels can be used for the horizontal and vertical directions. Information representing combinations of horizontal and vertical transform kernels can also be transmitted via a bitstream signal.
[0594] The execution units for the primary and secondary transformations will differ. For example, a primary transformation can be performed on an 8×8 block, and a secondary transformation can be performed on 4×4 sub-blocks within the transformed 8×8 block. In this case, the transformation coefficients of the remaining regions where the secondary transformation is not performed can also be set to 0.
[0595] Alternatively, a primary transformation can be performed on a 4×4 block, and a secondary transformation can be performed on an 8×8 region of the 4×4 block that includes the transformation.
[0596] Information indicating whether a secondary transformation is to be performed can be sent via signals through the bitstream.
[0597] Alternatively, the decision to perform a secondary transformation can be based on whether the horizontal and vertical transformation kernels are the same. For example, a secondary transformation can only be performed if the horizontal and vertical transformation kernels are the same. Alternatively, a secondary transformation can only be performed if the horizontal and vertical transformation kernels are different.
[0598] Alternatively, quadratic transformations are permitted only when the horizontal and vertical transformations utilize predefined transformation kernels. For example, quadratic transformations are permitted when the DCT2 transformation kernel is used for both horizontal and vertical transformations.
[0599] Alternatively, the decision to perform a quadratic transform can be based on the number of non-zero transform coefficients in the current block. For example, if the number of non-zero transform coefficients in the current block is less than or equal to a threshold, the quadratic transform can be set to not be used, and if the number of non-zero transform coefficients in the current block is greater than the threshold, the quadratic transform can be used. Alternatively, the quadratic transform can be set to be used only if the current block is coded with intra-frame prediction.
[0600] Based on the shape of the current block, the size or shape of the sub-block to be subjected to the secondary transformation can be determined.
[0601] Figure 38 and Figure 39 This is a diagram showing the sub-block to which a second transformation will be performed.
[0602] When the current block is square, after performing the primary transformation, a secondary transformation can be performed on the N×N sub-block to the upper left of the current block. For example, when the current block is an 8×8 coded block, after performing the primary transformation on the current block, a secondary transformation can be performed on the 4×4 sub-block to the upper left of the current block (see [reference]). Figure 38 ).
[0603] When the current block is a non-square shape with a width greater than four times its height, after performing the primary transformation, a secondary transformation can be performed on the (kN) × (4kN) sub-block at the top left of the current block. For example, when the current block is a non-square shape of size 16 × 4, a secondary transformation can be performed on the 2 × 8 sub-block at the top left corner of the current block after performing the primary transformation on the current block (see [reference]). Figure 39 (a)
[0604] When the current block is a non-square shape with a height greater than four times its width, after performing the primary transformation, a secondary transformation can be performed on the (4kN)×(kN) sub-block at the top left of the current block. For example, when the current block is a non-square shape of size 16×4, a secondary transformation can be performed on the 2×8 sub-block at the top left corner of the current block after performing the primary transformation on the current block (see [reference]). Figure 39 (b)
[0605] The decoder can perform the inverse of the second inverse transform (second inverse transform), and the result can be subjected to the inverse of the main transform (first inverse transform). The residual signal of the current block can be obtained as the result of the second inverse transform and the first inverse transform.
[0606] Quantization is used to reduce the energy of the block, and the quantization process involves dividing the transformation coefficients by a specific constant. This constant can be derived from quantization parameters, which can be defined as values between 1 and 63.
[0607] If transform and quantization are performed in the encoder, the decoder can obtain the residual block through inverse quantization and inverse transform. The decoder then adds the predicted block and the residual block together to obtain the reconstructed block of the current block.
[0608] Information indicating the transformation type of the current block can be transmitted via signals through the bitstream. This information may be an index, tu_mts_idx, indicating a combination of horizontal and vertical transformation types.
[0609] The vertical and horizontal transform kernels can be determined based on the transform type candidates specified by the index information tu_mts_idx. Tables 9 and 10 show the combinations of transform types based on tu_mts_idx.
[0610] Table 9
[0611]
[0612]
[0613] Table 10
[0614]
[0615] The transform type can be determined as DCT2, DST7, DCT8, or a skip transform. Alternatively, in addition to transform skipping, transform type combination candidates can be constructed using only the transform kernel.
[0616] When using Table 9, if tu_mts_idx is 0, transformation skipping can be applied in both the horizontal and vertical directions. If tu_mts_idx is 1, DCT2 can be applied in both the horizontal and vertical directions. If tu_mts_idx is 3, DCT8 can be applied in the horizontal direction and DCT7 in the vertical direction.
[0617] When using Table 10, DCT2 can be applied in both the horizontal and vertical directions when tu_mts_idx is 0. If tu_mts_idx is 1, transform skipping can be applied in both the horizontal and vertical directions. If tu_mts_idx is 3, DCT8 can be applied in the horizontal direction and DCT7 in the vertical direction.
[0618] Whether to encode index information can be determined based on at least one of the following: the size, shape, or number of non-zero coefficients of the current block. For example, if the number of non-zero coefficients is equal to or less than a threshold, index information is not sent, and a default transform type can be applied to the current block. The default transform type could be DST7. Alternatively, the default transform type may differ depending on the size, shape, or intra-prediction mode of the current block.
[0619] The threshold can be determined based on the size or shape of the current block. For example, the threshold can be set to 2 when the size of the current block is less than or equal to 32×32, and to 4 when the current block is greater than 32×32 (e.g., when the current block is a coded block of size 32×64 or 64×32).
[0620] Multiple lookup tables can be pre-stored in the encoder / decoder. In these multiple lookup tables, at least one of the following can be different: the index value assigned to a transform type combination candidate, the type of transform type combination candidate, or the number of transform type combination candidates.
[0621] The lookup table for the current block can be selected based on at least one of the following: the size and shape of the current block, the predictive coding mode, the intra-prediction mode, whether a second transform is applied, or whether the transform is skipped and applied to adjacent blocks.
[0622] For example, when the current block size is 4×4 or smaller, or when the current block is encoded by inter-frame prediction, the lookup table in Table 9 can be used, and when the current block size is greater than 4×4, or when the current block is encoded by intra-block prediction, the lookup table in Table 10 can be used.
[0623] Alternatively, information indicating any of the multiple lookup tables can be sent via signaling in the bitstream. The decoder can then select the lookup table for the current block based on this information.
[0624] As another example, the index assigned to a transform type combination candidate can be adaptively determined based on at least one of the following: the size and shape of the current block, the predictive coding mode, the intra-prediction mode, whether a second transform is applied, or whether a transform skip is applied to adjacent blocks. For example, when the current block size is 4×4, the index assigned to a transform skip can have a smaller value than the index assigned to a transform skip when the current block size is greater than 4×4. Specifically, when the current block size is 4×4, an index 0 can be assigned to a transform skip; when the current block size is greater than 4×4 but less than 16×16, an index greater than 0 (e.g., index 1) can be assigned to a transform skip. When the current block size is greater than 16×16, a maximum value (e.g., 5) can be assigned as the index for a transform skip.
[0625] Alternatively, when the current block is coded with inter-frame prediction, index 0 can be assigned to the transform skip. When the current block is coded with intra-frame prediction, an index greater than 0 (e.g., index 1) can be assigned to the transform skip.
[0626] Alternatively, when the current block is a 4×4 block coded with inter-frame prediction, index 0 can be skipped in the transform. On the other hand, when the current block is not coded with inter-frame prediction, or when the current block is larger than 4×4, an index with a value greater than 0 (e.g., index 1) can be skipped in the transform.
[0627] Transform type combination candidates that differ from those listed in Tables 9 and 10 can be defined and used. For example, a transform type combination candidate with a transform kernel such as DCT7, DCT8, or DST2 can be used to apply transform skipping to horizontal or vertical transforms, as well as other transforms. In this case, it can be determined whether to use transform skipping as a transform type candidate for the horizontal or vertical direction based on at least one of the current block size (e.g., width and / or height), shape, predictive coding mode, or intra-frame prediction mode.
[0628] Alternatively, information indicating whether a specified transform type candidate is available can be transmitted via signaling in the bitstream. For example, flags indicating whether a transform can be skipped as a transform type candidate for the horizontal and vertical directions can be transmitted via signaling. Based on these flags, it can be determined whether a specified transform type combination candidate is included among a plurality of transform type combination candidates.
[0629] Alternatively, information indicating whether to apply a transform type candidate to the current block can be sent via signaling in the bitstream. For example, a flag cu_mts_flag indicating whether to apply DCT2 to the horizontal and vertical directions can be sent via signaling. When cu_mts_flag is 1, DCT2 can be set as the transform kernel for both the vertical and horizontal directions. When cu_mts_flag is 0, DCT8 or DST7 can be set as the transform kernel for both the vertical and horizontal directions. Alternatively, when cu_mts_flag is 0, information tu_mts_idx can be sent via signaling to specify any one of several transform type combination candidates.
[0630] When the current block is a non-square with a width greater than its height or a non-square with a height greater than its width, the encoding of cu_mts_flag can be omitted, and the value of cu_mts_flag can be regarded as 0.
[0631] The number of available transform type combination candidates can be set differently depending on the size, shape, or intra-prediction mode of the current block. For example, more than three transform type combination candidates can be used when the current block is square, while two transform type combination candidates can be used when the current block is non-square. Alternatively, when the current block is square, only transform type combination candidates with different transform types in the horizontal and vertical directions can be utilized.
[0632] When there are three or more transform type combination candidates available for the current block, an index information tu_mts_idx indicating one of the transform type combination candidates can be sent using a signal. On the other hand, when there are two transform type combination candidates available for the current block, a flag mts_flag indicating any one of the transform type combination candidates can be sent using a signal. Table 11 below illustrates the process for encoding information for a specified transform type combination candidate based on the shape of the current block.
[0633] Table 11
[0634]
[0635] Based on the shape of the current block, the indices of the candidate transformation type combinations are rearranged (or reordered). For example, the indices assigned to the candidate transformation type combinations when the current block is a square can be different from those assigned when the current block is not a square. For example, when the current block is a square, the transformation type combinations can be selected based on Table 12 below, and when the current block is not a square, the transformation type combinations can be selected based on Table 13 below.
[0636] Table 12
[0637]
[0638] Table 13
[0639]
[0640] The transform type can be determined based on the number of non-zero coefficients in the horizontal or vertical direction of the current block. The number of non-zero coefficients in the horizontal direction represents the number of non-zero coefficients included in 1×N (where N is the width of the current block), and the number of non-zero coefficients in the vertical direction represents the number of non-zero coefficients included in N×1 (where N is the height of the current block). When the maximum value of the non-zero coefficients in the horizontal direction is less than or equal to a threshold, a primary transform type can be applied in the horizontal direction; when the maximum value of the non-zero coefficients in the horizontal direction is greater than the threshold, a secondary transform type can be applied in the horizontal direction. Similarly, when the maximum value of the non-zero coefficients in the vertical direction is less than or equal to a threshold, a primary transform type can be applied in the vertical direction; when the maximum value of the non-zero coefficients in the vertical direction is greater than the threshold, a secondary transform type can be applied in the vertical direction.
[0641] Figure 40 This is a diagram used to illustrate an example of determining the transformation type of the current block.
[0642] For example, when the current block is encoded via intra-frame prediction and the maximum value of the non-zero coefficients in the horizontal direction of the current block is 2 or less (see...). Figure 40 (a) can be used to determine the transformation type in the horizontal direction as DST7.
[0643] When the current block is encoded via intra-frame prediction and the maximum value of the non-zero coefficients in the vertical direction of the current block is greater than 2 (see...) Figure 40 (b) can determine the transformation type in the vertical direction as DCT2 or DCT8.
[0644] The residual coefficients can be encoded by transform unit or sub-transform unit. Here, residual coefficients refer to transform coefficients generated after transformation, transform skip coefficients generated after transform skipping, or quantized coefficients generated by quantizing the coefficients.
[0645] A transform unit can represent a block that performs a primary or secondary transform. A sub-transform unit represents a block smaller than a transform unit. For example, a sub-transform unit can be a block of size 4×4, 2×8, or 8×2.
[0646] At least one of the size or shape of the sub-transform unit can be determined based on the size or shape of the current block. For example, when the current block is a non-square with a width greater than its height, the sub-transform unit can also be set to a non-square with a width greater than its height (e.g., 8×2). When the current block is a non-square with a height greater than its width, the sub-transform unit can also be set to a non-square with a height greater than its width (e.g., 2×8). When the current block is a square, the sub-transform unit can also be set to a square (e.g., 4×4).
[0647] When the current block contains multiple sub-transform units, the sub-transform units can be encoded / decoded sequentially. Arithmetic coding with isoentropy can be used to encode the residual coefficients. The encoding / decoding method for the residual coefficients will be examined in detail below with reference to the accompanying diagram.
[0648] Figure 41 This is a flowchart illustrating a method for encoding residual coefficients.
[0649] In this embodiment, it is assumed that the current block contains more than one sub-transformation unit. Furthermore, it is assumed that the sub-transformation unit has a size of 4×4. However, this embodiment can still be applied when the size or shape of the sub-transformation unit differs from this.
[0650] It can be determined whether there are non-zero coefficients in the current block (S4101). Non-zero coefficients represent residual coefficients with an absolute value greater than 0. Information indicating the presence of non-zero coefficients in the current block can be sent and encoded using a signal. For example, this information could be a 1-bit flag CBF (Coded Block Flag).
[0651] When a non-zero coefficient exists in the current block, it can be determined whether a non-zero coefficient exists in each sub-transformation unit (S4102). Information indicating the presence of a non-zero coefficient in each sub-transformation unit can be sent and encoded using a signal. For example, this information can be a 1-bit coded sub-block flag (CSBF). The sub-transformation unit can be encoded according to the selected scan order.
[0652] When there are non-zero coefficients in a sub-transform unit, in order to encode the residual coefficients of the sub-transform unit, the residual coefficients within the sub-transform unit can be arranged in one dimension (S4103). The residual coefficients can be arranged in one dimension according to the selected scan order.
[0653] The scanning sequence may include at least one of diagonal scanning, horizontal scanning, vertical scanning, or their reverse scanning.
[0654] Figure 42 and Figure 43It is a graph showing the arrangement of residual coefficients according to different scanning orders.
[0655] Figure 42 (a) to Figure 42 (c) indicates diagonal scan, horizontal scan, and vertical scan. Figure 43 (a) to Figure 43 (c) indicates their reverse direction.
[0656] Depending on the selected scan order, the residual coefficients can be arranged in one dimension.
[0657] The scan order can be determined by considering at least one of the following: the size and shape of the current block, the intra-prediction mode, the transform kernel used for the first transform, or whether a second transform is applied. For example, when the current block is a non-square with a width greater than its height, a reverse horizontal scan can be used to encode the residual coefficients. On the other hand, when the current block is a non-square with a height greater than its width, a reverse vertical scan can be used to encode the residual coefficients.
[0658] Alternatively, rate-distortion optimization (RDO) can be calculated separately for multiple scan sequences, and the scan sequence with the lowest RDO can be determined as the scan sequence for the current block. In this case, information indicating the scan sequence for the current block can be sent and encoded using signals.
[0659] The scan order candidates available for encoding transform skip coefficients can be different from those available for encoding transform coefficients. For example, the scan order candidates available for encoding transform coefficients can be set in reverse order to the scan order candidates available for encoding transform skip coefficients.
[0660] For example, if the transform coefficients are encoded using any one of a reverse diagonal scan, a reverse horizontal scan, or a reverse vertical scan, then the transform skip coefficients can be encoded using any one of a diagonal scan, a horizontal scan, or a vertical scan.
[0661] Then, the position of the last non-zero coefficient in the scan sequence within the transform block can be encoded (S4103). The x-axis position and y-axis position of the last non-zero coefficient can be encoded separately.
[0662] Figure 44 An example of encoding the position of the last non-zero coefficient is shown.
[0663] like Figure 44In the example shown, when based on diagonal scanning, the residual coefficient located at the top right corner of the transform block can be set as the last non-zero coefficient. The x-axis coordinate LastX of the residual coefficient is 3, and the y-axis coordinate LastY is 0. LastX can be encoded by separating the prefix part last_sig_coeff_x_prefix and the suffix part last_sig_coeff_x_suffix. LastY can also be encoded by separating the prefix part last_sig_coeff_y_prefix and the suffix part last_sig_coeff_y_suffix.
[0664] When determining the position of the last non-zero coefficient, information indicating whether the residual coefficient is non-zero can be encoded for each residual coefficient in the scan sequence preceding the last non-zero coefficient (S4104). This information can be a 1-bit flag, sig_coeff_flag. When the residual coefficient is 0, the value of the non-zero coefficient flag sig_coeff_flag can be set to 0; when the residual coefficient is 1, the value of the non-zero coefficient flag sig_coeff_flag can be set to 1. For the last non-zero coefficient, the encoding of the non-zero coefficient flag sig_coeff_flag can be omitted. When the encoding of the non-zero coefficient flag sig_coeff_flag is omitted, the residual coefficient will be considered non-zero.
[0665] For transform skip coefficients, information about the position of the last non-zero coefficient may not need to be encoded. In this case, the non-zero coefficient flag sig_coeff_flag can be encoded for all transform skip coefficients within the block.
[0666] Alternatively, information indicating the position of the last non-zero coefficient can be transmitted and encoded using a signal. This information can be a 1-bit flag. The decoder can determine whether to decode the information indicating the position of the last non-zero coefficient based on the value of the flag. When the position of the last non-zero coefficient is not decoded, the non-zero coefficient flag can be decoded for all transform skip coefficients. On the other hand, when the position of the last non-zero coefficient is decoded, information indicating whether a transform skip coefficient is non-zero can be encoded for each transform skip coefficient whose scan order precedes the last non-zero coefficient.
[0667] The DC component of the transform coefficient or transform skip coefficient may not be set to zero. The DC component can represent the last sample in the scan sequence or the sample located at the top left of the block. For the DC component, the encoding of the non-zero coefficient flag sig_coeff_flag can be omitted from the residual coefficients. When the encoding of the non-zero coefficient flag sig_coeff_flag is omitted, the residual coefficients will be treated as non-zero.
[0668] When the residual coefficients are not zero, information indicating the absolute value of the residual coefficients and information indicating the sign of the residual coefficients can be encoded (S4105). The encoding process for information indicating the absolute value of the residual coefficients is through... Figure 45 More detailed explanation.
[0669] Residual coefficient encoding can be performed sequentially on all residual coefficients after the last non-zero coefficient and the sub-transform units included in the current block (S4106, S4107).
[0670] Figure 45 This is a flowchart of the process of encoding the absolute value of the residual coefficients.
[0671] When the absolute value of the residual coefficient is greater than 0, information indicating whether the residual coefficient or the residual coefficient is even or odd can be encoded (S4501). The residual coefficient can be defined as the value of the absolute value of the residual coefficient minus 1. This information can be a 1-bit parity flag. For example, a value of 0 for the syntax element `par_level_flag` indicates that the residual coefficient or the residual coefficient is even, and a value of 1 for the syntax element `par_level_flag` indicates that the residual coefficient or the residual coefficient is odd. The value of the syntax element `par_level_flag` can be determined based on the following equation 30.
[0672] Equation 30
[0673] par_level_flag=(abs(Tcoeff)-1)%2
[0674] In Equation 30, Tcoeff represents the residual coefficient, and abs() represents the absolute value function.
[0675] The adjusted residual coefficient can be derived by dividing the residual coefficient by 2 or by applying a shift operation to the residual coefficient (S4502). Specifically, the entropy of dividing the residual coefficient by 2 can be set as the adjusted residual coefficient, or the value of shifting the residual coefficient to the right by about 1 can be set as the adjusted residual coefficient.
[0676] For example, the adjusted residual coefficients can be derived based on the following equation 31.
[0677] Equation 31
[0678] ReRemLevel = RemLevel >> 1
[0679] In Equation 31, ReRemLevel represents the adjusted residual coefficient, and RemLevel represents the residual coefficient.
[0680] Then, the information indicating the magnitude of the adjustment residual coefficient can be encoded. This information can include whether the adjustment residual coefficient is greater than N. N can be an integer such as 1, 2, 3, or 4.
[0681] For example, information indicating whether the value of the adjusted residual coefficient is greater than 1 can be encoded (S4503). This information can be a 1-bit flag, rem_abs_gt1_flag.
[0682] When the residual coefficient is 0 or 1, the value of rem_abs_gt1_flag can be set to 0. That is, when the residual coefficient is 1 or 2, the value of rem_abs_gt1_flag can be set to 0. In this case, when the residual coefficient is 0, par_level_flag can be set to 0, and when the residual coefficient is 1, par_level_flag can be set to 1.
[0683] When the adjusted residual coefficient is 2 or greater (S4504), the value of rem_abs_gt1_flag is set to 1, and information indicating whether the adjusted residual coefficient is greater than 2 can be encoded (S4505). This information can be a 1-bit flag, rem_abs_gt2_flag.
[0684] When the residual coefficient is 2 or 3, the value of rem_abs_gt2_flag can be set to 0. That is, when the residual coefficient is 3 or 4, the value of rem_abs_gt2_flag can be set to 0. At this time, when the residual coefficient is 2, par_level_flag can be set to 0, and when the residual coefficient is 3, par_level_flag can be set to 1.
[0685] When the adjusted residual coefficient is greater than 2, the residual value information obtained by subtracting 2 from the adjusted residual coefficient can be encoded (S4506). That is, after subtracting 5 from the absolute value of the residual coefficient, the value of dividing the result by 2 can be encoded as the residual value information.
[0686] Although not illustrated, the residual coefficients can be further encoded using flags indicating whether the adjusted residual coefficients are greater than 3 (e.g., rem_abs_gt3_flag) or flags indicating whether the adjusted residual coefficients are greater than 4 (e.g., rem_abs_gt4_flag). In this case, the residual value can be set as the value obtained by subtracting the maximum value from the adjusted residual coefficients. The maximum value represents the maximum N value in rem_abs_gtN_flag.
[0687] Alternatively, a flag that compares the absolute value of the adjustment factor or residual coefficient with a specified value can be used instead of a flag that compares the value of the adjustment residual coefficient with a specified value. For example, instead of rem_abs_gt1_flag, you can use gr2_flag to indicate whether the absolute value of the residual coefficient is greater than 2, and instead of rem_abs_gt2_flag, you can use gr4_flag to indicate whether the absolute value of the residual coefficient is greater than 4.
[0688] Table 14 briefly illustrates the process of encoding residual coefficients using syntax elements.
[0689] Table 14
[0690]
[0691]
[0692] As described in the example, residual coefficients can be encoded using par_level_flag and at least one rem_abs_gtN_flag (where N is an integer such as 1, 2, 3, 4, etc.).
[0693] `rem_abs_gtN_flag` indicates whether the residual coefficients are greater than 2N. For residual coefficients of 2N-1 or 2N, `rem_abs_gt(N-1)_flag` is set to true, and `rem_abs_gtN_flag` is set to false. Additionally, for residual coefficients of 2N-1, `par_level_flag` can be set to 0, and for residual coefficients of 2N, `par_level_flag` can be set to 1. That is, for residual coefficients below 2N, `rem_abs_gtN_flag` and `par_level_flag` can be used for encoding.
[0694] For residual coefficients greater than 2MAX+1, the residual value information obtained by dividing the difference between N and 2MAX by 2 can be encoded. Here, MAX indicates the maximum value of N. For example, when using rem_abs_gt1_flag and rem_abs_gt2_flag, MAX can be 2.
[0695] The decoder can also be based on Figure 45 The residual coefficients are decoded in the order shown. Specifically, the decoder determines the position of the last non-zero coefficient and can decode sig_coeff_flag for each residual coefficient that is scanned before the last non-zero coefficient.
[0696] When `sig_coeff_flag` is true, the `par_level_flag` of the residual coefficients can be decoded. Additionally, the `rem_abs_gt1_flag` of the residual coefficients can be decoded. Then, based on the value of `rem_abs_gt(N-1)_flag`, `rem_abs_gtN_flag` can be further decoded. For example, when the value of `rem_abs_gt(N-1)_flag` is 1, `rem_abs_gtN_flag` can be decoded. Similarly, when the value of `rem_abs_gt1_flag` related to the residual coefficients is 1, the value of `rem_abs_gt2_flag` related to the residual coefficients can be further parsed. When the value of `rem_abs_gt(MAX)_flag` is 1, the residual value information can be decoded.
[0697] When the value of `rem_abs_gtN_flag` for the residual coefficient is 0, the value of the residual coefficient can be determined as 2N-1 or 2N based on the value of `par_level_flag`. Specifically, when `par_level_flag` is 0, the residual coefficient can be set to 2N-1, and when `par_level_flag` is 1, the residual coefficient can be set to 2N.
[0698] For example, when rem_abs_gt1_flag is 0, the absolute value of the residual coefficient can be set to 1 or 2 depending on the par_level_flag value. Specifically, when par_level_flag is 0, the absolute value of the residual coefficient is 1, and when par_level_flag is 1, the absolute value of the residual coefficient is 2.
[0699] For example, when rem_abs_gt2_flag is 0, the absolute value of the residual coefficient can be set to 3 or 4 based on the par_level_flag value. Specifically, when par_level_flag is 0, the absolute value of the residual coefficient is 3, and when par_level_flag is 1, the absolute value of the residual coefficient is 4.
[0700] When decoding the residual information, the residual coefficients can be set to 2(MAX+R)-1 or 2(MAX+R) based on the value of par_level_flag. Here, R indicates the value shown in the residual information. For example, when par_level_flag is 0, the residual coefficients can be set to 2(MAX+R)-1, and when par_level_flag is 1, the residual coefficients can be set to 2(MAX+R). For example, when MAX is 2, the residual coefficients can be derived based on the following equation 32.
[0701] Equation 32
[0702] parity_level_flag+Rem<<1+5
[0703] According to such Figure 45 The example shown illustrates that when encoding residual coefficients, the parity flag needs to be encoded for all non-zero coefficients. For instance, when the residual coefficient is 1, both `par_level_flag` and `rem_abs_gt1_flag` also need to be encoded. This encoding method, as described above, increases the number of bits required to encode residual coefficients with an absolute value of 1. To prevent this, information indicating whether the residual coefficient is greater than 1 should be encoded first, and then the parity flag can be encoded when the residual coefficient is greater than 1.
[0704] Figure 46 This is a flowchart of the process of encoding the absolute value of the residual coefficients.
[0705] For non-zero residual coefficients, information indicating whether the absolute value of the residual coefficient is greater than 1 can be encoded as gr1_flag (S4601). When the residual coefficient is 1, gr1_flag can be set to 0, and when the residual coefficient is greater than 1, gr1_flag can be set to 1.
[0706] When the absolute value of the residual coefficient is greater than 1, a parity flag indicating whether the residual coefficient is even or odd can be encoded (S4602, S4603). Specifically, the residual coefficient can be set to a value obtained by subtracting 2 from the residual coefficient. For example, the par_level_flag can be derived based on the following equation 33.
[0707] Equation 33
[0708] par_leyel_flag=(abs(Tcoeff)-2)%2
[0709] The adjusted residual coefficient can be derived by dividing the residual coefficient by 2 or shifting the residual coefficient to the right by 1, and information indicating whether the adjusted residual coefficient is greater than 1 can be encoded (S4604). For example, for a residual coefficient where gr1_flag is 1, rem_abs_gt1_flag indicating whether the adjusted residual coefficient is greater than 1 can be encoded.
[0710] When the residual coefficient is 0 or 1, the value of rem_abs_gt1_flag can be set to 0. That is, when the residual coefficient is 2 or 3, the value of rem_abs_gt1_flag can be set to 0. In this case, when adjusting the residual coefficient to 0, par_level_flag can be set to 0, and when the residual coefficient is 1, par_level_flag can be set to 1.
[0711] When the adjusted residual coefficient is 2 or higher (S4605), the value of rem_abs_gt1_flag can be set to 1, and information indicating whether the adjusted residual coefficient is greater than 2 can be encoded (S4606). For example, for a residual coefficient where rem_abs_gt1_flag is 1, rem_abs_gt2_flag indicating whether the adjusted residual coefficient is greater than 2 can be encoded.
[0712] When the residual coefficient is 2 or 3, the value of `rem_abs_gt2_flag` can be set to 1. That is, when the residual coefficient is 4 or 5, the value of `rem_abs_gt2_flag` can be set to 0. In this case, when the residual coefficient is 2, `par_level_flag` can be set to 0, and when the residual coefficient is 3, `par_level_flag` can be set to 1.
[0713] When the adjusted residual coefficient is greater than 2 (S4607), the residual value information obtained by subtracting 2 from the adjusted residual coefficient can be encoded (S4608). That is, after subtracting 6 from the absolute value of the residual coefficient, the result value divided by 2 can be encoded as the residual value information.
[0714] The residual coefficients can be further encoded using flags indicating whether the adjusted residual coefficients are greater than 3 (e.g., rem_abs_gt3_flag) or flags indicating whether the adjusted residual coefficients are greater than 4 (e.g., rem_abs_gt4_flag). In this case, the residual value can be set as the value obtained by subtracting the maximum value from the adjusted residual coefficients. The maximum value indicates the maximum N value in rem_abs_gtN_flag.
[0715] Alternatively, a flag that compares the absolute value of the adjustment factor or residual factor with a specified value can be used instead of a flag that compares the value of the adjustment residual factor with a specified value. For example, instead of rem_abs_gt1_flag, you can use gr3_flag to indicate whether the absolute value of the residual factor is greater than 3, and instead of rem_abs_gt2_flag, you can use gr5_flag to indicate whether the absolute value of the residual factor is greater than 5.
[0716] Table 15 briefly describes the process of encoding residual coefficients using syntax elements.
[0717] Table 15
[0718]
[0719] The decoder can also be based on Figure 46The residual coefficients are decoded in the order shown. Specifically, the decoder determines the position of the last non-zero coefficient and can decode sig_coeff_flag for each residual coefficient that is scanned before the last non-zero coefficient.
[0720] When sig_coeff_flag is true, the gr1_flag of the residual coefficients can be decoded. When gr1_flag is 0, the absolute value of the residual coefficients is determined to be 1. When gr1_flag is 1, the par_level_flag of the residual coefficients can be decoded. Then, the rem_abs_gt1_flag of the residual coefficients can be decoded. At this point, based on the value of rem_abs_gt(N-1)_flag, rem_abs_gtN_flag can be further decoded. For example, when the value of rem_abs_gt(N-1)_flag is 1, rem_abs_gtN_flag can be decoded. For example, when the value of rem_abs_gt1_flag related to the residual coefficients is 1, the rem_abs_gt2_flag related to the residual coefficients can be further parsed. When the value of rem_abs_gt(MAX)_flag is 1, the residual value information can be decoded.
[0721] When the value of `rem_abs_gtN_flag` for the residual coefficient is 0, the value of the residual coefficient can be determined as 2N or 2N+1 based on the value of `par_level_flag`. Specifically, when `par_level_flag` is 0, the residual coefficient can be set to 2N, and when `par_level_flag` is 1, the residual coefficient can be set to 2N+1.
[0722] For example, when rem_abs_gt1_flag is 0, the absolute value of the residual coefficient can be set to 2 or 3 depending on the par_level_flag value. Specifically, when par_level_flag is 0, the absolute value of the residual coefficient is 2, and when par_level_flag is 1, the absolute value of the residual coefficient is 3.
[0723] For example, when rem_abs_gt2_flag is 0, the absolute value of the residual coefficient can be set to 4 or 5 depending on the par_level_flag value. Specifically, when par_level_flag is 0, the absolute value of the residual coefficient is 4, and when par_level_flag is 1, the absolute value of the residual coefficient is 5.
[0724] When decoding residual information, the residual coefficient can be set to 2(MAX+R) or 2(MAX+R)+1 based on the value of par_level_flag. Here, R represents the value of the residual information. For example, when par_level_flag is 0, the residual coefficient can be set to 2(MAX+R), and when par_level_flag is 1, the residual coefficient can be set to 2(MAX+R)+1.
[0725] As another example, the parity flag can be encoded only when the residual coefficient value is greater than 2. For instance, after encoding information indicating whether the residual coefficient value is greater than 1 and information indicating whether the residual coefficient value is greater than 2, if it is determined that the residual coefficient value is greater than 2, the parity flag can be encoded for the residual coefficient.
[0726] Figure 47 This is a flowchart of the process of encoding the absolute value of the residual coefficients.
[0727] For non-zero residual coefficients, the information gr1_flag indicating whether the absolute value of the residual coefficient is greater than 1 can be encoded (S4701). When the residual coefficient is 1, gr1_flag can be set to 0, and when the residual coefficient is greater than 1, gr1_flag can be set to 1.
[0728] For residual coefficients with an absolute value greater than 1, the information gr2_flag indicating whether the absolute value of the residual coefficient is greater than 2 can be encoded (S4702, S4703). When the residual coefficient is 2, gr2_flag can be set to 0, and when the residual coefficient is greater than 2, gr2_flag can be set to 1.
[0729] When the absolute value is greater than 2, a parity flag indicating whether the residual coefficients are even or odd can be encoded (S4704, S4705). The residual coefficients can be set to a value obtained by subtracting 3 from the residual coefficients. For example, the par_level_flag can be derived based on the following equation 34.
[0730] Equation 34
[0731] par_level_flag=(abs(Tcoeff)-3)%2
[0732] The adjusted residual coefficient can be derived by dividing the residual coefficient by 2 or shifting the residual coefficient to the right by 1, and information indicating whether the adjusted residual coefficient is greater than 1 can be encoded. For example, for a residual coefficient where gr1_flag is 1, the rem_abs_gt1_flag indicating whether the adjusted residual coefficient is greater than 1 can be encoded (S4706).
[0733] When the residual coefficient is 0 or 1, the value of rem_abs_gt1_flag can be set to 0. That is, when the residual coefficient is 3 or 4, the value of rem_abs_gt1_flag can be set to 0. In this case, when adjusting the residual coefficient to 0, par_level_flag can be set to 0, and when the residual coefficient is 1, par_level_flag can be set to 1.
[0734] When the adjusted residual coefficient is 2 or higher (S4707), the value of rem_abs_gt1_flag can be set to 1, and information indicating whether the adjusted residual coefficient is greater than 2 can be encoded (S4708). For example, for a residual coefficient where rem_abs_gt1_flag is 1, rem_abs_gt2_flag indicating whether the adjusted residual coefficient is greater than 2 can be encoded.
[0735] When the residual coefficient is 2 or 3, the value of rem_abs_gt2_flag can be set to 1. That is, when the residual coefficient is 5 or 6, the value of rem_abs_gt2_flag can be set to 0. In this case, when the residual coefficient is 2, par_level_flag can be set to 0, and when the residual coefficient is 3, par_level_flag can be set to 1.
[0736] When the adjusted residual coefficient is greater than 2, the residual value obtained by subtracting 2 from the adjusted residual coefficient can be encoded (S4709, S4710). That is, after subtracting 7 from the absolute value of the residual coefficient, the result value divided by 2 can be encoded as the residual value.
[0737] The residual coefficients can be further encoded using flags indicating whether the adjusted residual coefficients are greater than 3 (e.g., rem_abs_gt3_flag) or flags indicating whether the adjusted residual coefficients are greater than 4 (e.g., rem_abs_gt4_flag). In this case, the residual value can be set as the value obtained by subtracting the maximum value from the adjusted residual coefficients. The maximum value indicates the maximum N value in rem_abs_gtN_flag.
[0738] Alternatively, a flag that compares the absolute value of the adjustment factor or residual coefficient with a specified value can be used instead of a flag that compares the value of the adjustment residual coefficient with a specified value. For example, instead of rem_abs_gt1_flag, you can use gr4_flag, which indicates whether the absolute value of the residual coefficient is greater than 4, and instead of rem_abs_gt2_flag, you can use gr6_flag, which indicates whether the absolute value of the residual coefficient is greater than 6.
[0739] The decoder is also pressed by AND. Figure 47 The order shown is the same, and the residual coefficients can be decoded. Specifically, the decoder determines the position of the last non-zero coefficient and can decode sig_coeff_flag for each residual coefficient that is scanned before the last non-zero coefficient.
[0740] When sig_coeff_flag is true, the gr1_flag of the residual coefficients can be decoded. When gr1_flag is 0, the absolute value of the residual coefficients is determined to be 1; when gr1_flag is 1, gr2_flag can be decoded. When gr2_flag is 0, the absolute value of the residual coefficients is determined to be 2; when gr2_flag is 1, the par_level_flag of the residual coefficients can be decoded. Then, the rem_abs_gt1_flag of the residual coefficients can be decoded. At this point, based on the value of rem_abs_gt(N-1)_flag, rem_abs_gtN_flag can be further decoded. For example, when the value of rem_abs_gt(N-1)_flag is 1, rem_abs_gtN_flag can be decoded. For example, when the value of rem_abs_gt1_flag related to the residual coefficients is 1, the value of rem_abs_gt2_flag related to the residual coefficients can be further parsed. When the value of rem_abs_gt(MAX)_flag is 1, the residual value information can be decoded.
[0741] When the value of rem_abs_gtN_flag for the residual coefficient is 0, the value of the residual coefficient can be determined as 2N+1 or 2(N+1) based on the value of par_level_flag. Specifically, when par_level_flag is 0, the residual coefficient can be set to 2N+1, and when par_level_flag is 1, the residual coefficient can be set to 2(N+1).
[0742] For example, when rem_abs_gt1_flag is 0, the absolute value of the residual coefficient can be set to 3 or 4 depending on the par_level_flag value. Specifically, when par_level_flag is 0, the absolute value of the residual coefficient is 3, and when par_level_flag is 1, the absolute value of the residual coefficient is 4.
[0743] For example, when rem_abs_gt2_flag is 0, the absolute value of the residual coefficient can be set to 5 or 6 based on the par_level_flag value. Specifically, when par_level_flag is 0, the absolute value of the residual coefficient is 5, and when par_level_flag is 1, the absolute value of the residual coefficient is 6.
[0744] When decoding residual information, the residual coefficient can be set to 2(MAX+R) or 2(MAX+R)+1 based on the value of par_level_flag. Here, R represents the value of the residual information. For example, when par_level_flag is 0, the residual coefficient can be set to 2(MAX+R), and when par_level_flag is 1, the residual coefficient can be set to 2(MAX+R)+1.
[0745] The number or type of comparison flags to be compared with the specified value can be determined based on at least one of the following: the size and shape of the current block, whether the transform is skipped, the transform kernel, the number of non-zero coefficients, or the position of the last non-zero coefficient. For example, when encoding transform coefficients, only `rem_abs_gt1_flag` can be used. On the other hand, when encoding transform skip coefficients, both `rem_abs_gt1_flag` and `rem_abs_gt2_flag` can be used.
[0746] Alternatively, within the diagonal direction of a 4×4 subtransformation, the number of `rem_abs_gt1_flag` flags can be set to a maximum of 8, and the number of `rem_abs_gt2_flag` flags can be set to a maximum of 1. Alternatively, when the number of non-zero coefficient flags is (16-N), the number of `rem_abs_gt1_flag` flags can be set to a maximum of 8+(N / 2), and the number of `rem_abs_gt2_flag` flags can be set to a maximum of 1+(N-(N / 2)).
[0747] If the reconstructed block of the current block is obtained, in-loop filtering can be used to reduce information loss during quantization and encoding. The in-loop filter can include at least one of a deblocking filter, a sample adaptive offset filter (SAO), or an adaptive loop filter (ALF). Hereinafter, the reconstructed block before applying the in-loop filter will be referred to as the first reconstructed block, and the reconstructed block after applying the in-loop filter will be referred to as the second reconstructed block.
[0748] A second reconstructed block can be obtained by applying at least one of a deblocking filter, SAO, or ALF to the first reconstructed block. In this case, SAO or ALF can be applied after the deblocking filter.
[0749] Deblocking filters are used to mitigate the image quality degradation (blocking artifact) that occurs at block boundaries when quantization is performed on a block-by-block basis. To apply a deblocking filter, the block strength (BS) between the first reconstructed block and its adjacent reconstructed blocks can be determined.
[0750] Figure 48 This is a flowchart illustrating the process of determining block strength.
[0751] exist Figure 48 In the example shown, P represents the first reconstructed block, and Q represents the adjacent reconstructed block. The adjacent reconstructed block can be adjacent to the left or top of the current block.
[0752] exist Figure 48 The example shown illustrates how to determine block strength by considering the predictive coding patterns of P and Q, whether non-zero transform coefficients are included, whether inter-frame prediction is performed using the same reference image, and whether the difference in motion vectors is greater than or equal to a threshold.
[0753] Based on the block strength, it can be determined whether a deblocking filter has been applied. For example, if the block strength is 0, filtering may not be performed.
[0754] SAO (Sound Analysis and Offset) is used to mitigate the ringing artifact that occurs when performing quantization in the frequency domain. SAO can be performed by adding or subtracting an offset determined by considering the pattern of the first reconstructed image. Methods for determining the offset include Edge Offset (EO) or Band Offset (BO). EO indicates a method of determining the offset of the current sample based on the pattern of surrounding pixels. BO indicates a method of applying a common offset to a set of pixels with similar brightness values within a region. Specifically, pixel brightness is divided into 32 equal intervals, and pixels with similar brightness values are grouped together. For example, four adjacent bands out of the 32 bands are grouped together, and samples belonging to those four bands can have the same offset applied.
[0755] ALF is a method for generating a second reconstructed image by applying a predefined filter of size / shape to a first reconstructed image or a reconstructed image with a deblocking filter applied. Equation 35 below shows an example of ALF application.
[0756] Equation 35
[0757]
[0758] You can select any of the predefined filter candidates at the image, coding tree unit, coding block, prediction block, or transform block level. Each filter candidate may have a different size or shape.
[0759] Figure 49 Predefined filter candidates are shown.
[0760] As in Figure 49 In the example shown, at least one of the following rhombuses can be selected: 5×5, 7×7, and 9×9.
[0761] Only 5×5 rhombuses can be used for chromaticity components.
[0762] Embodiments described with a focus on the decoding or encoding process are also included within the scope of this invention. Variations of multiple embodiments described in a predetermined order, in a different order than those described, are also included within the scope of this invention.
[0763] The embodiments have been described based on a series of steps or flowcharts, but this does not limit the chronological order of the invention, and they can be performed simultaneously or in a different order as needed. Furthermore, in the above embodiments, the structural elements constituting the block diagrams (e.g., units, modules, etc.) can also be implemented as hardware devices or software, and multiple structural elements can be combined to implement a single hardware device or software. The embodiments can be implemented in the form of program instructions, which can be executed by various computer components and recorded in a computer-readable recording medium. The computer-readable recording medium can individually or in combination include program instructions, data files, data structures, etc. Examples of computer-readable recording media can include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floppy optical disks; and hardware devices specifically configured to store and execute program instructions, such as ROMs, RAMs, and flash memory. The hardware device can be configured to operate as one or more software modules to perform the processing according to the invention, and vice versa.
[0764] Industrial applicability
[0765] This invention can be applied to electronic devices that encode / decode video.
Claims
1. A video decoding method, comprising the following steps: Parse the non-zero flag indicating whether the residual coefficients are non-zero from the bitstream; For the position of the last non-zero coefficient, when the non-zero flag is omitted, the residual coefficient is considered non-zero; When the non-zero flag indicates that the residual coefficient is non-zero, absolute value information is parsed from the bitstream, and the absolute value information is used to determine the absolute value of the residual coefficient; and Based on the absolute value information, the absolute value of the residual coefficient is determined. The absolute value information includes a residual coefficient comparison flag indicating whether the residual coefficient is greater than the first value. Parity flags are further parsed from the bitstream only when the residual coefficient is greater than the first value; wherein the parity flag indicates whether the value of the residual coefficient is even or odd. When the residual coefficient is greater than the first value, the first adjustment residual coefficient comparison flag is further analyzed. The first adjustment residual coefficient comparison flag indicates whether the adjustment residual coefficient derived by shifting the residual coefficient to the right by 1 bit is greater than the second value. When the adjusted residual coefficient is below the second value, the residual coefficient is determined to be 2N or 2N+1 according to the value of the parity flag, where N is the second value.
2. The video decoding method according to claim 1, wherein, When the adjusted residual coefficient is greater than the second value, the second adjusted residual coefficient comparison flag is further analyzed. The second adjusted residual coefficient comparison flag indicates whether the adjusted residual coefficient is greater than the third value.
3. The video decoding method according to claim 1, wherein, When the adjusted residual coefficient is greater than the second value, further analysis of the residual value information is performed. The residual value information is obtained by subtracting the second value from the adjusted residual coefficient.
4. A video encoding method, comprising the following steps: Encode the non-zero flag indicating whether the residual coefficient is non-zero; For the position of the last non-zero coefficient, when the non-zero flag is omitted, the residual coefficient is considered non-zero; as well as When the residual coefficient is non-zero, the absolute value information is encoded, and the absolute value information is used to determine the absolute value of the residual coefficient. The absolute value information includes a residual coefficient comparison flag indicating whether the residual coefficient is greater than the first value. The parity flag of the residual coefficient is further encoded only when the residual coefficient is greater than the first value; wherein the parity flag indicates whether the value of the residual coefficient is even or odd. When the residual coefficient is greater than the first value, the first adjustment residual coefficient comparison flag is further encoded. The first adjustment residual coefficient comparison flag indicates whether the adjustment residual coefficient derived by shifting the residual coefficient to the right by 1 is greater than the second value. When the adjusted residual coefficient is below the second value, the residual coefficient is determined to be 2N or 2N+1 according to the value of the parity flag.
5. The video encoding method according to claim 4, wherein, When the adjusted residual coefficient is greater than the second value, the second adjusted residual coefficient comparison flag is further encoded, and the second adjusted residual coefficient comparison flag indicates whether the adjusted residual coefficient is greater than the third value.
6. The video encoding method according to claim 4, wherein, When the adjusted residual coefficient is greater than the second value, the residual value information is further encoded. The residual value information is obtained by subtracting the second value from the adjusted residual coefficient.
7. A video decoder, characterized in that, include: A processor and memory for storing computer programs that can run on the processor. The processor is configured to: Parse the non-zero flag indicating whether the residual coefficients are non-zero from the bitstream; For the position of the last non-zero coefficient, when the non-zero flag is omitted, the residual coefficient is considered non-zero; When the non-zero flag indicates that the residual coefficient is non-zero, absolute value information is parsed from the bitstream, and the absolute value information is used to determine the absolute value of the residual coefficient; and Based on the absolute value information, the absolute value of the residual coefficient is determined. The absolute value information includes a residual coefficient comparison flag indicating whether the residual coefficient is greater than the first value. Parity flags are further parsed from the bitstream only when the residual coefficient is greater than the first value; wherein the parity flag indicates whether the value of the residual coefficient is even or odd. The processor is also configured to: When the residual coefficient is greater than the first value, the first adjustment residual coefficient comparison flag is further analyzed. The first adjustment residual coefficient comparison flag indicates whether the adjustment residual coefficient derived by shifting the residual coefficient to the right by 1 bit is greater than the second value. The processor is also configured to: When the residual coefficient is greater than the first value, the first adjustment residual coefficient comparison flag is further analyzed. The first adjustment residual coefficient comparison flag indicates whether the adjustment residual coefficient derived by shifting the residual coefficient one position to the right is greater than the second value.
8. The video decoder according to claim 7, wherein, The processor is also configured to: When the adjusted residual coefficient is greater than the second value, the second adjusted residual coefficient comparison flag is further analyzed. The second adjusted residual coefficient comparison flag indicates whether the adjusted residual coefficient is greater than the third value.
9. The video decoder according to claim 7, wherein, The processor is also configured to: When the adjusted residual coefficient is greater than the second value, further analysis of the residual value information is performed. The residual value information is obtained by subtracting the second value from the adjusted residual coefficient.
10. A video encoder, characterized in that, include: A processor and memory for storing computer programs that can run on the processor. The processor is configured to: Encode the non-zero flag indicating whether the residual coefficient is non-zero; For the position of the last non-zero coefficient, when the non-zero flag is omitted, the residual coefficient is considered non-zero; as well as When the residual coefficient is non-zero, the absolute value information is encoded, and the absolute value information is used to determine the absolute value of the residual coefficient. The absolute value information includes a residual coefficient comparison flag indicating whether the residual coefficient is greater than the first value. The parity flag of the residual coefficient is further encoded only when the residual coefficient is greater than the first value; wherein the parity flag indicates whether the value of the residual coefficient is even or odd. The processor is also configured to: When the residual coefficient is greater than the first value, the first adjustment residual coefficient comparison flag is further encoded. The first adjustment residual coefficient comparison flag indicates whether the adjustment residual coefficient derived by shifting the residual coefficient to the right by 1 is greater than the second value. The processor is also configured to: When the adjusted residual coefficient is below the second value, the residual coefficient is determined to be 2N or 2N+1 according to the value of the parity flag.
11. The video encoder according to claim 10, wherein, The processor is also configured to: When the adjusted residual coefficient is greater than the second value, the second adjusted residual coefficient comparison flag is further encoded, and the second adjusted residual coefficient comparison flag indicates whether the adjusted residual coefficient is greater than the third value.
12. The video encoder according to claim 10, wherein, The processor is also configured to: When the adjusted residual coefficient is greater than the second value, the residual value information is further encoded. The residual value information is obtained by subtracting the second value from the adjusted residual coefficient.
13. A video decoding device, comprising: A device for parsing non-zero flags indicating whether residual coefficients are non-zero from a bitstream; For the position of the last non-zero coefficient, when the non-zero flag is omitted, the residual coefficient is considered non-zero; When the non-zero flag indicates that the residual coefficient is non-zero, the apparatus for parsing absolute value information from the bitstream, the absolute value information being used to determine the absolute value of the residual coefficient; as well as A device for determining the absolute value of the residual coefficient based on the absolute value information. The absolute value information includes a residual coefficient comparison flag indicating whether the residual coefficient is greater than the first value. Parity flags are further parsed from the bitstream only when the residual coefficient is greater than the first value; wherein the parity flag indicates whether the value of the residual coefficient is even or odd. The video decoding device also includes: When the residual coefficient is greater than the first value, the device for further analyzing the first adjustment residual coefficient comparison flag indicates whether the adjustment residual coefficient derived by shifting the residual coefficient one position to the right is greater than the second value. The video decoding device also includes: When the adjusted residual coefficient is below the second value, the device for determining the residual coefficient as 2N or 2N+1 based on the value of the parity flag, where N is the second value.
14. The video decoding device according to claim 13, wherein, The video decoding device also includes: When the adjusted residual coefficient is greater than the second value, the device for further parsing the second adjusted residual coefficient comparison flag indicates whether the adjusted residual coefficient is greater than the third value.
15. The video decoding device according to claim 13, wherein, The video decoding device also includes: When the adjusted residual coefficient is greater than the second value, the device further analyzes the residual value information. The residual value information is obtained by subtracting the second value from the adjusted residual coefficient.
16. A video encoding device, comprising: A device for encoding a non-zero flag indicating whether the residual coefficient is non-zero; For the position of the last non-zero coefficient, when the non-zero flag is omitted, the residual coefficient is considered non-zero; as well as When the residual coefficient is non-zero, the apparatus encodes absolute value information, which is used to determine the absolute value of the residual coefficient. The absolute value information includes a residual coefficient comparison flag indicating whether the residual coefficient is greater than the first value. The parity flag of the residual coefficient is further encoded only when the residual coefficient is greater than the first value; wherein the parity flag indicates whether the value of the residual coefficient is even or odd. The video encoding device also includes: When the residual coefficient is greater than the first value, the apparatus further encodes the first adjustment residual coefficient comparison flag, wherein the first adjustment residual coefficient comparison flag indicates whether the adjustment residual coefficient derived by shifting the residual coefficient to the right by 1 is greater than the second value. The video encoding device also includes: When the adjusted residual coefficient is below the second value, the device determines the residual coefficient to be 2N or 2N+1 based on the value of the parity flag.
17. The video encoding device according to claim 16, wherein, The video encoding device also includes: When the adjusted residual coefficient is greater than the second value, the apparatus further encodes a second adjusted residual coefficient comparison flag, the second adjusted residual coefficient comparison flag indicating whether the adjusted residual coefficient is greater than a third value.
18. The video encoding device according to claim 16, wherein, The video encoding device also includes: When the adjusted residual coefficient is greater than the second value, the apparatus further encodes the residual value information. The residual value information is obtained by subtracting the second value from the adjusted residual coefficient.
19. A computer storage medium having stored thereon computer-executable instructions which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 3.
20. A computer storage medium having stored thereon computer-executable instructions which, when executed by a processor, implement the steps of the method according to any one of claims 4 to 6.