Video signal encoding / decoding method and apparatus

By partitioning image blocks and utilizing neighboring pixel information to determine the intra-frame prediction mode of local blocks, efficient image compression is achieved, solving the problem of insufficient utilization of multiple intra-frame prediction modes in existing technologies and improving encoding/decoding efficiency.

CN116915997BActive Publication Date: 2026-05-22IND ACAD COOP GRP OF SEJONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IND ACAD COOP GRP OF SEJONG UNIV
Filing Date
2017-04-28
Publication Date
2026-05-22

Smart Images

  • Figure CN116915997B_ABST
    Figure CN116915997B_ABST
Patent Text Reader

Abstract

A video signal encoding / decoding method and apparatus are provided. An image signal decoding method according to the present invention includes the steps of decoding information indicating whether a current block is encoded using multi-mode intra prediction, partitioning the current block into a plurality of local blocks when it is determined that the current block is encoded by multi-mode intra prediction, and obtaining an intra prediction mode of each of the plurality of local blocks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application filed on April 28, 2017, with application number 201780039690.3 and title "Video Signal Encoding / Decoding Method and Apparatus". Technical Field

[0002] This invention relates to a method and apparatus for encoding / decoding image signals. Background Technology

[0003] Recently, the demand for multimedia data such as video has increased rapidly on the Internet. However, the rate of development of channel bandwidth is struggling to keep up with the rapidly increasing volume of multimedia data. Summary of the Invention

[0004] Technical issues

[0005] One objective of this invention is to improve image compression efficiency by using multi-frame prediction modes during image encoding / decoding.

[0006] One objective of this invention is to improve image compression efficiency by efficiently encoding / decoding the multi-frame intra-prediction mode of the target encoding / decoding block during the encoding / decoding process of the image.

[0007] One objective of this invention is to improve image compression efficiency by efficiently encoding / decoding coefficients in local blocks.

[0008] Technical solution

[0009] According to a method and apparatus for decoding an image, the present invention can decode information indicating whether the current block is encoded using a multi-frame prediction mode; when it is determined that the current block is encoded using a multi-frame prediction mode, the current block is partitioned into multiple local blocks; and the intra-frame prediction mode of each of the multiple local blocks is obtained.

[0010] According to a method and apparatus for decoding an image according to the present invention, the following operations are performed: when a current block is partitioned into a plurality of local blocks, an inflection point is determined among neighboring pixels adjacent to the current block; slope information is obtained based on a plurality of pixels adjacent to the inflection point; and the partition shape of the current block is determined based on the inflection point and the slope information.

[0011] In a method and apparatus for decoding an image according to the present invention, an inflection point can be determined based on the inflection point value of each neighboring pixel, and an inflection point value can be generated based on the difference between adjacent pixels adjacent to the neighboring pixel.

[0012] In a method and apparatus for decoding an image according to the present invention, the following operations may be performed: when an intra-prediction mode of each of the plurality of local blocks is obtained, a first intra-prediction mode of a first local block is obtained; the difference between the first intra-prediction mode and a second intra-prediction mode of a second local block is decoded; and a second intra-prediction mode is obtained based on the difference.

[0013] In a method and apparatus for decoding an image according to the present invention, the intra-frame prediction mode of each of the plurality of local blocks may have different values.

[0014] According to a method and apparatus for encoding an image according to the present invention, it is possible to determine whether to encode the current block using a multi-intra-prediction mode; based on the determination result, information indicating whether the current block uses a multi-intra-prediction mode is encoded; when the current block is set to use a multi-intra-prediction mode, the current block is partitioned into a plurality of local blocks; and the intra-prediction mode of each of the plurality of local blocks is determined.

[0015] According to a method and apparatus for encoding an image according to the present invention, the following operations are performed: when a current block is divided into multiple local blocks, an inflection point is determined among the neighboring pixels adjacent to the current block; slope information is obtained based on the multiple pixels adjacent to the inflection point; and the partition shape of the current block is determined based on the inflection point and the slope information.

[0016] In a method and apparatus for encoding an image according to the present invention, an inflection point can be determined based on the inflection point value of each neighboring pixel, and an inflection point value can be generated based on the difference between adjacent pixels adjacent to the neighboring pixel.

[0017] According to a method and apparatus for encoding an image according to the present invention, a first intra-prediction mode of a first local block among a plurality of local blocks is determined; a second intra-prediction mode of the first local block among the plurality of local blocks is determined; and the difference between the first intra-prediction mode and the second intra-prediction mode is encoded.

[0018] In a method and apparatus for encoding an image according to the present invention, the intra-frame prediction mode of each of the plurality of local blocks may have different values.

[0019] Technical effect

[0020] According to the present invention, the compression efficiency of an image can be improved by using a multi-frame prediction mode during the encoding / decoding process.

[0021] According to the present invention, the compression efficiency of an image can be improved by efficiently encoding / decoding the multi-frame intra-prediction mode of the target encoding / decoding block during the encoding / decoding process of the image.

[0022] According to the present invention, the compression efficiency of an image can be improved by efficiently encoding / decoding the coefficients in local blocks. Attached Figure Description

[0023] Figure 1 This is a block diagram illustrating an apparatus for encoding an image according to an embodiment of the present invention.

[0024] Figure 2 This is a block diagram illustrating an apparatus for decoding an image according to an embodiment of the present invention.

[0025] Figure 3 This is a diagram used to explain the intra-frame prediction method using DC mode.

[0026] Figure 4 This is a diagram used to explain intra-frame prediction methods using planar modes.

[0027] Figure 5 This is a diagram used to explain intra-frame prediction methods using directional prediction modes.

[0028] Figure 6 This is a diagram illustrating a method for encoding the coefficients of a transform block as described in an embodiment of the present invention.

[0029] Figure 7 This is a diagram illustrating a method for encoding the maximum value of the coefficients of a local block as in an embodiment to which the present invention is applied.

[0030] Figure 8 This is a diagram illustrating a method for encoding a first threshold flag for a local block as in an embodiment to which the present invention is applied.

[0031] Figure 9 This is a diagram illustrating a method for decoding the coefficients of a transform block as described in an embodiment of the present invention.

[0032] Figure 10 This is a diagram illustrating a method for decoding the maximum value of the coefficients of a local block as in an embodiment to which the present invention is applied.

[0033] Figure 11 This is a diagram illustrating a method for decoding a first threshold flag for a local block as in an embodiment to which the present invention is applied.

[0034] Figure 12 This is a diagram illustrating a method for deriving a first / second threshold flag for the current local block as in an embodiment to which the present invention is applied.

[0035] Figure 13This is a diagram illustrating a method for determining the size / shape of a local block based on a merging flag, as in an embodiment of the present invention.

[0036] Figure 14 This is a diagram illustrating a method for determining the size / shape of a local block based on partition markers, as in an embodiment of the present invention.

[0037] Figure 15 This is a diagram illustrating a method for determining the size / shape of a local block based on partition index information, as in an embodiment of the present invention.

[0038] Figure 16 This is a diagram illustrating a method for partitioning transform blocks based on the positions of non-zero coefficients, as in an embodiment of the present invention.

[0039] Figure 17 This is a diagram illustrating a method for selectively partitioning a local region of a transform block as described in an embodiment of the present invention.

[0040] Figure 18 This is a diagram illustrating a method for partitioning a transform block based on DC / AC component attributes in the frequency domain, as in an embodiment of the present invention.

[0041] Figure 19 This is a flowchart illustrating the process of determining whether to use multi-frame prediction mode during the encoding process.

[0042] Figure 20 This is a diagram illustrating an example of a search inflection point.

[0043] Figure 21 This is a diagram showing the partition shape based on the shape of the current block.

[0044] Figure 22 This is an illustration used to explain an example of calculating the slope information of neighboring pixels.

[0045] Figure 23 This is a diagram illustrating an example of overlapping regions generated based on the partition shape of the current block.

[0046] Figure 24 This is a flowchart illustrating the process of encoding the intra-prediction mode of a local block.

[0047] Figure 25 This is a flowchart illustrating a method for encoding intra-prediction modes using MPM candidates.

[0048] Figure 26 This is a diagram showing the MPM candidates for the current block.

[0049] Figure 27This is a flowchart illustrating a method for decoding the intra-prediction mode of the current block.

[0050] Figure 28 This is a flowchart illustrating a method for decoding intra-prediction modes using MPM candidates. Detailed Implementation

[0051] Various modifications can be made to this invention, and various embodiments of the invention exist. Examples of various embodiments of the invention will now be provided and described in detail with reference to the accompanying drawings. However, the invention is not limited thereto, and exemplary embodiments can be constructed to include all modifications, equivalents, or alternatives to the technical concept and scope of the invention. In the described drawings, similar reference numerals refer to similar elements.

[0052] The terms "first," "second," etc., used in this specification may be used to describe various components, but these components shall not be construed as being limited by these terms. The terms are used only to distinguish one component from others. For example, without departing from the scope of the invention, a "first" component may be named a "second" component, and a "second" component may similarly be named a "first" component. The term "and / or" includes a combination of multiple items or any one of multiple items.

[0053] It will be understood that when an element is referred to in this description only as "connected to" or "coupled to" another element rather than "directly connected to" or "directly coupled to" another element, the element may be "directly connected to" or "directly coupled to" the other element, or, in the presence of other elements in between, be "connected to" or "coupled to" the other element. Conversely, it should be understood that when an element is referred to as "directly coupled to" or "directly connected to" another element, there are no intermediate elements.

[0054] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. Expressions used in the singular include plural expressions unless they have a distinct meaning in the context. It will be understood in this specification that terms such as “comprising,” “having,” etc., are intended to indicate the presence of the features, numbers, steps, actions, elements, components, or combinations thereof disclosed in the specification, and are not intended to exclude the possibility that one or more other features, numbers, steps, actions, elements, components, or combinations thereof may be present or added.

[0055] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the drawings, the same components are indicated by the same reference numerals, and repeated descriptions of the same components will be omitted.

[0056] Figure 1This is a block diagram illustrating an apparatus for encoding an image according to an embodiment of the present invention.

[0057] Reference Figure 1 The image encoding device 100 may include a screen partitioning module 110, prediction modules 120 and 125, a transformation module 130, a quantization module 135, a rearrangement module 160, an entropy encoding module 165, an inverse quantization module 140, an inverse transformation module 145, a filter module 150, and a memory 155.

[0058] exist Figure 1 The components shown are illustrated independently to represent different functionalities within the image encoding apparatus and do not imply that each component is constructed as a separate hardware or software unit. In other words, for convenience, each component includes every component listed, and at least two components of each component may be combined to form one component, or a component may be divided into multiple components to perform each function. Embodiments where each component is combined and embodiments where components are separated are also included within the scope of the invention without departing from its spirit.

[0059] Furthermore, some components may not be essential for performing the main functions of the invention, but may be optional components used only to improve its performance. The invention can be implemented by including only the essential components necessary to achieve the essence of the invention, excluding components used to improve performance. A structure that includes only the essential components and excludes optional components used only to improve performance can also be included within the scope of the invention.

[0060] The screen partitioning module 110 can partition the input screen into at least one block. Here, a block can mean a coding unit (CU), a prediction unit (PU), or a transform unit (TU). Partitioning can be performed based on at least one of a quadtree or a binary tree. A quadtree is a method of partitioning a high-level block into four low-level blocks, wherein the width and height of the four low-level blocks are half the size of the high-level block. A binary tree is a method of partitioning a high-level block into two low-level blocks, wherein the width and height of the two low-level blocks are half the size of the high-level block. Using binary tree-based partitioning, the blocks can have square shapes as well as non-square shapes.

[0061] In the embodiments of the present invention, the encoding unit may be used as a unit for performing encoding, or it may be used as a unit for decoding.

[0062] Prediction modules 120 and 125 may include an inter-frame prediction module 120 for performing inter-frame prediction and an intra-frame prediction module 125 for performing intra-frame prediction. It can be determined whether to perform inter-frame or intra-frame prediction for a prediction unit, and it can be determined based on specified information for each prediction method (e.g., intra-frame prediction mode, motion vector, reference frame, etc.). Here, the processing unit performing the prediction may be different from the processing unit that determines the prediction method and specified content. For example, the prediction method, prediction mode, etc., may be determined on a unit-by-unit basis, while prediction may be performed on a unit-by-unit basis.

[0063] The coding device can determine the optimal prediction mode for the coding block by using various schemes, such as rate-distortion optimization (RDO) for the residual block obtained by subtracting the source block from the prediction block. In one example, the RDO can be determined by Equation 1 below.

[0064] [Equation 1]

[0065] J(Φ, λ) = D(Φ) + λR(Φ)

[0066] In Equation 1 above, D represents the degradation due to quantization, R represents the rate of the compressed stream, and J represents the RD cost. Furthermore, Φ represents the encoding mode, and λ represents the Lagrange multiplier. λ can be used as a scaling factor to match the unit of error to the bit quantity. During encoding, the encoding device can determine the mode with the minimum RD cost as the optimal mode for the encoded block. Here, both the bit rate and the error are considered when calculating the RD cost.

[0067] In intra-frame prediction mode, DC mode, which is a non-directional prediction mode (or non-angle prediction mode), can use the average value of the neighboring pixels of the current block. Figure 3 This is a diagram used to explain the intra-frame prediction method using DC mode.

[0068] After the average value of neighboring pixels is filled into the prediction block, filtering can be performed on pixels located at the boundaries of the prediction block. In one example, a weighted sum filter with neighboring reference pixels can be applied to pixels located at the left or top boundary of the prediction block. For example, Equation 2 shows an example of generating prediction pixels for each region using DC mode. In Equation 1, regions R1, R2, and R3 are the outermost (i.e., boundary) regions of the prediction block, and a weighted sum filter can be applied to pixels included in these regions.

[0069] [Equation 2]

[0070]

[0071] R1 region)Pred[0][0]=(R[-1][0]+2*DC value+R[0][-1]+2)>>2

[0072] R2 region)Pred[x][0]=(R[x][-1]+3*DC value+2)>>2,x>0

[0073] R3 region)Pred[0][y]=(R[0][y]+3*DC value+2)>>2,y>0

[0074] R4 region) Pred[x][y] = DC value, x > 0, y > 0

[0075] In Equation 2, Wid represents the horizontal length of the prediction block, and Hei represents the vertical length of the prediction block. x and y represent the coordinates of each predicted pixel when the top-left position of the prediction block is defined as (0, 0). R represents neighboring pixels. For example, when in Figure 3 When pixel s is defined as R[-1][-1], pixels a to i can be represented as R[0][-1] to R[8][-1], and pixels j to r can be represented as R[-1][0] to R[-1][8]. Figure 3 In the example shown, the predicted pixel value Pred can be calculated for each of regions R1 to R4 according to the weighted filtering method shown in Equation 2.

[0076] In non-directional modes, planar mode is a method of generating predicted pixels for the current block by applying linear interpolation based on distance to neighboring pixels of the current block. For example, Figure 4 This is a diagram used to explain intra-frame prediction methods using planar modes.

[0077] For example, suppose in Figure 4 The Pred shown is predicted within an 8×8 coded block. In this case, the pixel e above Pred and the pixel r to the lower left of Pred can be copied to the bottom of Pred, and the vertical prediction value can be obtained by linear interpolation based on distance in the vertical direction. Furthermore, the pixel n to the left of Pred and the pixel i to the upper right of Pred can be copied to the rightmost side of Pred, and the horizontal prediction value can be obtained by linear interpolation based on distance in the horizontal direction. Subsequently, the average of the horizontal and vertical prediction values ​​can be determined as the value of Pred. Equation 3 is a formula representing the process of obtaining the predicted value Pred based on the planar pattern.

[0078] [Equation 3]

[0079]

[0080] In Equation 3, Wid represents the horizontal length of the prediction block, and Hei represents the vertical length of the prediction block. x and y represent the coordinates of each predicted pixel when the top-left position of the prediction block is defined as (0, 0). R represents neighboring pixels. For example, when in Figure 4 When pixel s is defined as R[-1][-1], pixels a to i can be represented as R[0][-1] to R[8][-1], and pixels j to r can be represented as R[-1][0] to R[-1][8].

[0081] Figure 5 This is a diagram used to explain intra-frame prediction methods using directional prediction modes.

[0082] Oriented prediction mode (or angle prediction mode) is a method of generating prediction samples from at least one or more pixels in the neighboring pixels of the current block located in any one of N predetermined directions.

[0083] Orientation prediction patterns can include horizontal orientation patterns and vertical orientation patterns. Here, a horizontal orientation pattern refers to a pattern with greater horizontal directionality than a prediction pattern pointing at a 45-degree angle to the upper left, and a vertical orientation pattern refers to a pattern with greater vertical directionality than a prediction pattern pointing at a 45-degree angle to the upper left. An orientation prediction pattern with a prediction direction pointing at a 45-degree angle to the upper left can be considered either a horizontal or vertical orientation pattern. Figure 5 In the image, horizontal orientation mode and vertical orientation mode are shown.

[0084] Reference Figure 5 There exist directions that do not match the integer pixel portion for each direction. In this case, after applying distance interpolation (such as linear interpolation, DCT-IF, cubic convolution interpolation, etc.) to the distance between the pixel and its neighboring pixels, the pixel value can be filled into the pixel position that matches the direction of the predicted block.

[0085] The residual values ​​(residual block or transform block) between the generated prediction block and the original block can be input into the transform module 130. The residual block is the smallest unit of the transform and quantization process. The partitioning method of the encoded block can be applied to the transform block. In one example, the transform block can be partitioned into four or two local blocks.

[0086] Prediction mode information, motion vector information, etc., used for prediction can be encoded by entropy encoding module 165 along with residual values ​​and sent to the decoding device. When a specific encoding mode is used, it is feasible to send the original block to the decoding device by encoding it instead of generating a prediction block through prediction modules 120 and 125.

[0087] The inter-frame prediction module 120 can predict the prediction unit based on information from at least one of the previous and subsequent frames of the current frame, or in some cases, it can predict the prediction unit based on information from some coded regions in the current frame. The inter-frame prediction module 120 may include a reference frame interpolation module, a motion prediction module, and a motion compensation module.

[0088] The reference image interpolation module can receive reference image information from the memory 155 and can generate pixel information in integer pixels or smaller than integer pixels from the reference image. In the case of luminance pixels, an interpolation filter based on an 8-tap DCT with different filtering coefficients can be used to generate pixel information in integer pixels or smaller than integer pixels in units of 1 / 4 pixels. In the case of chrominance signals, an interpolation filter based on a 4-tap DCT with different filtering coefficients can be used to generate pixel information in integer pixels or smaller than integer pixels in units of 1 / 8 pixels.

[0089] The motion prediction module can perform motion prediction based on a reference frame interpolated by the reference frame interpolation module. Various methods, such as Full-Search-Based Block Matching (FBMA), Three-Step Search (TSS), and New Three-Step Search (NTS), can be used to calculate motion vectors. Motion vectors can have values ​​in units of 1 / 2 or 1 / 4 pixels based on the interpolated pixels. The motion prediction module can predict the current prediction unit by changing the motion prediction method. Various methods, such as skipping, merging, and Improved Motion Vector Prediction (AMVP), can be used as motion prediction methods.

[0090] The encoding device can generate motion information for the current block based on motion estimates or motion information of neighboring blocks. Here, the motion information may include at least one of motion vectors, reference image indexes, and predicted directions.

[0091] The intra-frame prediction module 125 can generate prediction units based on reference pixel information adjacent to the current block, wherein the reference pixel information is pixel information in the current frame. When the neighboring block of the current prediction unit is the block to be inter-frame predicted and therefore the reference pixel is a pixel to be inter-frame predicted, the reference pixel included in the block to be inter-frame predicted can be replaced by the reference pixel information of the neighboring block to be intra-frame predicted. That is, when the reference pixel is unavailable, at least one of the available reference pixels can be used to replace the unavailable reference pixel information.

[0092] Intra-frame prediction modes may include directional prediction modes that use reference pixel information based on the prediction direction and non-directional prediction modes that do not use directional information when performing prediction. The mode used to predict luminance information may differ from the mode used to predict chrominance information. To predict chrominance information, intra-frame prediction mode information used to predict luminance information or predicted luminance signal information may be utilized.

[0093] In intra-frame prediction methods, prediction blocks can be generated after applying an adaptive intra-frame smoothing (AIS) filter to a reference pixel according to the prediction mode. The type of AIS filter applied to the reference pixel can be different. To perform the intra-frame prediction method, the intra-frame prediction mode of the current prediction unit can be predicted from the intra-frame prediction modes of prediction units adjacent to the current prediction unit. During the process of predicting the prediction mode of the current prediction unit using mode information predicted from neighboring prediction units, when the intra-frame prediction mode of the current prediction unit is the same as that of neighboring prediction units, predetermined flag information can be used to send information indicating that the prediction modes of the current prediction unit and neighboring prediction units are the same; and when the prediction mode of the current prediction unit is different from that of neighboring prediction units, entropy coding can be performed to encode the prediction mode information of the current block.

[0094] Furthermore, a residual block including information about residual values ​​can be generated based on the prediction units produced by prediction modules 120 and 125, wherein the residual values ​​are the differences between the prediction unit to which prediction is performed and the original block of that prediction unit. The generated residual block can be input to transformation module 130.

[0095] Transform module 130 can transform residual blocks containing residual data using transform methods such as Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), and Karhunen Loeve Transform (KLT). To simplify the use of transform methods, basis vectors are used to perform matrix operations. Here, various transform methods can be mixed and used in various ways within the matrix operations, depending on the prediction mode in which the prediction block is encoded. For example, when performing intra-frame prediction, depending on the intra-frame prediction mode, Discrete Cosine Transform can be used for the horizontal direction, and Discrete Sine Transform can be used for the vertical direction.

[0096] The quantization module 135 quantizes the values ​​transformed to the frequency domain by the transform module 130. That is, the quantization module 135 quantizes the transform coefficients of the transform block generated from the transform module 130 and generates a quantized transform block with quantized transform coefficients. Here, the quantization method may include dead-zone uniform threshold quantization (DZUTQ) or a quantization weighting matrix, etc. Improved quantization methods that improve upon these methods are also possible. The quantization coefficients may vary depending on the block size or importance of the image. The values ​​calculated by the quantization module 135 can be provided to the inverse quantization module 140 and the rearrangement module 160.

[0097] Transform module unit 130 and / or quantization module 135 may be selectively included in image encoding apparatus 100. That is, image encoding apparatus 100 may perform at least one of transformation and quantization on the residual data of the residual block, or may skip both transformation and quantization to encode the residual block. Even if no transformation or quantization or both transformation and quantization are performed in image encoding apparatus 100, the block provided as input to entropy encoding module 165 is generally referred to as a transform block (or quantized transform block).

[0098] The rearrangement module 160 can rearrange the coefficients of the quantized residual values.

[0099] The rearrangement module 160 can transform coefficients in two-dimensional block form into coefficients in one-dimensional vector form using a coefficient scanning method. For example, the rearrangement module 160 can use a predetermined scanning method to scan from DC coefficients to coefficients in the high-frequency domain in order to transform the coefficients into one-dimensional vector form.

[0100] Entropy coding module 165 can perform entropy coding based on the value calculated by rearrangement module 160. Entropy coding can use various coding methods, such as exponential Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC).

[0101] The entropy coding module 165 can encode various information (e.g., residual coefficient information and block type information of coding units, prediction mode information, partition unit information, prediction unit information), and send unit information, motion vector information, reference frame information, block interpolation information, filtering information, etc., from the rearrangement module 160 and prediction modules 120 and 125. In the entropy coding module 165, the coefficients of the transform block can be encoded into various types of flags, such as non-zero coefficients, coefficients with absolute values ​​greater than 1 or 2, and signs indicating the coefficients, on a local block basis. Coefficients that are not only encoded using flags can be encoded using the absolute value of the difference between the coefficients encoded via the flags and the coefficients of the actual transform block. The reference... Figure 6 The method for encoding the coefficients of the transform block is described in detail.

[0102] The entropy coding module 165 can entropy code the coefficients of the coding units input from the rearrangement module 160.

[0103] The inverse quantization module 140 can inverse quantize the value quantized by the quantization module 135, and the inverse transformation module 145 can inverse transform the value transformed by the transformation module 130.

[0104] Furthermore, the inverse quantization module 140 and the inverse transform module 145 can perform inverse quantization and inverse transform by reversely using the quantization and transform methods used in the quantization module 135 and the transform module 130. Additionally, when the transform module 130 and the quantization module 135 only perform quantization without performing transform, only inverse quantization is performed, while inverse transform may not be performed. When neither transform nor quantization is performed, the inverse quantization module 140 and the inverse transform module 145 may neither perform inverse transform nor inverse quantization, and the inverse quantization module 140 and the inverse transform module 145 may not be included in the image encoding device 100 and may be omitted.

[0105] The residual values ​​generated by the inverse quantization module 140 and the inverse transform module 145 can be combined with the prediction units predicted by the motion estimation module, motion compensation module and intra-frame prediction module included in the prediction modules 120 and 125 to generate a reconstruction block.

[0106] The filter module 150 may include at least one of a deblocking filter, an offset correction unit, and an adaptive loop filter (ALF).

[0107] Deblocking filters remove block distortion caused by boundaries between blocks in the reconstructed image. To determine whether to perform deblocking, the pixels included in several rows or columns within a block can be the basis for deciding whether to apply a deblocking filter to the current block. When a deblocking filter is applied to a block, a strong or weak filter can be applied as needed. Furthermore, horizontal and vertical filtering can be processed in parallel when applying a deblocking filter.

[0108] The offset correction module can correct the offset from the original image on a pixel-by-pixel basis in the image being deblocked. To perform offset correction on a specific image, it is possible to apply the offset by considering the edge information of each pixel, or to divide the image's pixels into a predetermined number of regions, determine the region to be offset, and apply the offset to the determined region.

[0109] Adaptive Loop Filtering (ALF) can be performed based on values ​​obtained by comparing the filtered reconstructed image with the original image. Pixels included in the image can be divided into predetermined groups, the filter to be applied to each group can be determined, and filtering can be performed separately for each group. Information regarding whether ALF is applied and the luminance signal can be transmitted according to the coding unit (CU). The shape and filter coefficients of the filter used for ALF can vary for each block. Moreover, a filter of the same shape (fixed shape) for ALF can be applied regardless of the characteristics of the target block.

[0110] The memory 155 can store the reconstructed blocks or frames calculated by the filter module 150, and the stored reconstructed blocks or frames can be provided to the prediction modules 120 and 125 when performing inter-frame prediction.

[0111] Figure 2 This is a block diagram illustrating an apparatus for decoding an image according to an embodiment of the present invention.

[0112] Reference Figure 2 The image decoding device 200 may include an entropy decoding module 210, a rearrangement module 215, an inverse quantization module 220, an inverse transform module 225, prediction modules 230 and 235, a filter module 240, and a memory 245.

[0113] When an image bitstream is input from an image encoding device, the input bitstream can be decoded according to the reverse processing of the image encoding device.

[0114] The entropy decoding module 210 can perform entropy decoding based on the reverse processing of entropy encoding performed by the entropy encoding module of the image encoding device. For example, corresponding to the method performed by the image encoding device, various methods such as exponential Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC) can be applied. In the entropy decoding module 210, the coefficients of the transform block can be decoded on a local block basis, based on various types of flags such as non-zero coefficients, coefficients with absolute values ​​greater than 1 or 2, and signs indicating coefficients. Coefficients not solely represented by flags can be decoded by combining coefficients represented by flags and coefficients transmitted by signals. (Refer to...) Figure 9 The method for decoding the coefficients of the transform block is described in detail.

[0115] The entropy decoding module 210 can decode information about inter-frame prediction and intra-frame prediction performed by the image coding device.

[0116] The rearrangement module 215 can perform a rearrangement on the bitstream entropy decoded by the entropy decoding module 210 based on a rearrangement method used in the image encoding apparatus. The rearrangement may include reconstructing and rearranging coefficients from one-dimensional vectors into coefficients in two-dimensional blocks. The rearrangement module 215 may receive information related to the coefficient scan performed in the image encoding apparatus and may perform the rearrangement via a method that scans the coefficients in reverse order based on the scan sequence performed in the image encoding apparatus.

[0117] The dequantization module 220 can perform dequantization based on the quantization parameters and block rearrangement coefficients received from the image encoding device.

[0118] The inverse transform module 225 can perform an inverse transform on the inverse-quantized transform coefficients according to a predetermined transform method. Here, the transform method can be determined based on the prediction method (inter-frame / intra-frame prediction), the block size / shape, information about the intra-frame prediction mode, etc.

[0119] Prediction modules 230 and 235 can generate prediction blocks based on information received from entropy decoding module 210 regarding the generation of prediction blocks and information on previously decoded blocks or images received from memory 245.

[0120] Prediction modules 230 and 235 may include a prediction unit determination module, an inter-frame prediction module, and an intra-frame prediction module. The prediction unit determination module may receive various information from the entropy decoding module 210, such as prediction unit information, prediction mode information of the intra-frame prediction method, and motion prediction information of the inter-frame prediction method. It can classify the current coding unit as a prediction unit and determine whether to perform inter-frame prediction or intra-frame prediction on the prediction unit. Using the information received from the image coding apparatus required for inter-frame prediction of the current prediction unit, the inter-frame prediction module 230 may perform inter-frame prediction on the current prediction unit based on information from at least one of the previous and subsequent frames of the current frame including the current prediction unit. Optionally, inter-frame prediction may be performed based on information from some pre-reconstructed regions in the current frame including the current prediction unit.

[0121] To perform inter-frame prediction, it can be determined for the coding unit which of the skip mode, merge mode, AMVP mode, and intra-block copy mode will be used as the motion prediction method for the prediction unit included in the coding unit.

[0122] Intra-prediction module 235 can generate prediction blocks based on pixel information in the current frame. When the prediction unit is a prediction unit to which intra-prediction is performed, intra-prediction can be performed based on the intra-prediction mode information of the prediction unit received from the image coding device. Intra-prediction module 235 may include an adaptive intra-smoothing (AIS) filter, a reference pixel interpolation module, and a DC filter. The AIS filter performs filtering on the reference pixels of the current block, and whether to apply the filter can be determined according to the prediction mode of the current prediction unit. AIS filtering can be performed on the reference pixels of the current block using the prediction mode of the prediction unit received from the image coding device and the AIS filter information. When the prediction mode of the current block is a mode in which AIS filtering is not performed, the AIS filter may not be applied.

[0123] When the prediction mode of the prediction unit is a prediction mode that performs intra-frame prediction based on pixel values ​​obtained by interpolating reference pixels, the reference pixel interpolation module can interpolate the reference pixels to produce reference pixels of integer pixels or less than integer pixels. When the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating reference pixels, the reference pixels may not be interpolated. When the prediction mode of the current block is DC mode, the DC filter can be used to generate the prediction block.

[0124] The reconstructed block or reconstructed screen can be provided to the filter module 240. The filter module 240 may include a deblocking filter, an offset correction module, and an ALF.

[0125] The image encoding device can receive information about whether a deblocking filter should be applied to the corresponding block or frame, and information about which filter, strong or weak, should be applied when the deblocking filter is applied. The image decoding device's deblocking filter can receive information about the deblocking filter from the image encoding module and can perform deblocking filtering on the corresponding block.

[0126] The offset correction module can perform offset correction on the reconstructed image based on the type and offset value information of the offset correction applied to the image during encoding.

[0127] The ALF can be applied to the coding unit based on information received from the image coding device regarding whether to apply the ALF, ALF coefficient information, etc. The ALF information can be provided by including it in a specific parameter set.

[0128] The memory 245 can store reconstructed screens or reconstructed blocks used as reference screens or reference blocks, and can provide reconstructed screens to the output module.

[0129] Figure 6 This is a diagram illustrating a method for encoding the coefficients of a transform block as described in an embodiment of the present invention.

[0130] The coefficients of a transform block can be encoded in a unit of predetermined blocks (hereinafter referred to as local blocks) in an image encoding apparatus. A transform block may include one or more local blocks. A local block may be an N×M block. Here, N and M are natural numbers, and N and M may be equal to or different from each other. That is, a local block may be a square block or a non-square block. The size / shape of the local block may be predefined as fixed (e.g., 4×4) in the image encoding apparatus, or the size / shape of the local block may be variably determined according to the size / shape of the transform block. Optionally, the image encoding apparatus may determine the optimal size / shape of the local block considering encoding efficiency and encode the local block. Information about the size / shape of the encoded local block can be transmitted by signal at the level of at least one of sequence, frame, strip, and block.

[0131] The order in which local blocks included in a transform block are encoded can be determined according to a predetermined scan type (hereinafter referred to as the first scan type) in the image encoding apparatus. Furthermore, the order in which coefficients included in a local block are encoded can be determined according to a predetermined scan type (hereinafter referred to as the second scan type). The first scan type and the second scan type can be the same or different. For the first / second scan type, diagonal scanning, vertical scanning, or horizontal scanning, etc., can be used. However, the invention is not limited to this, and one or more scan types with predetermined angles can be added. The first / second scan type can be determined based on at least one of the following: information related to the encoded block (e.g., maximum / minimum size, partitioning technique, etc.), the size / shape of the transform block, the size / shape of the local block, the prediction mode, information related to intra-frame prediction (e.g., the value, direction, angle, etc. of the intra-frame prediction mode), and information related to inter-frame prediction.

[0132] The image encoding apparatus can encode the position information of coefficients with non-zero values ​​(hereinafter referred to as non-zero coefficients) that first appear in the above encoding order within a transform block. Encoding is performed sequentially, starting from the local block containing the non-zero coefficients. Hereinafter, refer to... Figure 6 This will describe the process of encoding the coefficients of a local block.

[0133] The local block flag for the current local block can be encoded (S600). The local block flag can be encoded on a per-local-block basis. The local block flag can indicate whether there is at least one non-zero coefficient in the current local block. For example, when the local block flag is a first value, the first value can indicate that the current local block includes at least one non-zero coefficient, and when the local block flag is a second value, the second value can indicate that all coefficients in the current local block are 0.

[0134] The local block coefficient flag for the current local block can be encoded (S610). The local block coefficient flag can be encoded on a coefficient-by-coefficient basis. The local block coefficient flag can indicate whether a coefficient is non-zero. For example, when a coefficient is non-zero, the local block coefficient flag can be encoded as a first value, and when a coefficient is zero, the local block coefficient flag can be encoded as a second value. The local block coefficient flag can be selectively encoded based on the local block flag. For example, the current local block can be encoded for each coefficient of the local block only if there is at least one non-zero coefficient in the current local block (i.e., the local block flag is a first value).

[0135] A flag indicating whether the absolute value of the coefficient is greater than 1 (hereinafter referred to as the first flag) can be encoded (S620). The first flag can be selectively encoded based on the value of the local block coefficient flag. For example, when the coefficient is a non-zero coefficient (i.e., the local block coefficient flag is a first value), the first flag can be encoded by checking whether the absolute value of the coefficient is greater than 1. When the absolute value of the coefficient is greater than 1, the first flag is encoded as the first value; when the absolute value of the coefficient is not greater than 1, the first flag can be encoded as the second value.

[0136] A flag indicating whether the absolute value of the coefficient is greater than 2 (hereinafter referred to as the second flag) can be encoded (S630). The second flag can be selectively encoded based on the value of the first flag. For example, when the coefficient is greater than 1 (i.e., the first flag is a first value), the second flag can be encoded by checking whether the absolute value of the coefficient is greater than 2. When the absolute value of the coefficient is greater than 2, the second flag is encoded as the first value; when the absolute value of the coefficient is not greater than 2, the second flag can be encoded as the second value.

[0137] The number of at least one of the first and second flags can range from a minimum of 1 to a maximum of N×M. Optionally, at least one of the first and second flags can be a fixed number predefined in the image coding apparatus (e.g., one, two, or more). The number of the first / second flags can vary depending on the bit depth of the input image, the dynamic range of the raw pixel values ​​in a specific region of the image, the block size / depth, the partitioning technique (e.g., quadtree, binary tree), the transform technique (e.g., DCT, DST), whether the transform is skipped, the quantization parameters, the prediction mode (e.g., intra / inter-frame mode), etc. In addition to encoding the first / second flags, an nth flag indicating whether the absolute value of the coefficient is greater than n can also be encoded. Here, n can represent a natural number greater than 2. The number of the nth flag can be one, two, or more, and can be determined in the same / similar manner as the first / second flags described above.

[0138] Residual coefficients not encoded based on the first / second flag can be encoded in the current local block (S640). Here, encoding can be a process of encoding the coefficient value itself. Residual coefficients can be equal to or greater than two. Residual coefficients can be encoded based on the local block coefficient flag, or at least one of the first or second flags for the residual coefficients. For example, residual coefficients can be encoded as the value obtained by subtracting (local block coefficient flag + first flag + second flag) from the absolute value of the residual coefficients.

[0139] The symbols for the coefficients of a local block can be encoded (S650). Symbols can be encoded in a flag format, unit by coefficient. Symbols can be selectively encoded based on the values ​​of the local block coefficient flags described above. For example, symbols can be encoded only when the coefficients are non-zero (i.e., the local block coefficient flag is the first value).

[0140] As described above, each absolute value of the coefficients of a local block can be encoded by at least one of encoding the local block coefficient flag, encoding the first flag, encoding the second flag, or encoding the remaining coefficients.

[0141] Furthermore, the encoding of the coefficients of the local block described above may also include a process of specifying the range of coefficient values ​​belonging to the local block. Through this process, it can be confirmed whether at least one non-zero coefficient exists in the local block. This process can be implemented by at least one of (A) encoding the maximum value, (B) encoding the first threshold flag, and (C) encoding the second threshold flag, as will be described below. The above process can be implemented by being included in any one of steps S600 to S650 described above, or it can be implemented in a form that replaces at least one of steps S600 to S650. Hereinafter, reference will be made to... Figures 7 to 8 This section describes in detail the process of specifying the range of coefficient values ​​that belong to a local block.

[0142] Figure 7 This is a diagram illustrating a method for encoding the maximum value of the coefficients of a local block as in an embodiment to which the present invention is applied.

[0143] Reference Figure 7 The maximum value among the absolute values ​​of the coefficients in the current local block can be encoded (S700). The range of coefficient values ​​belonging to the current local block can be inferred from the maximum value. For example, when the maximum value is m, the coefficients of the current local block can fall within the range of 0 to m. The maximum value can be selectively encoded based on the value of the local block flag described above. For example, the maximum value can be encoded only if the current local block includes at least one non-zero coefficient (i.e., the local block flag is a first value). When all coefficients of the current local block are 0 (i.e., the local block flag is a second value), the maximum value can be derived as 0.

[0144] Furthermore, the maximum value determines whether the current local block includes at least one non-zero coefficient. For example, when the maximum value is greater than 0, the current local block includes at least one non-zero coefficient; when the maximum value is 0, all coefficients in the current local block can be 0. Therefore, encoding the maximum value can be performed instead of encoding the local block flag in S600.

[0145] Figure 8 This is a diagram illustrating a method for encoding a first threshold flag for a local block as in an embodiment to which the present invention is applied.

[0146] The first threshold flag of this invention can indicate whether all coefficients of a local block are less than a predetermined threshold. The number of thresholds can be N (N>=1), where the range of the thresholds can be {T0,T1,T2,...,T...} N-1 Let} represent the threshold. Here, the 0th threshold T0 represents the minimum value, and the (N-1)th threshold T N-1 Represents the maximum value, {T0,T1,T2,...,T N-1 The threshold values ​​can be ordered in ascending order. In an image coding apparatus, the number of thresholds can be predetermined. The image coding apparatus can determine the optimal number of thresholds considering coding efficiency and encode that number.

[0147] A threshold can be obtained by setting a minimum value of 1 and increasing that minimum value by n (n>=1). In an image coding apparatus, the threshold can be predetermined. The image coding apparatus can determine the optimal threshold by considering coding efficiency and then encode that threshold.

[0148] The threshold range can be determined differently based on the quantization parameter (QP). The QP can be set at at least one of the following levels: sequence, frame, strip, or transform block.

[0149] For example, when QP is greater than a predetermined QP threshold, it can be expected that the distribution of zero coefficients in the transform block will be higher. In this case, the range of the threshold can be determined as {3}, or the encoding process of the first / second threshold flag can be omitted, and the coefficients of the local block can be encoded by steps S600 to S650 described above.

[0150] When QP is less than a predetermined QP threshold, a higher distribution of non-zero coefficients in the transform block can be expected. In this case, the threshold range can be determined to be {3, 5} or {5, 3}.

[0151] In other words, the range of thresholds when QP is small may have a different number and / or size (e.g., a maximum value) than the range of thresholds when QP is large. The number of QP thresholds can be one, two, or more. The QP thresholds can be predetermined in the image coding apparatus. For example, the QP threshold may correspond to the median value of the range of QPs available in the image coding apparatus. Optionally, the image coding apparatus may determine the optimal QP threshold considering coding efficiency and encode that QP threshold.

[0152] Optionally, the threshold range can be determined differently based on the size / shape of the block. Here, a block can refer to an encoding block, a prediction block, a transform block, or a local block. The size can be represented by at least one of width, height, the sum of width and height, or the number of coefficients.

[0153] For example, when the block size is smaller than a predetermined threshold size, the threshold range can be determined to be {3}, or the encoding process of the first / second threshold flag can be omitted, and the coefficients of the local block can be encoded through steps S600 to S650 described above. When the block size is larger than the predetermined threshold size, the threshold range can be determined to be {3, 5} or {5, 3}.

[0154] In other words, the range of thresholds when the block size is small can have a different number and / or size (e.g., a maximum value) than the range of thresholds when the block size is large. The number of threshold sizes can be one, two, or more. The threshold size can be predetermined in the image coding apparatus. For example, the threshold size can be represented by a × b, where a and b are 2, 4, 8, 16, 32, 64, or larger, and a and b can be equal to or different from each other. Optionally, the image coding apparatus can determine the optimal threshold size considering coding efficiency and encode that threshold size.

[0155] Optionally, the range of the threshold can be determined differently based on the range of pixel values. The range of pixel values ​​can be represented by the maximum and / or minimum values ​​of pixels belonging to a predetermined region. Here, the predetermined region can mean at least one of a sequence, a frame, a strip, or a block.

[0156] For example, when the difference between the maximum and minimum values ​​of a pixel value range is less than a predetermined threshold difference, the threshold range is determined to be {3}, or the encoding process of the first / second threshold flag can be omitted, and the coefficients of the local block can be encoded through steps S600 to S650 described above. When the difference is greater than the predetermined threshold difference, the threshold range can be determined to be {3, 5} or {5, 3}.

[0157] In other words, the range of thresholds when the difference is small may have a different number and / or size (e.g., a maximum value) than the range of thresholds when the difference is large. The number of threshold differences may be one, two, or more. The threshold differences may be predetermined in the image coding apparatus. Optionally, the image coding apparatus may determine the optimal threshold difference considering coding efficiency and encode the threshold difference.

[0158] Reference Figure 8 It can determine whether the absolute value of all coefficients in the current local block is less than the current threshold (S800).

[0159] When the absolute value of all coefficients is not less than the current threshold, the first threshold flag can be encoded as "false" (S810). In this case, the current threshold (the i-th threshold) can be updated to the next threshold (the (i+1)-th threshold) (S820), and step S800 described above can be performed based on the updated current threshold. Optionally, when the absolute value of all coefficients is not less than the current threshold, the encoding process of the first threshold flag in step S810 can be omitted, and the current threshold can be updated to the next threshold.

[0160] When the current threshold reaches its maximum value or when the number of thresholds is 1, the current threshold can be updated by adding a predetermined constant to it. The predetermined constant can be an integer greater than or equal to 1. Here, the update can be repeated until the first threshold flag is encoded as "true". Based on the updated current threshold, step S800 can be executed. Optionally, the update process can be terminated when the current threshold reaches its maximum value or when the number of thresholds is 1.

[0161] When the absolute value of all coefficients is less than the current threshold, the first threshold flag can be encoded as "true" (S830).

[0162] As described above, when the first threshold flag for the i-th threshold is "true", this indicates that the absolute value of all coefficients in the local block is less than the i-th threshold. When the first threshold flag for the i-th threshold is "false", this indicates that the absolute value of all coefficients in the local block is greater than or equal to the i-th threshold. Based on the first threshold flag being "true", the range of coefficient values ​​belonging to the local block can be specified. That is, when the first threshold flag for the i-th threshold is "true", the coefficients belonging to the local block can fall within the range of 0 to (i-th threshold - 1).

[0163] Based on the encoded first threshold flag, at least one of the steps S600 to S650 described above can be omitted.

[0164] For example, when the threshold range is {3, 5}, at least one of the first threshold flag for threshold "3" or the first threshold flag for threshold "5" can be encoded. When the first threshold flag for threshold "3" is "true", the absolute value of all coefficients in the local block can fall within the range of 0 to 2. In this case, the coefficients of the local block can be encoded by performing the remaining steps other than at least one of steps S630 or S640 described above, or the coefficients of the local block can be encoded by performing the remaining steps other than at least one of S600, S630, or S640.

[0165] When the first threshold flag for threshold "3" is "false", the first threshold flag for threshold "5" can be encoded. When the first threshold flag for threshold "5" is "false", at least one of the absolute values ​​of the coefficients in the local block can be greater than or equal to 5. In this case, the coefficients of the local block can be encoded by performing steps S600 to S650 described above in the same manner, or the coefficients of the local block can be encoded by performing the remaining steps other than step S600.

[0166] When the first threshold flag for threshold "5" is "true", the absolute values ​​of all coefficients in the local block can fall within the range of 0 to 4. In this case, the coefficients of the local block can be encoded by performing steps S600 to S650 described above in the same manner, or the coefficients of the local block can be encoded by performing the remaining steps other than step S600.

[0167] Furthermore, the first threshold flag of the current local block can be derived based on the first threshold flag of another local block. In this case, the encoding process of the first threshold flag can be omitted, and will refer to... Figure 14 To describe this situation.

[0168] Figure 9 This is a diagram illustrating a method for decoding the coefficients of a transform block as described in an embodiment of the present invention.

[0169] In an image decoding apparatus, the coefficients of a transform block can be decoded in units of predetermined blocks (hereinafter referred to as local blocks). A transform block may include one or more local blocks. A local block may be an N×M block. Here, N and M are natural numbers, and N and M may be equal to or different from each other. That is, a local block may be a square block or a non-square block. In an image decoding apparatus, the size / shape of a local block may be predefined as fixed (e.g., 4×4), may be variably determined according to the size / shape of the transform block, or may be variably determined based on information about the size / shape of the local block transmitted by a signal. Information about the size / shape of the local block may be transmitted by signal at the level of at least one of sequence, frame, strip, or block.

[0170] In an image decoding apparatus, the order in which local blocks belonging to a transform block are decoded can be determined according to a predetermined scan type (hereinafter referred to as the first scan type). Furthermore, the order in which coefficients belonging to local blocks are decoded can be determined according to a predetermined scan type (hereinafter referred to as the second scan type). The first scan type and the second scan type can be the same or different. For the first / second scan type, diagonal scanning, vertical scanning, horizontal scanning, etc., can be used. However, the invention is not limited to this, and one or more scan types with predetermined angles can also be added. The first / second scan type can be determined based on at least one of the following: information related to the coded block (e.g., maximum / minimum size, partitioning technique, etc.), the size / shape of the transform block, the size / shape of the local block, the prediction mode, information related to intra-frame prediction (e.g., the value, direction, angle, etc. of the intra-frame prediction mode), or information related to inter-frame prediction.

[0171] The image decoding device can decode the position information of the coefficients (hereinafter referred to as non-zero coefficients) that first appear in the above decoding order within the transform block. Decoding can be performed sequentially starting from the local block based on the position information. Hereinafter, reference will be made to... Figure 9 Describe the process of decoding the coefficients of a local block.

[0172] The local block flag for the current local block can be decoded (S900). The local block flag can be decoded on a local block-by-local block basis. The local block flag can indicate whether there is at least one non-zero coefficient in the current local block. For example, when the local block flag is a first value, the first value can indicate that the current local block includes at least one non-zero coefficient, and when the local block flag is a second value, the second value can indicate that all coefficients in the current local block are 0.

[0173] The local block coefficient flag for the current local block can be decoded (S910). The local block coefficient flag can be decoded on a coefficient-by-coefficient basis. The local block coefficient flag can indicate whether a coefficient is non-zero. For example, when the local block coefficient flag is a first value, it can indicate that the coefficient is non-zero, and when the local block coefficient flag is a second value, it can indicate that the coefficient is zero. The local block coefficient flag can be selectively decoded based on the local block flag. For example, the current local block can be decoded for each coefficient of the local block only when there is at least one non-zero coefficient in the current local block (i.e., the local block flag is a first value).

[0174] A flag indicating whether the absolute value of a coefficient is greater than 1 (hereinafter referred to as the first flag) can be decoded (S920). The first flag can be selectively decoded based on the value of the local block coefficient flag. For example, when the coefficient is a non-zero coefficient (i.e., the local block coefficient flag is the first value), the first flag can be decoded to check whether the absolute value of the coefficient is greater than 1. When the first flag is the first value, the absolute value of the coefficient is greater than 1; when the first flag is the second value, the absolute value of the coefficient can be 1.

[0175] A flag indicating whether the absolute value of the coefficient is greater than 2 (hereinafter referred to as the second flag) can be decoded (S930). The second flag can be selectively decoded based on the value of the first flag. For example, when the coefficient is greater than 1 (i.e., the first flag is a first value), the second flag can be decoded to check whether the absolute value of the coefficient is greater than 2. When the second flag is a first value, the absolute value of the coefficient is greater than 2; when the second flag is a second value, the absolute value of the coefficient can be 2.

[0176] The number of at least one of the first and second flags can range from a minimum of 1 to a maximum of N×M. Optionally, at least one of the first and second flags can be a fixed number predefined in the image decoding apparatus (e.g., one, two, or more). The number of the first / second flags can vary depending on the bit depth of the input image, the dynamic range of the raw pixel values ​​in a specific region of the image, the block size / depth, the partitioning technique (e.g., quadtree, binary tree), the transform technique (e.g., DCT, DST), whether the transform is skipped, the quantization parameters, the prediction mode (e.g., intra-frame / inter-frame mode), etc. In addition to the first / second flags, an nth flag indicating whether the absolute value of the coefficient is greater than n can be additionally decoded. Here, n can represent a natural number greater than 2. The number of the nth flag can be one, two, or more, and can be determined in the same / similar manner as the first / second flags described above.

[0177] Remaining coefficients in the current local block that have not been decoded based on the first / second flag can be decoded (S940). Here, decoding can be a process of decoding the coefficient values ​​themselves. Remaining coefficients can be equal to or greater than two.

[0178] The symbols for the coefficients of a local block can be decoded (S950). Symbols can be decoded in units of coefficients according to the flag format. Symbols can be selectively decoded based on the values ​​of the local block coefficient flags described above. For example, a symbol can be decoded only when the coefficients are non-zero (i.e., the local block coefficient flag is the first value).

[0179] Furthermore, the decoding of the coefficients of the local block described above may also include a process of specifying the range of coefficient values ​​belonging to the local block. Through this process, it can be confirmed whether at least one non-zero coefficient exists in the local block. This process can be implemented by at least one of (A) decoding the maximum value, (B) decoding the first threshold flag, and (C) decoding the second threshold flag, as will be described below. The above process can be implemented by being included in any one of steps S900 to S950 described above, or it can be implemented in a form that replaces at least one of steps S900 to S950. Hereinafter, reference will be made to... Figures 10 to 11 This section describes in detail the process of specifying the range of coefficient values ​​that belong to a local block.

[0180] Figure 10 This is a diagram illustrating a method for decoding the maximum value of the coefficients of a local block as in an embodiment to which the present invention is applied.

[0181] Reference Figure 10 The information indicating the maximum value among the absolute values ​​of the coefficients in the current local block can be decoded (S1000). The range of coefficient values ​​belonging to the current local block can be inferred from the maximum value of said information. For example, when the maximum value is m, the coefficients of the current local block may fall within the range of 0 to m. The information indicating the maximum value can be selectively decoded based on the value of the local block flag described above. For example, the current local block can be decoded only if the current local block includes at least one non-zero coefficient (i.e., the local block flag is a first value). When all coefficients of the current local block are 0 (i.e., the local block flag is a second value), the information indicating the maximum value can be deduced as 0.

[0182] Furthermore, by considering the maximum value of the information, it can be determined whether the current local block includes at least one non-zero coefficient. For example, when the maximum value is greater than 0, the current local block includes at least one non-zero coefficient; when the maximum value is 0, all coefficients of the current local block can be 0. Therefore, decoding of the maximum value can be performed instead of decoding the local block flag in S900.

[0183] Figure 11 This is a diagram illustrating a method for decoding a first threshold flag for a local block as in an embodiment to which the present invention is applied.

[0184] The first threshold flag of this invention can indicate whether all coefficients of a local block are less than a predetermined threshold. The number of thresholds can be N (N>=1), where the range of the thresholds can be {T0,T1,T2,...,T...} N-1 Let} represent the threshold value. Here, the 0th threshold T0 can represent the minimum value, and the (N-1)th threshold T N-1 It can represent the maximum value, {T0,T1,T2,...,T N-1The threshold values ​​can be arranged in ascending order. The number of thresholds can be predetermined in the image decoding device, or it can be determined based on information about the number of thresholds transmitted by a signal.

[0185] A threshold can be obtained by setting a minimum value of 1 and increasing that minimum value by n (n>=1). The threshold can be set in the image decoding device, or it can be determined based on information about the threshold transmitted via a signal.

[0186] The threshold range can be determined differently based on the quantization parameter (QP). The QP can be set at at least one of the following levels: sequence, frame, strip, and transform block.

[0187] For example, when QP is greater than a predetermined QP threshold, the range of the threshold can be determined as {3}, or the decoding process of the first / second threshold flag can be omitted, and the coefficients of the local block can be decoded through steps S900 to S950 described above.

[0188] Furthermore, when QP is less than a predetermined QP threshold, the threshold range can be determined to be {3, 5} or {5, 3}.

[0189] In other words, the range of thresholds when QP is small may have a different number and / or size (e.g., a maximum value) than the range of thresholds when QP is large. The number of QP thresholds can be one, two, or more. The QP thresholds can be set in the image decoding apparatus. For example, a QP threshold may correspond to an intermediate value of the range of QPs available in the image decoding apparatus. Alternatively, the QP threshold may be determined based on information about the QP threshold transmitted by the image encoding apparatus via a signal.

[0190] Optionally, the threshold range can be determined differently based on the size / shape of the block. Here, a block can refer to an encoding block, a prediction block, a transform block, or a local block. The size can be represented by at least one of width, height, the sum of width and height, and the number of coefficients.

[0191] For example, when the block size is smaller than a predetermined threshold size, the threshold range can be determined to be {3}, or the decoding process of the first / second threshold flag can be omitted, and the coefficients of the local block can be decoded through steps S900 to S950 described above. Furthermore, when the block size is larger than the predetermined threshold size, the threshold range can be determined to be {3, 5} or {5, 3}.

[0192] In other words, the range of thresholds when the block size is small may have a different number and / or size (e.g., a maximum value) than the range of thresholds when the block size is large. The number of threshold sizes can be one, two, or more. The threshold size can be set in the image decoding device. For example, the threshold size can be represented by a × b, where a and b are 2, 4, 8, 16, 32, 64, or larger, and a and b can be equal to or different from each other. Optionally, the threshold size can be determined based on information about the threshold size transmitted by the image encoding device via a signal.

[0193] Optionally, the threshold range can be determined differently based on the range of pixel values. The range of pixel values ​​can be represented by the maximum and / or minimum values ​​of pixels belonging to a predetermined region. Here, the predetermined region can refer to at least one of a sequence, a frame, a strip, and a block.

[0194] For example, when the difference between the maximum and minimum values ​​of the pixel value range is less than a predetermined threshold difference, the threshold range is determined to be {3}, or the decoding process of the first / second threshold flag can be omitted, and the coefficients of the local block can be decoded through steps S900 to S950 described above. Furthermore, when the difference is greater than the predetermined threshold difference, the threshold range can be determined to be {3, 5} or {5, 3}.

[0195] In other words, the range of thresholds when the difference is small may have a different number and / or size (e.g., a maximum value) than the range of thresholds when the difference is large. The number of threshold differences may be one, two, or more. The threshold differences may be set in the image decoding device or determined based on information about the threshold differences transmitted by the image encoding device via signals.

[0196] Reference Figure 11 The first threshold flag for the current threshold can be decoded (S1100).

[0197] The first threshold flag indicates whether the absolute values ​​of all coefficients in a local block are less than the current threshold. For example, when the first threshold flag is "false," it indicates that the absolute values ​​of all coefficients in the local block are greater than or equal to the current threshold. Conversely, when the first threshold flag is "true," it indicates that the absolute values ​​of all coefficients in the local block are less than the current threshold.

[0198] When the first threshold is "false", the current threshold (the i-th threshold) can be updated to the next threshold (the (i+1)-th threshold) (S1110), and the steps S1100 described above can be performed based on the updated current threshold.

[0199] When the current threshold reaches its maximum value or when the number of thresholds is 1, the current threshold can be updated by adding a predetermined constant to it. The predetermined constant can be an integer greater than or equal to 1. This update can be repeated until the first threshold that is "true" is decoded. Optionally, the update process can be terminated when the current threshold reaches its maximum value or when the number of thresholds is 1.

[0200] As in Figure 11 As shown, when the first threshold flag is "true", the decoding of the first threshold flag may no longer be performed.

[0201] As described above, when the first threshold flag for the i-th threshold is "true", this indicates that the absolute value of all coefficients in the local block is less than the i-th threshold. Furthermore, when the first threshold flag for the i-th threshold is "false", this indicates that the absolute value of all coefficients in the local block is greater than or equal to the i-th threshold. Based on the first threshold flag being "true", the range of coefficient values ​​belonging to the local block can be specified. That is, when the first threshold flag for the i-th threshold is "true", the coefficients belonging to the local block can fall within the range of 0 to (i-th threshold - 1).

[0202] Based on the first threshold flag of the decoding, at least one of the steps S900 to S950 described above can be omitted.

[0203] For example, when the threshold range is {3, 5}, at least one of the first threshold flag for threshold "3" and the first threshold flag for threshold "5" can be decoded. When the first threshold flag for threshold "3" is "true", the absolute values ​​of all coefficients in the local block can fall within the range of 0 to 2. In this case, the coefficients of the local block can be decoded by performing the remaining steps other than at least one of steps S930 and S940 described above, or by performing the remaining steps other than at least one of steps S900, S930, and S940 described above.

[0204] When the first threshold flag for threshold "3" is "false", the first threshold flag for threshold "5" can be decoded. When the first threshold flag for threshold "5" is "false", at least one of the absolute values ​​of the coefficients in the local block can be greater than or equal to 5. In this case, the coefficients of the local block can be decoded by performing steps S900 to S950 described above in the same manner, or by performing the remaining steps other than step S900.

[0205] Furthermore, when the first threshold flag for threshold "5" is "true", the absolute values ​​of all coefficients in the local block can fall within the range of 0 to 4. In this case, the coefficients of the local block can be decoded by performing steps S900 to S950 as described above in the same manner, or by performing the remaining steps other than step S900.

[0206] Furthermore, the first threshold flag of the current local block can be derived based on the first threshold flag of another local block. In this case, the encoding process for the first threshold flag can be omitted, and will refer to... Figure 12 To describe this situation.

[0207] Figure 12 This is a diagram illustrating a method for deriving a first / second threshold flag for the current local block as in an embodiment to which the present invention is applied.

[0208] In this embodiment, it is assumed that the transform block 1100 is 8×8, the local block is 4×4, and the block that first appears with a non-zero coefficient is 1220. The local blocks of the transform block are encoded / decoded in the order of 1240, 1220, 1230, and 1210 according to the scan type.

[0209] Within the current local block, a first threshold flag for a specific threshold can be derived based on the first threshold flag of a previous local block. For example, based on a first threshold flag that was "false" in a previous local block, the first threshold flag of the current local block can be derived as "false". Here, it is assumed that {3, 5, 7} is used as the range of thresholds.

[0210] Specifically, since the first local block 1240 in the encoding / decoding sequence has an earlier encoding / decoding order than the local block 1220 to which the first non-zero coefficient belongs, the first threshold flag may not be encoded / decoded. In the second local block 1220 in the encoding / decoding sequence, since the first threshold flag for threshold "3" is "true", only the first threshold flag for threshold "3" can be encoded / decoded. In the third local block 1230 in the encoding / decoding sequence, since the first threshold flag for threshold "3" is "false" and the first threshold flag for threshold "5" is "true", the first threshold flags for thresholds "3" and "5" can be encoded / decoded respectively. In the last local block 1210 in the encoding / decoding sequence, the first threshold flag for threshold "3" is "false" and the first threshold flag for threshold "5" is "false". Here, since the first threshold flag for threshold “3” in the previous local block 1230 is “false”, it can be expected that the current local block 1210 has at least one coefficient with an absolute value equal to or greater than 3, and the first threshold flag for threshold “3” can be derived as “false”.

[0211] In the current local block, a first threshold flag for a specific threshold can be derived based on the first threshold flag of a previous local block. For example, based on a first threshold flag that was "false" in a previous local block, the first threshold flag of the current local block can be derived as "false".

[0212] The following is for reference Figures 13 to 18 The method for determining the local blocks of the transform block will be described in detail.

[0213] An image encoding apparatus can determine local blocks of predetermined size / shape that constitute a transform block, and can encode information about the size / shape of the local blocks. An image decoding apparatus can determine the size / shape of the local blocks based on the encoded information (a first method). Optionally, the size / shape of the local blocks can be determined by predetermined rules in the image encoding / decoding apparatus (a second method). Information indicating whether the size / shape of the local blocks has been determined by one of the first and second methods can be transmitted by signal at least one layer of video, sequence, frame, strip, and block. A block can represent an encoded block, a prediction block, or a transform block.

[0214] The size of a local block within a transform block can be equal to or smaller than the size of the transform block. The shape of the transform block / local block can be square or non-square. The shape of the transform block can be the same as or different from the shape of the local block.

[0215] Information about the shape of the transform block can be encoded. This information may include at least one of the following: whether the transform block uses only squares, only non-squares, or both. Information can be transmitted by signal at at least one layer of video, sequence, picture, strip, and block. A block can represent an coded block, a prediction block, or a transform block. Information about the size of the transform block can be encoded. This information may include at least one of minimum size, maximum size, partition depth, and a maximum / minimum value for the partition depth. Information can be transmitted by signal at at least one layer of video, sequence, picture, strip, and block.

[0216] Information about the shape of a local block can be encoded. This information may include at least one of the following: whether the shape of the local block uses only squares, only non-squares, or both. Information can be transmitted by signaling at at least one layer of video, sequence, picture, strip, and block. A block can represent an coded block, a prediction block, or a transform block. Information about the size of the local block can be encoded. This information may include at least one of minimum size, maximum size, partition depth, and a maximum / minimum value for the partition depth. Information can be transmitted by signaling at at least one layer of video, sequence, picture, strip, and block.

[0217] Figure 13 This is a diagram illustrating a method for determining the size / shape of a local block based on a merging flag, as in an embodiment of the present invention.

[0218] The image encoding device can check and merge the best-shaped local blocks by using RDO from the smallest local block to the largest local block.

[0219] See attached document Figure 13 The RD value can be calculated for transform block 1301, which includes multiple local blocks 1 to 16. A local block can be a local block of a predetermined minimum size in the image coding apparatus. Transform block 1302 is the case where four local blocks 13 to 16 of transform block 1301 are merged into one local block, and the RD value for transform block 1302 can be calculated. Transform block 1303 is the case where four local blocks 9 to 12 of transform block 1301 are merged into one local block, and the RD value for transform block 1303 can be calculated. Transform block 1304 is the case where four local blocks 5 to 8 of transform block 1301 are merged into one local block, and the RD value for transform block 1304 can be calculated.

[0220] The image encoding apparatus can use a quadtree method to calculate the RD cost when merging local blocks in the transform block to reach the maximum size of the local block. Based on the RD cost, the optimal merging is determined, and a merging flag indicating the optimal merging can be encoded. The image decoding apparatus can determine the size / shape of the local blocks in the transform block based on the encoded merging flag.

[0221] For example, it can be assumed that transform block 1304 is the optimal merge, the minimum size of the local block is equal to the size of local block "1" of transform block 1304, and the maximum size of the local block is equal to the size of local block "5" of transform block 1304. In this case, the image encoding device can encode the merge flag "false" indicating that transform block 1301 is not the optimal merge. Furthermore, the merge flag "false" indicating that the state of the four local blocks 10 to 13 of transform block 1304 not being merged is the optimal merge can be encoded, and the merge flag "false" indicating that the state of the four local blocks 6 to 9 of transform block 1304 not being merged is the optimal merge can be encoded. Additionally, the merge flag "true" indicating that the state of the four local blocks 6 to 9 of transform block 1304 being merged into one local block is the optimal merge can be encoded, and the merge flag "false" indicating that the state of the four local blocks 1 to 4 of transform block 1304 not being merged is the optimal merge can be encoded. In other words, the image encoding device can generate a bit stream "00010" by encoding, and the image decoding device can decode the bit stream to determine the merged shape of the transform block 1304.

[0222] Figure 13The merging order of local blocks is not restricted, but local blocks can be merged in different orders. The merging can be performed within a predetermined block size / shape range in the image encoding / decoding device. The shape of the merged local block can be square or non-square. The shape of the merged local block can be determined based on the encoding order or scanning order of the local blocks.

[0223] For example, when the encoding order of local blocks in a transform block is diagonal, merging square shapes can be used. Optionally, when the encoding order of local blocks in a transform block is vertical, merging non-square shapes of vertical rectangles can be used. Optionally, when the encoding order of local blocks in a transform block is horizontal, merging non-square shapes of horizontal rectangles can be used.

[0224] Figure 14 This is a diagram illustrating a method for determining the size / shape of a local block based on partition markers, as in an embodiment of the present invention.

[0225] The image encoding device can check the transform block with the optimal partition shape by using RDO from the largest local block to the smallest local block.

[0226] Reference Figure 14 The RD value for transform block 1401, which includes a local block 1, can be calculated. Local block 1 can be a local block of a predetermined maximum size in the image coding apparatus. Transform block 1302 is the case where the maximum-sized local block is divided into four local blocks 1 to 4, and the RD value for transform block 1302 can be calculated. Transform block 1403 is the case where local block 1 of transform block 1402 is divided into four local blocks 1 to 4, and the RD value for transform block 1403 can be calculated. Transform block 1404 is the case where local block 1 of transform block 1403 is further divided into four local blocks 1 to 4, and the RD value for transform block 1404 can be calculated.

[0227] Image encoding apparatuses can use quadtrees to calculate the RD cost when partitioning local blocks in a transform block to achieve the minimum local block size. Optimal partitioning can be determined based on the RD cost, and partition flags indicating the optimal partitioning can be encoded. Image decoding apparatuses can determine the size / shape of local blocks in a transform block based on the encoded partition flags.

[0228] For example, it can be assumed that transform block 1403 is an optimal partition, the minimum size of the local blocks is equal to the size of local block "1" of transform block 1403, and the maximum size of the local blocks is equal to the size of transform block 1403. In this case, the image encoding device can encode the partition flag "true" indicating that transform block 1401 is not an optimal partition. Furthermore, the partition flag "true" indicating that the state of local block "1" of transform block 1402 being partitioned into four local blocks is an optimal partition can be encoded. The partition flag "false" indicating that the state of the remaining local blocks 2 to 4 of transform block 1402 not being partitioned into four local blocks is an optimal partition can be encoded. That is, the image encoding device can generate a bitstream "11000" through encoding, and the image decoding device can decode the bitstream to determine the partition shape of transform block 1403.

[0229] Figure 14 The partitioning order of local blocks is not restricted, but local blocks can be partitioned in different orders. The partitioning described above can be performed within a predetermined block size / shape range in the image encoding / decoding device. The shape of the partitioned local blocks can be square or non-square. The shape of the partitioned local blocks can be determined based on the encoding order or scanning order of the local blocks.

[0230] For example, when the encoding order of local blocks in a transform block is diagonal, square-shaped partitions can be used. Optionally, when the encoding order of local blocks in a transform block is vertical, vertically elongated non-square-shaped partitions can be used. Optionally, when the encoding order of local blocks in a transform block is horizontal, horizontally elongated non-square-shaped partitions can be used.

[0231] Figure 15 This is a diagram illustrating a method for determining the size / shape of a local block based on partition index information, as in an embodiment of the present invention.

[0232] The image coding device can determine which partition shape of a local block is optimal by using the RDO (Real-Depth Orientation) from when all local blocks of the transform block have the maximum size to when all local blocks have the minimum size.

[0233] Reference Figure 15 Transform block 1501 consists of a local block 1, and the RD cost in this case can be calculated. Local block 1 can be a local block of a predefined maximum size in the image coding device. Transform block 1502 is the case where transform block 1501 is divided into four local blocks 1 to 4, and the RD cost in this case can be calculated. Transform block 1503 is the case where each local block of transform block 1502 is further divided into four local blocks, and the RD cost in this case can be calculated.

[0234] As described above, within the range of local blocks from the largest to the smallest size, the RD cost can be calculated when partitioning the transform block into local blocks of the same size. The optimal partition is determined based on the RD cost, and partition index information indicating the optimal partition can be encoded. The image decoding apparatus can determine the size / shape of the local blocks in the transform block based on the encoded partition index information.

[0235] For example, when transform block 1501 is the optimal partition, the image encoding device can encode "0" as the partition index information; when transform block 1502 is the optimal partition, the image encoding device can encode "1" as the partition index information; and when transform block 1503 is the optimal partition, the image encoding device can encode "2" as the partition index information. The image decoding device can determine the size / shape of local blocks in the transform block based on the encoded partition index information.

[0236] Figure 16 This is a diagram illustrating a method for partitioning transform blocks based on the positions of non-zero coefficients, as in an embodiment of the present invention.

[0237] In a transform block, the size / shape of a local block can be determined based on the position of the first non-zero coefficient that appears in the encoding / decoding order.

[0238] The size / shape of a local block can be determined as the size / shape of the block including the position of the first occurrence of a non-zero coefficient (hereinafter referred to as the first reference block). The first reference block can be the block with the smallest size among the blocks including the position of the first occurrence of a non-zero coefficient and the position of the lower right corner coefficient of the transform block. Here, the first reference block may belong to the range of minimum and maximum sizes of local blocks predetermined in the image encoding / decoding apparatus. The transform block can be partitioned according to the determined size / shape of the local block.

[0239] For example, when transform block 1601 is 16×16 and the position of the first non-zero coefficient in the encoding / decoding order is (12, 12) relative to the top-left corner (0, 0) of the transform block, the size of the local block including coefficient (12, 12) can be determined to be 4×4, and transform block 1601 can be partitioned as follows: Figure 16 The diagram shows 16 4×4 local blocks. Optionally, when transform block 1602 is 16×16 and the position of the first non-zero coefficient in the encoding / decoding order is (8, 8) relative to the top left corner (0, 0) of the transform block, the size of the local block including coefficient (8, 8) can be determined to be 8×8, and transform block 1602 can be partitioned as shown in... Figure 16The diagram shows four 8×8 local blocks. Alternatively, when the transform block 1603 is 16×16 and the position of the first non-zero coefficient in the encoding / decoding order is (8,0) relative to the top left corner (0,0) of the transform block, the size of the local block including the coefficient (8,0) can be determined to be 8×16, and the transform block 1603 can be partitioned into two 8×16 local blocks.

[0240] Optionally, the size / shape of the local block can be determined as the size / shape of the block excluding the position of the first occurrence of a non-zero coefficient (hereinafter referred to as the second reference block). The second reference block can be the block with the largest size among the blocks excluding the position of the first occurrence of a non-zero coefficient and the position of the lower right corner coefficient of the transform block. Here, the second reference block may belong to the range of minimum and maximum sizes of local blocks predetermined in the image encoding / decoding device. The transform block can be partitioned according to the determined size / shape of the local block.

[0241] For example, when transform block 1601 is 16×16 and the position of the first non-zero coefficient in the encoding / decoding order is (12, 11) relative to the top left corner (0, 0) of the transform block, the size of the local block excluding coefficient (12, 11) can be determined to be 4×4, and transform block 1601 can be partitioned into 16 local blocks of size 4×4. Optionally, when transform block 1602 is 16×16 and the position of the first non-zero coefficient in the encoding / decoding order is (7, 13) relative to the top left corner (0, 0) of the transform block, the size of the local block excluding coefficient (7, 13) can be determined to be 8×8, and transform block 1602 can be partitioned into 4 local blocks of size 8×8. Optionally, when the transform block 1603 is 16×16 and the position of the first non-zero coefficient in the encoding / decoding order is (6, 14) relative to the upper left corner (0, 0) of the transform block, the size of the local block excluding the coefficient (6, 14) can be determined to be 8×16, and the transform block 1603 can be partitioned into two local blocks of size 8×16.

[0242] Figure 17 This is a diagram illustrating a method for selectively partitioning a local region of a transform block as in an embodiment to which the present invention is applied.

[0243] In the transform block, the remaining region besides the local regions can be partitioned into local blocks of predetermined size / shape. Here, the local region can be specified based on the position (a, b) of the first occurrence of a non-zero coefficient. For example, the local region may include at least one of a region with an x-coordinate greater than a and a region with a y-coordinate greater than b. The size / shape of the local blocks can be determined in the same / similar manner as in at least one embodiment described above, and its detailed description will be omitted.

[0244] For example, when transform block 1701 is 16×16 and the first non-zero coefficient in the encoding / decoding order is located at (11, 11) relative to the top-left corner (0, 0) of the transform block, the region to the right of the x-coordinate at (11, 11) and the region below the y-coordinate at (11, 11) can be excluded from the local block setting range, and these regions do not need to be encoded / decoded. Here, the remaining region of transform block 1701 can be partitioned into four 6×6 local blocks.

[0245] Optionally, in the transform block, the remaining region besides the local region can be partitioned into local blocks with predetermined sizes / shapes. Here, the local region can be specified based on the position (a, b) of the first occurrence of a non-zero coefficient and the maximum coordinate value (c, d) of the transform block. The maximum coordinate value (c, d) can be the position of the lower right corner coefficient of the transform block. For example, the differences “(ca)” and “db” between the position (a, b) of the first occurrence of a non-zero coefficient and the maximum coordinate value (c, d) of the transform block can be calculated respectively. The position (e, f) of the minimum offset from the difference can be determined relative to the position of the lower right corner coefficient of the transform block. Here, the local region can include at least one of a region with an x-coordinate greater than e and a region with a y-coordinate greater than f. Here, the size / shape of the local block can be determined in the same / similar manner as in at least one embodiment described above, and its detailed description will be omitted.

[0246] For example, when transform block 1701 is 16×16 and the position of the first non-zero coefficient in the encoding / decoding order is (11, 8) relative to the top left corner (0, 0) of the transform block, the difference "4" between the x-coordinate "11" of the corresponding coefficient and the maximum x-coordinate "15" of the transform block can be calculated, and the difference "7" between the y-coordinate "8" of the corresponding coefficient and the maximum y-coordinate "15" of the transform block can be calculated. The position offset by the minimum value "4" of the difference relative to the bottom right corner coefficient of the transform block can be determined as (11, 11). In this case, the region to the right of the x-coordinate (11, 11) and the region below the y-coordinate (11, 11) can be excluded from the local block setting range, and these regions do not need to be encoded / decoded. Here, the remaining region of transform block 1701 can be partitioned into four 6×6 local blocks.

[0247] Optionally, the transform block can be divided into multiple regions based on predetermined boundary lines. There can be one, two, or more boundary lines. The boundary lines have a slope of a predetermined angle, which can fall within the range of 1 to 90 degrees. The boundary lines can include the positions of the first non-zero coefficients appearing in the encoding / decoding order among the coefficients in the transform block. Specifically, relative to the boundary lines, the transform block can be divided into a first region and a second region. Here, the first region can be divided into local blocks of a predetermined size / shape, while the second region may not be divided into local blocks of a predetermined size / shape. That is, coefficients belonging to the first region can be encoded / decoded based on local blocks of a predetermined size / shape, and the encoding / decoding of coefficients belonging to the second region can be skipped. The first region can refer to the region located above, to the left, or to the upper left relative to the boundary line. In this case, the first region may also include the local block region containing the boundary line. The second region can refer to the region located below, to the right, or to the lower right relative to the boundary line.

[0248] For example, assume that transform block 1402 is 16×16, the first non-zero coefficient in the encoding / decoding order is located at (10, 6) relative to the top-left corner (0, 0) of the transform block, and all local blocks of transform block 1402 have been partitioned into 4×4 units. In this case, relative to (10, 6), the last pixel position in the direction 45 degrees to the top-right corner is (15, 1), and the last pixel position in the direction 45 degrees to the bottom-left corner is (1, 15). Relative to the boundary line connecting the two pixel positions, local blocks 1 to 10 in the top-left region can be determined as local blocks to be encoded / decoded, and local blocks in the bottom-right region can be determined as local blocks not to be encoded / decoded. Here, the shape of the local blocks in the top-left region can be determined as quadrilaterals and / or triangles.

[0249] For example, assume that transform block 1403 is 16×16, the first non-zero coefficient appearing in the encoding / decoding order is at position (8, 7) relative to the top-left corner (0, 0) of the transform block, and all local blocks of transform block 1403 have been partitioned into 4×4 units. In this case, relative to (8, 7), the last pixel position in the direction 45 degrees to the top-right corner is (15, 0), and the last pixel position in the direction 45 degrees to the bottom-left corner is (0, 15). Relative to the boundary line connecting the two pixel positions, local blocks 1 to 10 in the top-left region can be determined as local blocks to be encoded / decoded, and local blocks in the bottom-right region can be determined as local blocks not to be encoded / decoded. Here, in the case of local blocks including the boundary line, only coefficients in the top-left region relative to the boundary line can be included in the local block, and coefficients in the bottom-right region can be excluded from the local block. The shape of the local blocks in the upper left region can be determined as quadrilaterals (local blocks 1 to 5, 8) and / or triangles (local blocks 6, 7, 9, 10).

[0250] Figure 18 This is a diagram illustrating a method for partitioning a transform block based on DC / AC component attributes in the frequency domain, as in an embodiment of the present invention.

[0251] In the frequency domain, the partition shape of the transform block can be determined by considering the properties of the DC / AC components included in the transform block. These properties can refer to component location, distribution, concentration, strength, and weakness, and can be determined based on the transform method of the transform block (e.g., DCT, DST, etc.).

[0252] In a transform block, the region where the AC components are least concentrated can be partitioned into local blocks larger than the remaining regions. For example, in a 16×16 transform block 1801, the AC components may be mostly located in local blocks "5" to "13". Here, an 8×8 local block may be assigned only to local block 13 where the AC components are least concentrated, and a 4×4 local block may be assigned to the remaining regions.

[0253] Optionally, in a transform block, the region where the AC components are least concentrated can be partitioned into local blocks smaller than the remaining region. For example, in a 16×16 transform block 1802, the AC components may be mostly located in local blocks "2" to "7". Here, 4×4 local blocks may be assigned only to local blocks "4" to "7" where the AC components are least concentrated, and 8×8 local blocks may be assigned to the remaining region.

[0254] Optionally, in a transform block, the region where the DC component is most concentrated can be partitioned into local blocks smaller than the size of the remaining regions. For example, in a 16×16 transform block 1803, the DC component may be mostly located in local blocks "1" to "4". Here, 4×4 local blocks may be assigned only to local blocks "1" to "4" where the DC component is most concentrated, and 8×8 local blocks may be assigned to the remaining regions.

[0255] Optionally, in a transform block, the region where the DC components are most concentrated can be partitioned into local blocks larger than the remaining regions. For example, in a 16×16 transform block 1804, the DC components may be mostly located in local block "1". Here, an 8×8 local block may be assigned only to local block "1" where the DC components are most concentrated, and a 4×4 local block may be assigned to the remaining regions.

[0256] In the above embodiments, the DC / AC component regions in the transform block are assumed to be based on a four-part transform block. However, the present invention is not limited to this, and the same / similar approach can be applied to transform blocks divided into N (N ≥ 1) or 2 to the power of N equal parts.

[0257] Depending on the quantization parameter (QP) of the transform block, all or some local blocks in the transform block can be selectively encoded / decoded. For example, when the QP of the transform block is greater than a predetermined QP threshold, only one local block in the transform block can be encoded / decoded. Conversely, when the QP of the transform block is less than a predetermined QP threshold, all local blocks in the transform block can be encoded / decoded.

[0258] Here, a local block can be specified by at least one of a predetermined vertical line and a horizontal line. The vertical line may be located at a distance *a* from the left boundary of the transform block in a direction to the left of the left boundary, and the horizontal line may be located at a distance *b* from the upper boundary of the transform block in a direction below the left boundary of the transform block. *a* and *b* are natural numbers and may be the same or different from each other. The local region may be a region located to the left of the vertical line and / or above the horizontal line. The positions of the vertical / horizontal lines may be predetermined in the image encoding / decoding apparatus or may be variably determined considering the size / shape of the transform block. Optionally, the image encoding apparatus may encode information specifying the local region (e.g., information for specifying the positions of the vertical / horizontal lines) and transmit this information by signal, and the image decoding apparatus may specify the local region based on the information transmitted by signal. The boundary of the specified local region may or may not contact the boundary of the local block in the transform block.

[0259] For example, a local region can be a local block within the DC component set or further include N (N ≥ 1) local blocks adjacent to it. Alternatively, a local region can be specified by a vertical line crossing the upper boundary of the transform block at point n and / or a horizontal line crossing the left boundary of the transform block at point m. n and m are natural numbers and can be the same or different from each other.

[0260] The number of QP thresholds can be one, two, or more. The QP thresholds can be predetermined in the image coding apparatus. For example, the QP thresholds may correspond to the median value of the range of QPs available in the image coding / decoding apparatus. Optionally, the image coding apparatus may determine the optimal QP thresholds considering coding efficiency, and the QP thresholds may be encoded.

[0261] Optionally, depending on the size of the transform block, all or some local blocks within the transform block may be selectively encoded / decoded. For example, when the size of the transform block is equal to or greater than a predetermined threshold size, only local regions within the transform block may be encoded / decoded. Conversely, when the size of the transform block is less than the predetermined threshold size, all local blocks within the transform block may be encoded / decoded.

[0262] Here, a local region can be specified by at least one of a predetermined vertical line and a horizontal line. The vertical line may be located at a distance *a* from the left boundary of the transform block in the direction to the left of the left boundary, and the horizontal line may be located at a distance *b* from the upper boundary of the transform block in the direction below the left boundary of the transform block. *a* and *b* are natural numbers and may be the same or different from each other. *a* may fall within the range of 0 to the width of the transform block, and *b* may fall within the range of 0 to the height of the transform block. The local region may be a region located to the left of the vertical line and / or above the horizontal line. The positions of the vertical / horizontal lines may be predetermined in the image encoding / decoding apparatus or may be variably determined considering the size / shape of the transform block. Optionally, the image encoding apparatus may encode information specifying the local region (e.g., information for specifying the positions of the vertical / horizontal lines) and transmit this information by signaling, and the image decoding apparatus may specify the local region based on the information transmitted by signaling. The boundary of the specified local region may or may not contact the boundary of a local block in the transform block.

[0263] For example, a local region can be a local block of a region where DC components are concentrated, or further include N (N ≥ 1) local blocks adjacent to each other. Alternatively, a local region can be specified by a vertical line crossing the upper boundary of the transform block at point n and / or a horizontal line crossing the left boundary of the transform block at point m. n and m are natural numbers and can be the same or different from each other.

[0264] The number of threshold sizes can be one, two, or more. The threshold sizes can be predetermined in the image coding apparatus. For example, the threshold size can be represented by c × d, where c and d are 2, 4, 8, 16, 32, 64, or larger, and c and d can be equal to or different from each other. Optionally, the image coding apparatus can determine the optimal threshold size considering coding efficiency and encode the threshold size.

[0265] Next, we will describe in detail the method of encoding / decoding blocks using multi-frame prediction mode.

[0266] A coded block can be partitioned into at least one prediction block, and each prediction block can be further partitioned into at least one local block (or prediction local block) through an additional partitioning process. Different intra-frame prediction modes can be used to encode each local block. That is, multiple intra-frame prediction modes (or multiple modes) can be used to partition a coded block or prediction block into multiple prediction blocks or multiple local blocks. A prediction block can be partitioned into multiple local blocks according to a predetermined mode. Here, the partition shape of the prediction block can be predetermined and used in the image encoding and decoding apparatus.

[0267] The encoding / decoding of multi-frame prediction mode information will be described in detail below with reference to the accompanying drawings.

[0268] Figure 19 This is a flowchart illustrating the process of determining whether to use a multi-frame intra-prediction mode during encoding. In this embodiment, it is assumed that the current block represents a prediction block encoded in the current intra-prediction mode. In some cases, the current block may be a coded block, a transform block, or a local block generated by partitioning the prediction block.

[0269] First, the encoding device may search for points where the pixel values ​​of neighboring pixels adjacent to the current block change significantly (hereinafter referred to as "inflection points") (S601). An inflection point may refer to a point where the pixel value change between a neighboring pixel and its adjacent neighboring pixels is equal to or greater than a predetermined threshold, or the point where the pixel value change between a neighboring pixel and its adjacent neighboring pixels is the largest.

[0270] Figure 20 This is an example diagram used to explain the search inflection point. For ease of explanation, assume the current block is a 4×4 size prediction block.

[0271] The encoding device can calculate the degree of change in pixel value for each neighboring pixel adjacent to the current block (hereinafter referred to as "inflection point value"). The inflection point value can be calculated based on the amount of change (or difference) between the neighboring pixels adjacent to the current block and the neighboring pixels adjacent to the neighboring pixels, or it can be calculated based on the amount of change (or difference) between the neighboring pixels adjacent to the neighboring pixels.

[0272] Here, neighboring pixels adjacent to the current block include at least one of the following: pixels adjacent to the top boundary of the current block, pixels adjacent to the left boundary of the current block, and pixels adjacent to the corners of the current block (e.g., top left, top right, and bottom left corners). For example, in Figure 20 In the example shown, pixels a through k are illustrated as neighboring pixels of the current block.

[0273] For other examples, the encoding device may calculate the inflection point values ​​of the remaining neighboring pixels among the neighboring pixels adjacent to the current block, excluding pixels adjacent to the corners of the current block (e.g., at least one of the pixels adjacent to the top-left corner, the bottom-left corner, and the top-right corner). Optionally, the encoding device may calculate the inflection point values ​​of the remaining neighboring pixels among the neighboring pixels adjacent to the current block, excluding the rightmost or bottommost pixel. For example, in Figure 20 In the example shown, the encoding device can calculate the inflection point values ​​of the remaining pixels from neighboring pixels a to k, excluding pixels f and k.

[0274] The number or range of neighboring pixels can vary depending on the size and shape of the current prediction block. Therefore, the number or range of neighboring pixels to which the inflection point value is calculated can also vary depending on the size and shape of the current block.

[0275] Equation 4 shows an example of how to calculate the inflection point value.

[0276] [Equation 4]

[0277] If (nearest neighbor pixel == top left neighbor pixel)

[0278] Inflection point value = |(-1 × the pixel below the current neighboring pixel) + (1 × the pixel to the right of the current neighboring pixel)|

[0279] Otherwise, if (nearest neighbor pixel == leftmost neighbor pixel)

[0280] Inflection point value = |(-1 × the upper pixel of the current neighboring pixel) + (1 × the lower velocity of the current neighboring pixel)|

[0281] Otherwise, if (nearest neighbor pixel == top nearest neighbor pixel)

[0282] Inflection point value = |(-1 × the left pixel of the current neighboring pixel) + (1 × the right pixel of the current neighboring pixel)|

[0283] otherwise

[0284] Inflection point value = 0

[0285] In Equation 4, the current neighboring pixel refers to the neighboring pixel whose inflection point value is being calculated among the neighboring pixels adjacent to the current block. The inflection point value of the current neighboring pixel can be calculated as a value obtained by applying an absolute value to the difference between the neighboring neighboring pixels adjacent to the current neighboring pixel. For example, the neighboring pixel adjacent to the upper side of the current block (e.g., in the case of the neighboring pixel) can be calculated based on the difference between the neighboring pixel adjacent to the left and the neighboring pixel adjacent to the right of the neighboring pixel. Figure 20 The inflection point value of pixels "b to e" in the block. The neighboring pixels to the left of the current block (e.g., in the...) can be calculated based on the difference between the neighboring pixels above and below the current pixel. Figure 20 The inflection point value of pixels "g to j" in the block. The neighboring pixels to the top left corner of the current block (e.g., in the...) can be calculated based on the difference between the neighboring pixel to the lower right of the neighboring pixel and the neighboring pixel to the right of the neighboring pixel. Figure 20 The inflection point value of pixel "a" in the image.

[0286] As described above, some neighboring pixels of the current block (e.g., pixels f and k among the neighboring pixels that are not adjacent to the boundary of the current block) can be excluded from the calculation of the inflection point value. The inflection point values ​​of the neighboring pixels of the current block that are not subject to inflection point value calculation can be set to a predetermined value (e.g., 0).

[0287] Once the inflection point values ​​for neighboring pixels have been calculated, the encoding device can select an inflection point based on the calculated inflection point values. For example, the encoding device can set the neighboring pixel with the largest inflection point value, or the neighboring pixel with the largest inflection point value among those with inflection point values ​​equal to or greater than a threshold, as the inflection point. For example, in Figure 20 In the example shown, when the inflection point value of the neighboring pixel b is the largest, the encoding device can... Figure 20 The point “O” shown in the figure is set as the inflection point.

[0288] When multiple neighboring pixels have the same maximum inflection point value, the encoding device can recalculate the inflection point value for the multiple neighboring pixels, or select one of these neighboring pixels as the inflection point based on a predetermined priority.

[0289] For example, when multiple neighboring pixels have the same maximum inflection point value, the encoding device can recalculate the inflection point value of these pixels by adjusting the number or position of the neighboring pixels used to recalculate the inflection point value. For example, the encoding device can increase the number of neighboring pixels adjacent to the neighboring pixel in both directions (or unidirectionally) by 1 when recalculating the inflection point value of a neighboring pixel. Therefore, when the neighboring pixel with the maximum inflection point value is a neighboring pixel adjacent to the upper side of the current block, the inflection point value of the neighboring pixel can be recalculated using two neighboring pixels adjacent to the left side of the neighboring pixel and two neighboring pixels adjacent to the right side of the neighboring pixel.

[0290] As another example, when multiple neighboring pixels have the same maximum inflection point value, the encoding device may select the neighboring pixel closest to the pixel at a specific position as the inflection point. For instance, the encoding device may give higher priority to the inflection point that is closer to the top-left neighboring pixel of the current block.

[0291] As another example, when there are multiple neighboring pixels with the same maximum inflection point value, the encoding device can recalculate the inflection point value and determine the inflection point based on the recalculated inflection point value. When there are multiple neighboring pixels with the maximum inflection point value even after recalculating the inflection point value, the encoding device can select the inflection point based on priority.

[0292] In the above example, it was described that inflection points can be determined based on neighboring pixels adjacent to the current block and neighboring pixels adjacent to neighboring pixels. Besides the example described above, inflection points can also be determined based on the coding parameters of neighboring blocks adjacent to the current block. Here, coding parameters are parameters used to encode neighboring blocks and may include prediction-related information, such as intra-prediction modes, motion vectors, etc. For example, the boundary of a neighboring block can be set as an inflection point if the difference in intra-prediction modes between neighboring blocks adjacent to the current block is equal to or greater than a predetermined threshold, or if the difference in motion vectors between neighboring blocks adjacent to the current block is equal to or greater than a predetermined threshold. Accordingly, when the difference in coding parameters between neighboring blocks is large, it is predictable that a sharp change occurs at the boundary of the neighboring block, thereby improving encoding / decoding efficiency by using the boundary of the neighboring block as an inflection point. The decoding device can also use the coding parameters of neighboring blocks to determine inflection points.

[0293] When the inflection point is determined, the encoding device can determine the partition shape of the current block based on the inflection point (S602).

[0294] Figure 21 This is a diagram showing the partition shape based on the shape of the current block.

[0295] exist Figure 21In the example shown, the current block can be partitioned into two or more local blocks based on at least one reference line. Here, the local blocks created by partitioning the current block can have shapes other than squares, such as triangles or trapezoids. Figure 21 In the text, reference numeral 2101 is an example of 18 partition lines (or 18 partition shapes) existing when the current block is a square. Figure 21 Reference numerals 2102 and 2103 indicate that there are 20 partition lines (or 20 partition shapes) when the current block is not square. However, the partition shape of the current block is not limited to the illustrative example. In addition to the illustration shown, various shapes of diagonal partitions or bisections can be applied.

[0296] The encoding device can select a partition line with a defined inflection point as the starting point or a partition line with a starting point closest to the defined inflection point, and can partition the current block according to the selected partition line. For example, when the position of the inflection point does not match the starting point to the left or top of the partition line, the encoding device can change the position of the inflection point to the starting point of the closest partition line.

[0297] The encoding device can then calculate the slope information of the pixels adjacent to the inflection point in order to determine what shape to divide the current block into from the inflection point.

[0298] For example, Figure 22 This is an example diagram used to explain the calculation of slope information for neighboring pixels. For ease of explanation, assume the current block size is 4×4 and the inflection point is b1.

[0299] The encoding device can use N pixels selected based on inflection points to calculate slope information. For example, as in Figure 22 In the example shown, the encoding device can be configured to include a 3×3 block with inflection point b1 and calculate slope information based on the pixels included in the configured 3×3 block.

[0300] When the inflection point of the current block is a neighboring pixel located above the current block, the inflection point can be included in the bottom row of the 3×3 block. Furthermore, when the inflection point of the current block is a neighboring pixel located to the left of the current block, the inflection point can be included in the rightmost column of the 3×3 block. In other words, when the inflection point of the current block is at the top, the 3×3 block can be configured primarily by the pixels neighboring to the inflection point and the pixels above the inflection point. Similarly, when the inflection point of the current block is on the left, the 3×3 block can be configured primarily by the pixels neighboring to the inflection point and the pixels to the left of the inflection point.

[0301] The encoding device can use neighboring pixels adjacent to the inflection point to calculate the horizontal and vertical slopes, and can calculate slope information for the current block based on the horizontal and vertical slopes. For example, Equation 5 illustrates a series of procedures for calculating slope information.

[0302] Equation 5

[0303] f x (x,y)=I(x+1,y)-I(x-1,y)

[0304] f y (x,y)=I(x,y+1)=I(x,y-1)

[0305]

[0306] In equation 5, f x (x,y) represents the degree of inclination in the horizontal direction, f y (x, y) represents the degree of inclination in the vertical direction. The slope of the current block can be obtained by applying the arctangent of the slope values ​​in the vertical and horizontal directions. The outline direction of the current block can be determined by the slope value of the current block.

[0307] In the example above, the slope value can be calculated based on neighboring pixels adjacent to the current block and neighboring pixels adjacent to neighboring pixels. As another example, the slope of the current block can be determined based on the coding parameters of neighboring blocks adjacent to the current block. For example, the slope value can be determined based on the intra-prediction mode (or intra-prediction mode angle) of the neighboring block adjacent to the inflection point, the motion vector of the neighboring block adjacent to the inflection point, etc. When the inflection point is located at the boundary between neighboring blocks, the coding device can encode index information indicating which neighboring block among the neighboring blocks located at the inflection point is used to determine the slope value. Here, the index information can be information indicating one of a plurality of neighboring blocks.

[0308] When the slope of the current block is calculated, the encoding device can determine the optimal partition shape for the current block by considering the position and degree of inflection of the inflection point. Here, the optimal partition shape can be a partition shape using a partition line whose inflection angle is equal to the inflection angle of the current block, or it can be a partition shape using a partition line with an inflection angle most similar to the inflection angle of the current block (i.e., the partition line whose inflection angle differs least from the inflection angle of the current block). When multiple partition shapes with an inflection angle most similar to the inflection angle of the current block exist, the encoding device can select the partition shape with the higher priority according to a predetermined priority.

[0309] The encoding device can encode information used to determine the partition shape of the current block and send the encoded information to the decoding device via a bitstream. The information used to determine the partition shape of the current block can be encoded via predictive block units or upper-layer headers. Here, upper-layer headers can refer to coding block layers, stripe layers, picture layers, sequence layers, video layers, etc.

[0310] Information used to determine the partition shape of the current block may include at least one of an index identifying the partition shape of the current block, the location of inflection points, and slope information. For example, the location of the inflection points may be encoded by a predictive block cell or an upper-level header and sent to the decoding device.

[0311] As another example, the decoding device can determine the partition shape of the current block in the same way as the encoding device. For instance, the decoding device can deduce inflection points based on the amount of pixel value change between neighboring blocks, use the N pixels around the derived inflection points to calculate slope information, and then determine one of the predetermined partitioning patterns as the partition shape of the current block.

[0312] Here, the encoding device may encode information about whether to use a multi-intra-prediction mode (i.e., whether to perform intra-prediction by partitioning a current block into multiple local blocks) and transmit the encoded information via a bitstream. The decoding device may encode this information and partition the current block into multiple local blocks only when the information indicates that intra-prediction will be performed by partitioning the current block into multiple local blocks.

[0313] When the partition shape for the current block is determined, the encoding device can determine the intra-prediction mode for each local block included in the current block (S603). Here, the encoding device can assign different intra-prediction modes to each local block.

[0314] Based on the rate-distortion cost (RD cost) calculated by performing rate-distortion optimization (RDO) for each intra-prediction mode of the local block, the coding device can determine the intra-prediction mode to be included in the current block. For example, the coding device can determine the intra-prediction mode with the minimum RD cost value to be included in each local block of the current block as the intra-prediction mode for each local block.

[0315] During RDO execution, the encoding device can calculate the RD value of all available intra-prediction modes, or it can calculate the RD value of a subset of multiple intra-prediction modes. The subset of intra-prediction modes can be set and used in both the image encoding and decoding devices. Optionally, the subset of intra-prediction modes may include only intra-prediction modes that use only reference pixels adjacent to the current block.

[0316] As another example, a subset of intra-prediction modes among multiple intra-prediction modes may include at least one of the intra-prediction modes of neighboring blocks adjacent to the current block and predetermined additional intra-prediction modes. For instance, when the intra-prediction mode of a neighboring block is a non-directional mode, RDO can be performed from candidates such as all non-directional modes, some directional prediction modes with high selection frequencies (e.g., vertical prediction mode, horizontal prediction mode, etc.). Alternatively, when the intra-prediction mode of a neighboring block is a directional mode, RDO can be performed from candidates such as all non-directional modes, the intra-prediction mode of a neighboring block, and intra-prediction modes with a direction similar to the intra-prediction mode of a neighboring block (e.g., intra-prediction modes whose difference from the intra-prediction mode of a neighboring block is equal to or less than a threshold).

[0317] The encoding device can perform intra-prediction on the current block using a selected intra-prediction mode. Here, when the current block is divided into local blocks along the diagonal, the partition lines may not match the pixel boundaries, and overlapping pixels (i.e., overlapping regions) may occur between local blocks.

[0318] For example, Figure 23 This is a diagram illustrating an example of overlapping regions generated based on the partition shape of the current block.

[0319] In and in Figure 23 In the case of the blocks corresponding to reference numerals 2301 and 2302 shown, no overlapping area occurs between local blocks because the partition lines match the pixel boundaries. However, in the case of the blocks corresponding to reference numerals 2303 to 2305, since there are parts where the partition lines do not match the pixel boundaries (i.e., the parts where the partition lines pass through pixels), overlapping pixels (i.e., overlapping areas) between local blocks can be considered.

[0320] Similarly, when there are overlapping regions between local blocks, the predicted values ​​of pixels included in the overlapping regions can be obtained by averaging the predicted values ​​generated from the intra-prediction results of the local blocks including the overlapping regions, or by linear interpolation of the predicted values. For example, assuming the current block is divided into a first local block and a second local block, the predicted values ​​of pixels jointly included in the first and second local blocks can be determined as the average of a first predicted value calculated using the intra-prediction mode of the first local block and a second predicted value calculated using the intra-prediction mode of the second local block, or the value of linear interpolation of the first and second predicted values.

[0321] As another example, an intra-prediction mode of any one of the local blocks that includes the overlapping region can be used to generate predicted values ​​for pixels included in the overlapping region. Here, which block of the local blocks to use can be determined based on a predetermined priority between local blocks, a predetermined priority between intra-prediction modes, the position of the pixel included in the overlapping region, etc. For example, an intra-prediction mode with a higher predetermined priority among the intra-prediction modes of the local blocks that include the overlapping region can be used to generate predicted values ​​for pixels included in the overlapping region.

[0322] Next, the encoding of the intra-prediction mode for each local block in the coding apparatus will be described in detail.

[0323] Figure 24 This is a flowchart illustrating the process of encoding the intra prediction mode of a local block. For ease of explanation, the intra prediction modes of a local block are classified into primary prediction modes and secondary prediction modes. Here, the primary prediction mode may indicate the intra prediction mode of one of the local blocks, and the secondary prediction mode may indicate the intra prediction mode of another local block. Optionally, the primary and secondary prediction modes can be determined based on the priority of the intra prediction modes, the size of the local block, the position of the local block, etc. The secondary prediction mode may be determined as an intra prediction mode different from the primary prediction mode.

[0324] First, the encoding device can encode the operation information for multi-frame prediction (S1101). Here, the multi-mode operation information indicates whether it is optimal to use a single intra-frame prediction mode or to use multiple intra-frame prediction modes for the current block.

[0325] When the multi-frame prediction mode is true (S1102), the coding device can encode the intra-frame prediction mode determined as the primary prediction mode (S1103), and then encode the intra-frame prediction modes determined as secondary prediction modes (S1104). The primary prediction mode can be encoded using an encoding process with MPM (most probable mode) candidates, or it can be encoded without using MPM candidates. The secondary prediction mode can be encoded using an encoding process with MPM candidates, or it can be encoded without using MPM candidates.

[0326] The encoding device can encode the difference between the primary prediction mode and the secondary prediction mode. For example, the encoding device can use MPM candidates to encode the primary prediction mode and can also encode the difference between the primary and secondary prediction modes. In this case, the decoding device can use the MPM candidates to derive the primary prediction mode and obtain the secondary prediction mode through the difference between the primary and secondary prediction modes.

[0327] When the multi-frame prediction mode is false (S1102), the coding device can encode the single intra-frame prediction mode of the current block (S1105). For example, the coding device can use the MPM candidate to encode the intra-frame prediction mode of the current block.

[0328] Reference Figure 25 The encoding of the intra-prediction mode for the current block using MPM candidates will be described in detail.

[0329] Figure 25 This is a flowchart illustrating a method for encoding intra-prediction modes using MPM candidates. In this embodiment, the intra-prediction mode of the current block may refer to one of a primary prediction mode, a secondary prediction mode, and a single prediction mode.

[0330] Reference Figure 25 First, the coding device can determine the MPM candidate for the current block (S1210). The coding device can determine the MPM candidate for the current block based on the intra-prediction mode of the neighboring blocks adjacent to the current block.

[0331] For example, Figure 26 This is a diagram illustrating the determination of MPM candidates for the current block. Figure 26 In this context, L can indicate the intra prediction mode with the highest usage frequency among the intra prediction modes of the left-adjacent block to the left of the current block, and A can indicate the intra prediction mode with the highest usage frequency among the intra prediction modes of the upper-adjacent block to the upper of the current block. Optionally, L can indicate the intra prediction mode of the left-adjacent block at a specified position, and A can indicate the intra prediction mode of the upper-adjacent block at a specified position.

[0332] exist Figure 26 The diagram illustrates three MPM candidates for the current block, which may include at least one of L, A, intra-prediction modes with a similar orientation to L (i.e., L-1 and L+1), a non-directional mode (planar, DC), and a predetermined directional mode (vertical mode). However, the invention is not limited to the illustrative example, and MPM candidates for the current block may be generated by methods different from those shown. For example, the current block may include more than three MPM candidates.

[0333] When neighboring blocks adjacent to the current block are encoded in multi-frame intra-prediction modes, the MPM candidate for the current block can be determined by considering all multi-frame intra-prediction modes of the neighboring blocks, or by considering only one of the multi-frame intra-prediction modes of the neighboring blocks. For example, the dominant prediction mode among the multi-frame intra-prediction modes of the neighboring blocks can be considered to determine the MPM candidate for the current block.

[0334] Optionally, the primary prediction modes of neighboring blocks can be used to determine the MPM candidates for the primary prediction mode of the current block, and the secondary prediction modes of neighboring blocks can be used to determine the MPM candidates for the secondary prediction mode of the current block.

[0335] The encoding device can determine whether there is an MPM candidate with the same intra-prediction mode as the current block, and can encode operation information according to the determination result (S1202). Here, the operation information can be a 1-bit flag (e.g., an MPM flag). For example, when there is an MPM candidate with the same intra-prediction mode as the current block, the encoding device can encode the MPM flag as "true". When there is no MPM candidate with the same intra-prediction mode as the current block, the encoding device can encode the MPM flag as "false".

[0336] When there is an MPM candidate with the same intra-prediction mode as the current block (S1203), the encoding device can encode the index information of the MPM candidate with the same intra-prediction mode as the current block (S1204).

[0337] When no MPM candidate is available that matches the intra-prediction mode of the current block (S1203), the encoding device may encode the remaining modes that indicate the optimal intra-prediction mode for the current block among the remaining intra-prediction modes other than the MPM candidates (S1205). Specifically, the encoding device may encode the remaining modes by allocating as many bits as the number of remaining intra-prediction modes obtained by subtracting the number of MPM candidates from the total number of intra-prediction modes (or the intra-prediction modes available for the current block).

[0338] The secondary prediction mode is set to a different value than the primary prediction mode, and the remaining modes for the secondary prediction mode can be encoded based on the remaining intra-prediction modes other than the MPM candidate and the primary prediction mode.

[0339] Next, a method for decoding the optimal intra-frame prediction mode in a decoding device will be described.

[0340] Figure 27 This is a flowchart illustrating a method for decoding the intra-prediction mode of the current block.

[0341] Reference Figure 27 First, the decoding device can decode the multi-mode operation information from the bitstream (S1401). Here, the multi-mode operation information can indicate whether the current block is encoded using a multi-frame prediction mode.

[0342] When it is determined that the current block is encoded using a multi-frame intra-prediction mode (S1402), the decoding device can partition the current block into multiple local blocks. Here, the decoding device can partition the current block based on the block partitioning mode information sent by the encoding device using a signal, or it can calculate the inflection point and slope, select the partitioning mode corresponding to the calculated inflection point and slope, and then partition the current block.

[0343] When the current block is partitioned into multiple local blocks, the decoding device can decode the primary prediction mode of the current block (S1403), and then decode the secondary prediction mode (S1404). As described in the encoding process, the primary or secondary prediction mode can be decoded using MPM candidates, or it can be decoded without using MPM candidates. Optionally, after decoding the primary prediction mode, the secondary prediction mode can be obtained based on the difference between the primary and secondary prediction modes.

[0344] When it is determined that the current block is not encoded using multi-frame prediction mode, the decoding device may decode the current block using single-frame prediction mode (S1405). Here, MPM candidates can be used to decode the current block using single-frame prediction mode.

[0345] Figure 28 This is a flowchart illustrating a method for decoding intra-prediction modes using MPM candidates. In this embodiment, the intra-prediction mode of the current block can refer to one of the primary prediction mode, secondary prediction mode, and single prediction mode.

[0346] First, the decoding device can determine the MPM candidates for the current block (S1501). The decoding device can determine the MPM candidates for the current block based on the intra-prediction modes of the neighboring blocks adjacent to the current block. (See reference...) Figure 26 The generation of MPM candidates for the current block is described in detail, but the specific details will be omitted.

[0347] Subsequently, the decoding device may decode information indicating whether there is an MPM candidate with the same intra-prediction mode as the current block (S1502). This information may be a 1-bit flag, but is not limited to this.

[0348] When it is determined that there is an MPM candidate with the same intra-prediction mode as the current block (S1503), the decoding device can decode the information specifying the MPM candidate with the same intra-prediction mode as the current block (i.e., the MPM index information) (S1504). In this case, the intra-prediction mode of the current block can be determined as the intra-prediction mode specified by the MPM index information.

[0349] Furthermore, when it is determined that there are no MPM candidates with the same intra-prediction mode as the current block (S1503), the decoding device can decode the remaining mode information (S1505) and determine the intra-prediction mode of the current block based on the decoded remaining mode information. Here, the remaining mode information may be encoded by excluding MPM candidates from the intra-prediction modes available for the current block.

[0350] Although the exemplary methods of this disclosure are shown through a series of steps for clarity of explanation, they are not intended to limit the order in which the steps are performed, and each step may be performed simultaneously or in a different order if necessary. To implement the methods according to this disclosure, it is possible to additionally include other steps in the illustrative steps, exclude some steps and include the remaining steps, or exclude some steps and include other steps.

[0351] The various embodiments disclosed herein are not intended to describe all possible combinations exhaustively, but are illustrative representations of aspects of this disclosure, and the features described in the various embodiments may be applied independently or in combination of two or more.

[0352] Furthermore, the various embodiments of this disclosure can be implemented by hardware, firmware, software, or a combination thereof. Hardware implementations can be performed by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, etc.

[0353] The scope of this disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, instructions, etc.) and non-transitory computer-readable media executable on a device or computer, wherein the software or machine-executable instructions perform operations of methods according to various embodiments on the device or computer, and such software or instructions are stored on the non-transitory computer-readable medium.

[0354] Industrial applications

[0355] This invention can be used to encode / decode images.

Claims

1. A method for decoding video signals using a decoding device, comprising: A candidate list of the current block is generated based on the neighboring blocks adjacent to the current block, the neighboring blocks including the upper neighboring block and the left neighboring block; Using the decoding device, a reference line for partitioning the current block is determined, the reference line being one of a plurality of candidate reference lines, each of the plurality of candidate reference lines being distinguished based on at least one of a predetermined slope or a predetermined position; Using the decoding device, the current block is partitioned into two or more sub-blocks based on the determined reference line; Using the decoding device, the current block is predicted based on the candidate list and the prediction information of each sub-block in the sub-block of the partition, so as to obtain the predicted block of the current block; as well as Using the decoding device, the current block is reconstructed based on the predicted block and the residual block of the current block obtained from the bitstream. Wherein, at least one of the sub-blocks is triangular in shape. The prediction information includes first prediction information for predicting the first sub-block within the sub-block and second prediction information for predicting the second sub-block within the sub-block. Specifically, the first prediction information for predicting the first sub-block is determined based on index information, wherein the index information specifies one of multiple candidates in the candidate list and is transmitted via a bitstream signal. Specifically, the second prediction information for predicting the second sub-block is determined based on index information, wherein the index information specifies one of multiple candidates in the candidate list and is transmitted via a bitstream signal. The number of the plurality of candidates is greater than 3.

2. The method according to claim 1, wherein, The reference line is determined based on two positions on the side of the current block, one of which is located on a different side from the other of the two positions.

3. The method according to claim 2, wherein, The reference line includes the position of the upper right sample of the current block and the position of the lower left sample of the current block.

4. The method according to claim 1, wherein, The reference line is determined based on first information used to determine the partition type of the current block.

5. The method according to claim 4, wherein, The first information includes information about at least one of the slope or the position of the reference line.

6. The method according to claim 1, wherein, The current block includes the first sub-block and the second sub-block, and Specifically, the predicted value of the boundary region between the first sub-block and the second sub-block is obtained by interpolating the first predicted value of the first sub-block and the second predicted value of the second sub-block.

7. A method for encoding video signals using an encoding device, comprising: A candidate list of the current block is generated based on the neighboring blocks adjacent to the current block, the neighboring blocks including the upper neighboring block and the left neighboring block; Using the encoding device, a reference line for partitioning the current block is determined, the reference line being one of a plurality of candidate reference lines, each of the plurality of candidate reference lines being distinguished based on at least one of a predetermined slope or a predetermined position; Using the encoding device, the current block is divided into two or more sub-blocks based on the reference line to obtain the predicted block of the current block; Using the encoding device, the current block is predicted based on the candidate list and prediction information for each sub-block in the partition; as well as Using the encoding device, the residual block of the current block is encoded into a bitstream, the residual block being obtained based on the predicted block and the original block of the current block. Wherein, at least one of the sub-blocks is triangular in shape. The prediction information includes first prediction information for predicting the first sub-block within the sub-block and second prediction information for predicting the second sub-block within the sub-block. Specifically, the first prediction information for predicting the first sub-block is determined based on index information, whereby the index information specifies one of multiple candidates in the candidate list and is encoded as a bitstream. Specifically, the second prediction information for predicting the second sub-block is determined based on index information, whereby the index information specifies one of multiple candidates in the candidate list and is encoded as a bitstream. The number of the plurality of candidates is greater than 3.

8. The method according to claim 7, wherein, The reference line is determined based on two positions on the side of the current block, one of which is located on a different side from the other of the two positions.

9. The method according to claim 8, wherein, The reference line includes the position of the upper right sample of the current block and the position of the lower left sample of the current block.

10. The method according to claim 7, wherein, The first information used to determine the partition type of the current block is encoded into a bit stream based on a defined reference line.

11. The method according to claim 10, wherein, The first information includes information about at least one of the slope or the position of the reference line.

12. The method according to claim 7, wherein, The current block includes the first sub-block and the second sub-block, and Specifically, the predicted value of the boundary region between the first sub-block and the second sub-block is obtained by interpolating the first predicted value of the first sub-block and the second predicted value of the second sub-block.

13. A method for transmitting a bit stream generated by an encoding method, the encoding method comprising: A candidate list of the current block is generated based on the neighboring blocks adjacent to the current block, the neighboring blocks including the upper neighboring block and the left neighboring block; Using an encoding device, a reference line for partitioning the current block is determined, the reference line being one of a plurality of candidate reference lines, each of the plurality of candidate reference lines being distinguished based on at least one of a predetermined slope or a predetermined position; Using the encoding device, the current block is partitioned into two or more sub-blocks based on the reference line; Using the encoding device, the current block is predicted based on the candidate list and the prediction information of each sub-block in the sub-block of the partition, so as to obtain the predicted block of the current block; as well as Using the encoding device, the residual block of the current block is encoded into a bitstream, the residual block being obtained based on the predicted block and the original block of the current block. Wherein, at least one of the sub-blocks is triangular in shape. The prediction information includes first prediction information for predicting the first sub-block within the sub-block and second prediction information for predicting the second sub-block within the sub-block. Specifically, the first prediction information for predicting the first sub-block is determined based on index information, whereby the index information specifies one of multiple candidates in the candidate list and is encoded as a bitstream. Specifically, the second prediction information for predicting the second sub-block is determined based on index information, whereby the index information specifies one of multiple candidates in the candidate list and is encoded as a bitstream. The number of the plurality of candidates is greater than 3.

14. The transmission method according to claim 13, wherein, The reference line is determined based on two positions on the side of the current block, one of which is located on a different side from the other of the two positions.

15. The sending method according to claim 14, wherein, The reference line includes the position of the upper right sample of the current block and the position of the lower left sample of the current block.

16. The sending method according to claim 13, wherein, First information for determining the partition type of the current block is encoded into a bit stream based on a defined reference line, the first information including information about at least one of the slope or position of the reference line.

17. The sending method according to claim 13, wherein, The current block includes the first sub-block and the second sub-block, and Specifically, the predicted value of the boundary region between the first sub-block and the second sub-block is obtained by interpolating the first predicted value of the first sub-block and the second predicted value of the second sub-block.

Citation Information

Patent Citations

  • Video coding / decoding method and device

    CN101610413A

  • Interframe prediction encoding method, interframe prediction decoding method and equipment

    CN101873500A