Video encoding / decoding method and apparatus
By dividing video blocks into sub-blocks and determining intra-frame prediction modes, the method enhances video encoding/decoding performance and accuracy, addressing the inefficiencies in existing methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INST OF IMAGE TECH INC
- Filing Date
- 2026-02-04
- Publication Date
- 2026-04-28
AI Technical Summary
Existing video encoding/decoding methods lack performance improvements, particularly in intra-frame prediction, necessitating enhanced techniques for video processing efficiency.
The method involves dividing current blocks into sub-blocks, forming candidate groups, and determining intra-frame prediction modes based on these sub-block divisions to enhance encoding/decoding performance and accuracy.
This approach improves encoding/decoding performance and accuracy by enabling efficient intra-frame prediction at the sub-block level, adapting encoding orders, and constructing candidate groups effectively.
Smart Images

Figure 2026071326000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video encoding / decoding method and apparatus.
Background Art
[0002] With the spread of the Internet and mobile terminals and the development of information and communication technologies, the use of multimedia data has been increasing rapidly. Therefore, there is a growing need to improve the performance and efficiency of video processing systems in order to perform various services and operations by video prediction in various systems. However, the research and development results that can meet such an atmosphere are not sufficient.
[0003] As described above, in the conventional video encoding / decoding method and apparatus, there is a situation where performance improvement for video processing, particularly video encoding or video decoding, is required.
Summary of the Invention
Problems to be Solved by the Invention
[0004] An object of the present invention is to provide an intra-frame prediction method and apparatus. Another object of the present invention is to provide an intra-frame prediction method and apparatus in units of sub-blocks. Another object of the present invention is to provide a method and apparatus for determining the division and encoding order in units of sub-blocks.
Means for Solving the Problems
[0005] The video encoding / decoding method and apparatus according to the present invention constitute a candidate group regarding the division form of the current block, determine the division form of the current block into sub-blocks based on the candidate group and candidate index, induce an intra-frame prediction mode for the current block unit, and can perform intra-frame prediction of the current block based on the intra-frame prediction mode of the current block and the division form of the sub-blocks.
Effects of the Invention
[0006] According to the present invention, encoding / decoding performance can be improved by in-screen prediction at the sub-block level. Furthermore, according to the present invention, the accuracy of prediction can be improved by efficiently constructing a candidate group of sub-block division patterns. In addition, according to the present invention, the encoding / decoding efficiency of in-screen prediction can be improved by adaptively applying the encoding order at the sub-block level. [Brief explanation of the drawing]
[0007] [Figure 1] This is a conceptual diagram of a video encoding and decoding system according to an embodiment of the present invention. [Figure 2] This is a block diagram of a video encoding device according to one embodiment of the present invention. [Figure 3] This is a block diagram of an image decoding device according to one embodiment of the present invention. [Figure 4] This is an illustrative diagram showing various division configurations that can be obtained with the block division section of the present invention. [Figure 5] This is an illustrative diagram showing an in-screen prediction mode according to one embodiment of the present invention. [Figure 6] This is an illustrative diagram illustrating the reference pixel configuration used for in-screen prediction according to one embodiment of the present invention. [Figure 7] This is a conceptual diagram showing the target block for in-screen prediction according to one embodiment of the present invention, and the blocks adjacent to that target block. [Figure 8] This shows various sub-block division patterns obtainable from coded blocks. [Figure 9] This is an illustrative diagram of a reference pixel region used based on the in-screen prediction mode according to one embodiment of the present invention. [Figure 10] An example of an encoding order that can be had in a diagonal upright prediction mode according to one embodiment of the present invention is shown. [Figure 11] An example of a coding sequence that can be had in horizontal mode according to one embodiment of the present invention is shown. [Figure 12]An example of a coding order that can be had in a diagonal downright prediction mode according to one embodiment of the present invention is shown. [Figure 13] An example of a coding sequence that can be had in vertical mode according to one embodiment of the present invention is shown. [Figure 14] An example of an encoding sequence that can be had in a diagonal down-left direction according to one embodiment of the present invention is shown. [Figure 15] This is an illustrative diagram of an encoding order considering the in-screen prediction mode and division configuration according to one embodiment of the present invention. [Modes for carrying out the invention]
[0008] The video encoding / decoding method and apparatus according to the present invention constitute a candidate group of division patterns for the current block, determine the division pattern of the current block into subblocks based on the candidate group and candidate index, induce an in-screen prediction mode for each current block, and perform in-screen prediction of the current block based on the in-screen prediction mode of the current block and the division pattern of the subblocks.
[0009] A mode for carrying out the invention.
[0010] While the present invention can be modified in various ways and has many embodiments, we will attempt to illustrate and describe in detail specific embodiments with the drawings. However, this should not be understood as limiting the present invention to specific embodiments, but rather as including all modifications, equivalents, or substitutes that fall within the spirit and technical scope of the present invention.
[0011] Terms such as first, second, A, and B can be used to describe various components, but the components should not be limited to these terms. These terms are only used for the purpose of distinguishing one component from another. For example, within the scope not departing from the scope of the rights of the present invention, the first component can be named the second component, and similarly, the second component can be named the first component. The term "and / or" includes combinations of a plurality of related description items or any one of a plurality of related description items.
[0012] When a component is referred to as being "connected" or "coupled" to another component, it should be understood that it can be directly connected or coupled to the other component, but there can also be other components in between. On the other hand, when a component is referred to as being "directly connected" or "directly coupled" to another component, it should be understood that there are no other components in between.
[0013] The terms used in the present invention are only used to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In the present invention, terms such as "comprising" or "having" are intended to specify the presence of the features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should be understood that the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof is not precluded in advance.
[0014] Unless otherwise defined, all terms used herein, including technical or scientific terms, mean the same as those commonly understood by a person of ordinary skill in the technical field to which the present invention belongs. Terms defined as in a commonly used dictionary should be interpreted as having a meaning consistent with the context of the related art, and should not be interpreted in an idealized or overly formal sense unless clearly defined in the present invention.
[0015] Typically, it can be composed of one or more color spaces depending on the color format of the video. It can be composed of one or more pictures having a certain size or one or more pictures having other sizes depending on the color format. As an example, in the YCbCr color configuration, color formats such as 4:4:4, 4:2:2, 4:2:0, monochrome (composed only of Y) can be supported. As an example, in the case of YCbCr4:2:0, it can be composed of one luminance component (in this example, Y) and two chrominance components (in this example, Cb / Cr). Here, the composition ratio of the chrominance component and the luminance component can have a 1:2 ratio horizontally and vertically. As an example, in the case of 4:4:4, the horizontal and vertical lengths can have the same composition ratio. When composed of one or more color spaces as in the above example, the picture can perform division into each color space.
[0016] Videos can be classified into I, P, B, etc. according to the video type (e.g., picture type, slice type, tile type, etc.). The I video type can mean a video that is encoded / decoded by itself without using a reference picture, the P video type can mean a video that is encoded / decoded using a reference picture but only allows forward prediction, and the B video type can mean a video that is encoded / decoded using a reference picture and allows forward / backward prediction. However, depending on the encoding / decoding settings, some of the above types can be combined (combining P and B) or other configured video types can also be supported.
[0017] The diverse encoding / decoding information generated in this invention can be processed explicitly or implicitly. Here, explicit processing can be understood as generating selection information representing one of a group of candidates for encoding information in the form of a sequence, slice, tile, block, subblock, etc., and recording this in a bitstream, and then the decoder parsing the related information at the same level as the encoder to restore the decoded information. Here, implicit processing can be understood as processing the encoding / decoding information in the encoder and decoder using the same process, rules, etc.
[0018] Figure 1 is a conceptual diagram of a video encoding and decoding system according to an embodiment of the present invention.
[0019] Referring to Figure 1, the video encoding device 105 and decoding device 100 can be a user terminal such as a personal computer (PC), notebook computer, personal digital assistant (PDA), portable multimedia player (PMP), PlayStation Portable (PSP), wireless communication terminal (Wireless Communication Terminal), smartphone (Smartphone), or TV, or a server terminal such as an application server or service server. They can also include a variety of devices such as a communication modem for communicating with various devices or wired / wireless networks, memory 120, 125 for storing various programs and data for encoding or decoding video, or processor 110, 115 for executing programs, performing calculations, and controlling.
[0020] Furthermore, the video encoded into a bitstream by the video encoding device 105 can be transmitted to the video decoding device 100 in real time or non-real time via a wired wireless network (Network) such as the Internet, a short-range wireless communication network, a wireless LAN network, a Wi-Fi network, or a mobile communication network, or via various communication interfaces such as a cable or a Universal Serial Bus (USB), where it can be decoded and restored as video for playback. In addition, the video encoded into a bitstream by the video encoding device 105 can be transmitted from the video encoding device 105 to the video decoding device 100 via a computer-readable recording medium.
[0021] The aforementioned video encoding device and video decoding device can be separate devices, but they can be implemented as a single video encoding / decoding device. In that case, some components of the video encoding device are substantially the same technical elements as some components of the video decoding device, and can be implemented to include at least the same structure or perform at least the same function.
[0022] Therefore, in the detailed explanations of the following technical elements and their operating principles, redundant explanations of corresponding technical elements will be omitted. Also, since the video decoding device corresponds to a computer device that applies the video encoding method performed by the video encoding device to decoding, the following explanation will focus on the video encoding device.
[0023] A computer device may include a memory for storing a program or software module that embodies a video encoding method and / or a video decoding method, and a processor linked to the memory for executing the program. Here, a video encoding device may be called an encoder, and a video decoding device may be called a decoder.
[0024] Figure 2 is a block diagram of a video encoding device according to one embodiment of the present invention.
[0025] Referring to Figure 2, the video encoding device 20 may include a prediction unit 200, a subtraction unit 205, a conversion unit 210, a quantization unit 215, an inverse quantization unit 220, an inverse conversion unit 225, an addition unit 230, a filter unit 235, an encoded picture buffer 240, and an entropy encoding unit 245.
[0026] The prediction unit 200 can be implemented using a software module called a prediction module, and can generate predicted blocks for blocks to be encoded using either intra-prediction or inter-prediction. The prediction unit 200 can generate predicted blocks by predicting the current block to be encoded in the video. In other words, the prediction unit 200 can generate predicted blocks having predicted pixel values for each pixel of the current block to be encoded in the video, generated by predicting the pixel value of each pixel using intra-prediction or inter-prediction. Furthermore, the prediction unit 200 can transmit information necessary for generating predicted blocks, such as information about the prediction mode, such as intra-prediction mode or inter-prediction mode, to the encoding unit so that it can encode information about the prediction mode. Here, the processing unit in which prediction is performed, the prediction method, and the processing unit in which the specific content is determined can be determined by the encoding / decoding settings. For example, the prediction method and prediction mode can be determined at the prediction unit, and the execution of the prediction can be performed at the conversion unit.
[0027] In the inter-screen prediction unit, the motion prediction method can be divided into a motion model and a non-motion model. In the case of a motion model, prediction is performed considering only translation, while in the case of a non-motion model, prediction can be performed considering not only translation but also rotation, perspective, zoom in / out, and other movements. Assuming unidirectional prediction, a motion model may require one motion vector, while a non-motion model may require one or more motion vectors. In the case of a non-motion model, each motion vector can be information applied to a pre-defined position of the current block, such as the upper-left vertex or upper-right vertex of the current block, and the position of the area to be predicted in the current block can be obtained on a pixel-by-pixel or sub-block basis using the corresponding motion vector. Depending on the motion model, some processes described later can be applied commonly in the inter-screen prediction unit, while some processes can be applied individually.
[0028] The inter-screen prediction unit may include a reference picture configuration unit, a motion estimation unit, a motion compensation unit, a motion information determination unit, and a motion information encoding unit. The reference picture configuration unit may include previously encoded pictures (L0, L1) centered on the current picture in a reference picture list. A prediction block can be obtained from the reference pictures included in the reference picture list, and the current video can also be composed of a reference picture depending on the encoding settings and may be included in at least one of the reference picture lists.
[0029] In the inter-screen prediction unit, the reference picture configuration unit may include a reference picture interpolation unit, which can perform interpolation for small numbers of pixels depending on the interpolation accuracy. For example, an interpolation filter based on an 8-tap DCT can be applied for the luminance component, and an interpolation filter based on a 4-tap DCT can be applied for the chrominance component.
[0030] In the inter-screen prediction unit, the motion estimation unit searches for blocks with a high correlation to the current block using a reference picture, and can use various methods such as FBMA (Full search-based block matching algorithm) and TSS (Three step search). The motion compensation unit refers to the process of acquiring predicted blocks through the motion estimation process.
[0031] In the inter-screen prediction unit, the motion information determination unit can perform the process of selecting the optimal motion information for the current block, and the motion information can be encoded by motion information encoding modes such as Skip Mode, Merge Mode, and Competition Mode. The modes can be configured by combining modes supported by the motion model, and examples of such modes may be Skip Mode (movement), Skip Mode (non-movement), Merge Mode (movement), Merge Mode (non-movement), Competition Mode (movement), and Competition Mode (non-movement). Depending on the encoding settings, some of the modes may be included in the candidate group.
[0032] The motion information coding mode can obtain predicted values of motion information (motion vector, reference picture, predicted direction, etc.) for the current block using at least one candidate block, and can generate optimal candidate selection information when supporting two or more candidate blocks. In skip mode (no residual signal) and merge mode (with residual signal), the predicted values can be used directly as motion information for the current block, while in competition mode, difference information between the motion information of the current block and the predicted values can be generated.
[0033] The candidate group for predicting the motion information of the current block can have an adaptive and diverse configuration depending on the motion information coding mode. The candidate group can include motion information of blocks spatially adjacent to the current block (e.g., blocks to the left, above, upper left, upper right, lower left, etc.), motion information of blocks temporally adjacent to the current block, and motion information of combinations of spatial and temporal candidates.
[0034] The aforementioned temporally adjacent blocks include other blocks in the video that correspond to (or are equivalent to) the current block, and can refer to blocks located to the left, right, above, below, upper left, upper right, lower left, lower right, etc., with respect to the block in question. The aforementioned combined motion information can refer to information obtained as an average, median, etc., from the motion information of spatially adjacent blocks and the motion information of temporally adjacent blocks.
[0035] A priority order can exist for constructing a group of candidate motion information prediction values. The procedure to be included in constructing the candidate prediction value group can be determined by the priority order, and the candidate group construction can be completed when the number of candidates (determined by the motion information coding mode) is satisfied by the priority order. Here, the priority order can be determined in the following order: motion information of spatially adjacent blocks, motion information of temporally adjacent blocks, and motion information of combinations of spatial and temporal candidates, but other variations are also possible.
[0036] For example, among spatially adjacent blocks, the candidate group can include blocks in the order of left-top-upper-upper-right-lower-left-upper-left, and among temporally adjacent blocks, the candidate group can include blocks in the order of lower-right-middle-right-lower-bottom.
[0037] The subtraction unit 205 can generate a residual block by subtracting the predicted block from the current block. In other words, the subtraction unit 205 can generate a residual block, which is a residual signal in block form, by calculating the difference between the pixel value of each pixel in the current block to be encoded and the predicted pixel value of each pixel in the predicted block generated by the prediction unit. The subtraction unit 205 can also generate residual blocks using units other than block units obtained by the block division unit described later.
[0038] The conversion unit 210 can convert signals belonging to the spatial domain into signals belonging to the frequency domain, and the signals obtained through the conversion process are called transformed coefficients. For example, a residual block having a residual signal transmitted from the subtraction unit can be converted to obtain a transformed block having transformed coefficients. The input signal is determined by the encoding setting, and this is not limited to residual signals.
[0039] The transformation unit can transform residual blocks using transformation techniques such as the Hadamard Transform, Discrete Sine Transform (DST-Based Transform), and Discrete Cosine Transform (DCT-Based Transform), but is not limited to these; a variety of improved and modified transformation techniques can be used.
[0040] The system can support at least one of the transformation techniques, and can support at least one detailed transformation technique for each transformation technique. Here, the detailed transformation technique may be a transformation technique configured such that a portion of the basis vectors differs in each transformation technique.
[0041] For example, in the case of DCT, it can support one or more detailed conversion techniques from DCT-1 to DCT-8, and in the case of DST, it can support one or more detailed conversion techniques from DST-1 to DST-V8. A group of candidate conversion techniques can be formed by combining some of the aforementioned detailed conversion techniques. As an example, DCT-2, DCT-8, and DST-7 can be combined into a group of candidate conversion techniques to perform the conversion.
[0042] The transformation can be performed in the horizontal or vertical direction. For example, by performing a one-dimensional transformation in the horizontal direction using the DCT-2 transformation technique and a one-dimensional transformation in the vertical direction using the DST-7 transformation technique, a total two-dimensional transformation can be performed, thereby converting pixel values in the spatial domain to the frequency domain.
[0043] The conversion can be performed using a single fixed conversion technique, or by adaptively selecting the conversion technique based on the encoding / decoding settings. In the adaptive case, the conversion technique can be selected using explicit or implicit methods. In the explicit case, conversion technique selection information or conversion technique set selection information applied horizontally and vertically can be generated in units such as blocks. In the implicit case, encoding settings can be defined by video type (I / P / B), color components, block size, shape, in-screen prediction mode, etc., thereby allowing the selection of a predefined conversion technique.
[0044] Furthermore, the aforementioned partial conversion may be omitted depending on the encoding settings. In other words, one or more horizontal / vertical units may be omitted explicitly or implicitly.
[0045] Furthermore, the conversion unit can transmit the information necessary to generate the conversion block to the encoding unit for encoding, record the resulting information into a bitstream, transmit it to the decoder, and the decoder's decoding unit can parse the information and use it in the inverse conversion process.
[0046] The quantization unit 215 can quantize the input signal. Here, the signal obtained through the quantization process is called the quantized coefficient. For example, a residual block having residual conversion coefficients transmitted from the conversion unit can be quantized to obtain a quantized block having quantized coefficients. The input signal is determined by the encoding settings, which are not limited to residual conversion coefficients.
[0047] The quantization unit can quantize the transformed residual blocks using quantization techniques such as Dead Zone Uniform Threshold Quantization and Quantization Weighted Matrix, but is not limited to these; it can use a variety of quantization techniques that are improvements and modifications of these.
[0048] The quantization process can be omitted depending on the encoding settings. For example, the quantization process (including the reverse process) can be omitted depending on the encoding settings (e.g., a quantization parameter of 0, i.e., a lossless compression environment). As another example, the quantization process can be omitted if the compression performance due to quantization is not achieved due to the characteristics of the video. Here, the region in the quantization block (M×N) where the quantization process is omitted can be the entire region or a part of it (M / 2×N / 2, M×N / 2, M / 2×N, etc.), and the quantization omission selection information can be determined implicitly or explicitly.
[0049] The quantization unit can transmit the information necessary to generate quantization blocks to the encoding unit for encoding, record the resulting information into a bitstream, transmit it to the decoder, and the decoder's decoding unit can parse the information and use it in the dequantization process.
[0050] In the above example, the explanation was based on the assumption that the residual block is transformed and quantized by a transformation unit and a quantization unit. However, it is also possible to transform the residual signal to generate a residual block with transformation coefficients and not perform the quantization process. Furthermore, it is possible to perform only the quantization process without converting the residual signal of the residual block into transformation coefficients, or even to perform neither the transformation nor the quantization process. This can be determined by the encoder settings.
[0051] The inverse quantization unit 220 inversely quantizes the residual blocks that have been quantized by the quantization unit 215. That is, the inverse quantization unit 220 inversely quantizes the sequence of quantized frequency coefficients to generate residual blocks that have frequency coefficients.
[0052] The inverse transform unit 225 inversely transforms the residual block that has been inversely quantized by the inverse quantization unit 220. That is, the inverse transform unit 225 inversely transforms the frequency coefficients of the inversely quantized residual block to generate a residual block with pixel values, i.e., a restored residual block. Here, the inverse transform unit 225 can perform the inverse transform by using the same transformation method used in the transformation unit 210 in reverse.
[0053] The adder 230 adds the predicted block predicted by the prediction 200 and the residual block restored by the inverse transform 225 to restore the current block. The restored current block is stored in the encoded picture buffer 240 as a reference picture (or reference block) and can be used as a reference picture when encoding the next block or other blocks and pictures in the future.
[0054] The filter unit 235 may include one or more post-processing filter steps, such as a deblocking filter, SAO (Sample Adaptive Offset), and ALF (Adaptive Loop Filter). A deblocking filter can remove block distortion that occurs at the boundaries between blocks from the restored picture. An ALF can perform filtering based on a comparison between the restored image and the original image after the blocks have been filtered through the deblocking filter. An SAO can restore the offset difference from the original image on a pixel-by-pixel basis for residual blocks to which the deblocking filter has been applied. Such post-processing filters can be applied to the restored picture or blocks.
[0055] The encoded picture buffer 240 can store blocks or pictures restored via the filter unit 235. The restored blocks or pictures stored in the encoded picture buffer 240 can be provided to the prediction unit 200, which performs in-screen prediction or inter-screen prediction.
[0056] The entropy coding unit 245 generates a quantization coefficient sequence by scanning the generated quantization frequency coefficient sequence using various scanning methods, and outputs it by encoding it using entropy coding techniques. The scan pattern can be set to one of various patterns such as zigzag, diagonal, or raster. It can also generate coded data containing coded information transmitted from each component and output it as a bitstream.
[0057] Figure 3 is a block diagram of an image decoding device according to one embodiment of the present invention.
[0058] Referring to Figure 3, the video decoding device 30 may include an entropy decoding unit 305, a prediction unit 310, an inverse quantization unit 315, an inverse transform unit 320, an adder / subtractor 325, a filter 330, and a decoding picture buffer 335.
[0059] Furthermore, the prediction unit 310 may also include an in-screen prediction module and an inter-screen prediction module.
[0060] First, once the video bitstream transmitted from the video encoding device 20 is received, it can be transmitted to the entropy decoding unit 305.
[0061] The entropy decoding unit 305 can decode the bitstream, decode the quantized coefficients, and decode the decoded data which includes the decoded data transmitted to each component.
[0062] The prediction unit 310 can generate prediction blocks based on the data transmitted from the entropy decoding unit 305. Here, a reference picture list can also be constructed using the default configuration technique based on the reference video stored in the decoded picture buffer 335.
[0063] The inter-screen prediction unit may include a reference picture configuration unit, a motion compensation unit, and a motion information decoding unit, some of which can perform the same process as the encoder, and others which can perform the reverse induction process.
[0064] The inverse quantization unit 315 can inverse quantize the quantization conversion coefficients that are provided as a bitstream and decoded by the entropy decoding unit 305.
[0065] The inverse transform unit 320 can generate residual blocks by applying inverse DCT, inverse integer transform, or a similar inverse transform technique to the transform coefficients.
[0066] Here, the inverse quantization unit 315 and the inverse transformation unit 320 can be implemented in various ways by reversing the processes performed by the transformation unit 210 and the quantization unit 215 of the video encoding device 20 described above. For example, the same processes and inverse transformations shared with the transformation unit 210 and the quantization unit 215 can be used, and the transformation and quantization processes can be reversed using information about the transformation and quantization processes (e.g., transformation size, transformation shape, quantization type, etc.) from the video encoding device 20.
[0067] The remaining blocks after the inverse quantization and inverse transformation processes can be added to the predicted blocks derived by the prediction unit 310 to generate the reconstructed image blocks. Such addition can be performed by the adder / subtractor 325.
[0068] Filter 330 can also apply a deblocking filter to the restored video blocks to remove blocking phenomena as needed, and other loop filters can be additionally used before and after the decoding process to improve video quality.
[0069] The restored and filtered video blocks can be stored in the decoded picture buffer 335.
[0070] Although not shown in the drawings, the video encoding / decoding device may further include a picture division unit and a block division unit.
[0071] The picture division unit can divide (or partition) a picture into at least one processing unit such as a color space (e.g., YCbCr, RGB, or XYZ), tile, slice, or basic coding unit (or maximum coding unit, Coding Tree Unit, CTU), and the block division unit can divide a basic coding unit into at least one processing unit (e.g., coding, prediction, transformation, quantization, entropy, and in-loop filter units).
[0072] Basic coding units can be obtained by dividing a picture into fixed intervals in the horizontal and vertical directions. Based on this, divisions such as tiles and slices can be performed, but are not limited to these. While such division units as tiles and slices can consist of integer multiples of the basic coding block, exceptional cases may occur for division units located at the boundaries of the image. For this reason, adjustments to the basic coding block size may be necessary.
[0073] For example, a picture can be divided into basic encoding units and then split into those units, or a picture can be divided into those units and then split into basic encoding units. This invention explains the case where the partitioning and division order of each unit are as described above, but is not limited to this, and the latter case is also possible depending on the encoding / decoding settings. In the latter case, it is possible to modify the basic encoding unit so that the size is adaptive depending on the division unit (tile, etc.). That is, it means that it is possible to support basic encoding blocks having different sizes for each division unit.
[0074] In this invention, the example described below will be based on the assumption that the picture is divided into basic coding units. This basic assumption may mean that the picture is not divided into tiles or slices, or that the picture is a single tile or a single slice. However, as mentioned above, it should be understood that the various embodiments described below can also be applied similarly or with modifications to cases where each division unit (tile, slice, etc.) is first divided and then divided into basic coding units by the acquired units (i.e., cases where each division unit is not an integer multiple of the basic coding unit).
[0075] Within the aforementioned division units, a slice can consist of a bundle of at least one block that is consecutive according to a scan pattern, and a tile can consist of a rectangular bundle of spatially adjacent blocks. Other additional division units can be supported and defined accordingly. Slices and tiles can be division units that support purposes such as parallel processing, and for this purpose, references between division units can be restricted (i.e., not possible).
[0076] A slice can generate division information for each unit using information about the starting position of consecutive blocks, and in the case of tiles, it can generate information about horizontal and vertical dividing lines or position information for the tiles (e.g., top-left, top-right, bottom-left, bottom-right positions).
[0077] Here, slices and tiles can be divided into multiple units depending on the encoding / decoding settings.
[0078] For example, some units This can be a unit containing configuration information that affects the encoding / decoding process (i.e., including a tile header or slice header), and some units It can be a unit that does not include configuration information. Or, a part of the unit A unit can be a unit that cannot reference other units during the encoding / decoding process, and some units It can be a unit that can be referenced. Also, some units is other units Can a hierarchical relationship be established that includes some units? is other units It is possible to have an equal relationship with them.
[0079] Here, A and B can be a slice and a tile (or a tile and a slice). Alternatively, A and B can consist of a slice and one tile. For example, A can be a slice / tile of type 1, and B can be a slice / tile of type 2.
[0080] Here, Type 1 and Type 2 can each be one slice or tile. Alternatively, Type 1 can be multiple slices or tiles (a set of slices or tiles) (including Type 2), and Type 2 can be one slice or tile.
[0081] As already mentioned above, the present invention is described assuming that the picture consists of a single slice or tile, but if two or more division units occur, the above description can be applied to the embodiments described later. Furthermore, A and B are examples of characteristics that division units may have, and it is also possible to construct examples by combining A and B from each example.
[0082] On the other hand, the block division unit can obtain information about the basic coding unit from the picture division unit, and the basic coding unit can mean the basic (or starting) unit for prediction, transformation, quantization, etc., in the video coding / decoding process. Here, the basic coding unit can consist of one luminance basic coding block (or maximum coding block; Coding Tree Block; CTB) and two basic chrominance coding blocks, depending on the color format (YCbCr in this example), and the size of each block can be determined by the color format. Then, a coding block (Coding Block; CB) can be obtained through the division process. A coding block can be understood as a unit that cannot be divided into further coding blocks by certain limitations, and can be set as the starting unit for division into lower units. In this invention, blocks are not limited to squares but can be understood as a broad concept that includes various shapes such as triangles and circles. For the sake of explanation, we will assume the case of a rectangle.
[0083] The following discussion focuses on a single color component, but it should be understood that this can be applied to other color components in proportion to the ratio determined by the color format (for example, in the case of YCbCr 4:2:0, the ratio of the horizontal to vertical lengths of the luminance and chrominance components is 2:1). Furthermore, while it is possible to perform block divisions that are dependent on other color components (for example, Cb / Cr depending on the Y block division result), it should be understood that independent block divisions are possible for each color component. Additionally, while a single common block division setting (considering its proportionality to the length ratio) can be used, it is also necessary to consider and understand the possibility of using individual block division settings for each color component.
[0084] In the block division section, blocks can be represented as M×N, and the maximum and minimum values of each block can be obtained within the range. For example, if the maximum value of a block is determined to be 256×256 and the minimum value to be 4×4, then 2 m ×2 n A block of size (in this example, m and n are integers from 2 to 8) or 2 m ×2 m A block of size (in this example, m and n are integers from 2 to 128) or an m x m block (in this example, m and n are integers from 4 to 256) can be obtained. Here, m and n can be the same or different, and one or more ranges can be generated that support the blocks such as the maximum and minimum values.
[0085] For example, information about the maximum and minimum size of a block can be generated, and information about the maximum and minimum size of a block can be generated with certain division settings. In the former case, this can be range information for the maximum and minimum sizes that can occur within the video, and in the latter case, this can be information about the maximum and minimum sizes that can occur with certain division settings. Here, the division settings can be defined by video type (I / P / B), color component (YCbCr, etc.), block type (encoding / prediction / conversion / quantization, etc.), division type (Index or Type), division method (Tree method such as QT, BT, TT, etc., Index method such as SI2, SI3, SI4, etc.).
[0086] Furthermore, there can be restrictions on the aspect ratio (block shape) that a block can have, and boundary value conditions can be set for this. Here, only blocks less than or equal to an arbitrary boundary value (k) can be supported, where k can be defined by the aspect ratio such as A / B (A being longer or the same in terms of aspect ratio, and B being any other value), and can be one or more real numbers such as 1, 5, 2, 3, 4, etc. As in the example above, it is possible to support restrictions on the shape of a single block in the image, or to support one or more restrictions by setting divisions.
[0087] In summary, the feasibility of supporting block splitting can be determined by the scope and conditions described above, as well as the splitting settings described later. For example, if the block conditions for supporting candidate blocks (child blocks) resulting from the splitting of a block (parent block) are met, the split can be supported; otherwise, the split cannot be supported.
[0088] The block division section can be configured in relation to each component of the video encoding device and the decoding device, and the size and shape of the blocks can be determined through this process. Here, the blocks to be configured can be defined differently depending on the component; for example, a prediction block for the prediction section, a transformation block for the transformation section, and a quantization block for the quantization section. However, it is not limited to this, and additional block units can be defined by other components. In this invention, the case where the input and output are rectangular in each component will be described in detail, but some components can have inputs / outputs of other shapes (e.g., triangles).
[0089] The size and shape of the initial (or starting) block of the block division unit can be determined from the higher-level unit. The initial block can be divided into smaller blocks, and once the optimal size and shape are determined by the block division, that block can be determined as the initial block of the lower-level unit. Here, the higher-level unit can be a coding block, and the lower-level unit can be a prediction block or a transformation block, but is not limited to these, and various variations are possible. Once the initial block of the lower-level unit is determined as in the above example, a division process to find the block of the optimal size and shape can be carried out, similar to the higher-level unit.
[0090] In summary, the block division unit can divide a basic coding block (or maximum coding block) into at least one coding block, and can divide a coding block into at least one prediction block / transformation block / quantization block. Furthermore, it can divide a prediction block into at least one transformation block / quantization block, and can divide a transformation block into at least one quantization block. Here, some blocks can have a dependent relationship with other blocks (i.e., defined by a higher unit and a lower unit) or an independent relationship. For example, a prediction block can be a higher unit of a transformation block or an independent unit of a transformation block, and various relationship settings are possible depending on the type of block.
[0091] The encoding / decoding settings can be used to determine whether or not higher-level and lower-level units can be combined. Here, the coupling between units means that the division from the higher-level unit to the lower-level unit is not performed, but rather the encoding / decoding process (e.g., prediction unit, transformation unit, inverse transformation unit, etc.) of the lower-level unit to the block (size and shape) of the higher-level unit is performed. In other words, it can be said that the division process is shared among multiple units, and the division information is generated in one unit (e.g., the higher-level unit) within that process.
[0092] For example, (when a coding block is combined with a prediction block and a transformation block), the coding block can perform the prediction process, transformation, and inverse transformation process.
[0093] For example, (when a coding block is combined with a prediction block), the prediction process can be performed using the coding block, and the transformation and inverse transformation processes can be performed using a transformation block that is identical to or smaller than the coding block.
[0094] For example, (when a coding block is combined with a transform block), the prediction process can be performed using a prediction block that is identical to or smaller than the coding block, and the transform and inverse transform processes can be performed using the coding block.
[0095] For example, (when a prediction block is combined with a transformation block), the prediction process can be carried out using a prediction block that is identical to or smaller than the coding block, and the transformation and inverse transformation processes can be carried out using the prediction block.
[0096] For example, (when no blocks are combined), the prediction process can be performed using prediction blocks that are identical to or smaller than the coding blocks, and the transformation and inverse transformation processes can be performed using transformation blocks that are identical to or smaller than the coding blocks.
[0097] The above examples illustrate a variety of cases regarding coding, prediction, and transformation blocks, but are not limited to these.
[0098] The combination of the aforementioned units can support fixed settings in the video, or it can support adaptive settings considering various encoding / decoding elements. Here, the encoding / decoding elements may include video type, color components, encoding mode (Intra / Inter), division settings, block size / shape / position, aspect ratio, prediction-related information (e.g., in-screen prediction mode, inter-screen prediction mode, etc.), transformation-related information (e.g., transformation technique selection information, etc.), quantization-related information (e.g., quantization domain selection information, quantized transformation coefficient encoding information, etc.), and so on.
[0099] As described above, when a block of the optimal size and shape is found, mode information (e.g., segmentation information) can be generated for it. The mode information, along with information generated in the components to which the block belongs (e.g., prediction-related information and transformation-related information), can be recorded in the bitstream and transmitted to the decoder, where it can be parsed in units of the same level and used in the video decoding process.
[0100] The following describes the division method, and for the sake of explanation, we assume that the initial block is a square. However, this method can be applied identically or similarly to the case where the initial block is a rectangle, and is not limited to this case.
[0101] The block partitioning unit can support various types of partitioning. For example, it can support tree-based partitioning or index-based partitioning, and can support other methods as well. Tree-based partitioning can determine the partitioning form based on various types of information (e.g., feasibility of partitioning, tree type, partitioning direction, etc.), and index-based partitioning can determine the partitioning form based on predetermined index information.
[0102] Figure 4 is an illustrative diagram showing various division configurations that can be obtained with the block division unit of the present invention. In this example, it is assumed that the division configuration shown in Figure 4 is obtained by a single division operation (or process), but it is not limited to this, and it is also possible to obtain it by multiple division operations. In addition, additional division configurations not shown in Figure 4 are possible.
[0103] (Partitioning of the tree base)
[0104] In the tree-based partitioning of the present invention, the tree scheme can support quad trees (Quad Tree.QT), binary trees (Binary Tree.BT), terminally trees (Ternary Tree.TT), and the like. Supporting one tree scheme can be called single-tree partitioning, and supporting two or more tree schemes can be called multiple-tree partitioning.
[0105] QT refers to a method (n) in which the block is divided into two horizontally and two vertically (i.e., four divisions); BT refers to a method (b-g) in which the block is divided into two horizontally or vertically; and TT refers to a method (h-m) in which the block is divided into three horizontally or vertically.
[0106] Here, in the case of QT, it is also possible to support a method (o, p) that divides into four sections by limiting the division direction to one direction, either horizontally or vertically. In the case of BT, it is possible to support only methods with equal sizes (b, c), only methods with unequal sizes (d~g), or a combination of both. In the case of TT, it is possible to support only methods with a division biased in a specific direction (e.g., 1:1:2, 2:1:1 in the left-to-right or top-to-bottom direction) (h, j, k, m), only methods with a central division (e.g., 1:2:1) (i, l), or a combination of both. It is also possible to support a method (q) that divides into four sections each in the horizontal and vertical directions (i.e., 16 sections).
[0107] Furthermore, the tree scheme can support either a method of z-partitioning limited to the horizontal partitioning direction (b, d, e, h, i, j, o), a method of z-partitioning limited to the vertical partitioning direction (c, f, g, k, l, m, p), or a combination of both. Here, z can be an integer of 2 or more, such as 2, 3, or 4.
[0108] In this invention, we will explain assuming that QT supports n, BT supports b and c, and TT supports i and l.
[0109] The encoding / decoding settings can support one or more schemes within the tree partition. For example, it can support QT, QT / BT, or QT / BT / TT.
[0110] The above example is for a case where the basic tree partitioning is QT, and BT and TT are included in the additional partitioning scheme depending on whether other trees are supported, but various modifications are possible. Here, information about whether other trees are supported (bt_enabled_flag, tt_enabled_flag, bt_tt_enabled_flag, etc., which can have a value of 0 or 1, where 0 means no support and 1 means support) can be implicitly determined by the encoding / decoding settings or can be explicitly determined in units such as sequences, pictures, slices, and tiles.
[0111] The partitioning information can include information about whether partitioning is possible (tree_part_flag, or qt_part_flag, bt_part_flag, tt_part_flag, bt_tt_part_flag; can have a value of 0 or 1, where 0 means no partitioning and 1 means partitioning). Additionally, depending on the partitioning method (BT and TT), information about the partitioning direction (dir_part_flag, or bt_dir_part_flag, tt_dir_part_flag, bt_tt_dir_part_flag; can have a value of 0 or 1, where 0 means horizontal and 1 means vertical) can be added, and this information can be generated when partitioning is performed.
[0112] When supporting multiple tree partitions, various partition information configurations are possible. Next, we will explain an example of how partition information is structured at a single depth level (i.e., although one or more partition depths are supported and recursive partitioning is possible, this example is for the sake of explanation).
[0113] In example (1), we check the information regarding whether or not the division is possible. If the division is not to be carried out, the division is terminated.
[0114] If a partition is to be performed, the selection information for the partition type (for example, tree_idx; 0 for QT, 1 for BT, 2 for TT) is checked. At this point, the partition direction information is further checked based on the selected partition type, and the process moves to the next stage (if additional partitioning is possible for reasons such as the partition depth not reaching the maximum, the process starts again from the beginning; if partitioning is not possible, the partitioning is terminated).
[0115] In example (2), information regarding the feasibility of partitioning for some tree schemes (QT) is checked, and the process moves to the next stage. If partitioning is not performed at this stage, information regarding the feasibility of partitioning for some tree schemes (BT) is checked. If partitioning is not performed at this stage, information regarding the feasibility of partitioning for some tree schemes (TT) is checked. If partitioning is not performed at this stage, the partitioning process is terminated.
[0116] If a partial tree-based (QT) partitioning is to be performed, proceed to the next stage. If a partial tree-based (BT) partitioning is to be performed, check the partitioning direction information and proceed to the next stage. If a partial tree-based (TT) partitioning is to be performed, check the partitioning direction information and proceed to the next stage.
[0117] In example (3), we check the feasibility of partitioning for some tree schemes (QT). If partitioning is not performed, we check the feasibility of partitioning for some tree schemes (BT and TT). If partitioning is not performed, we terminate the partitioning process.
[0118] If a partitioning method using a specific tree scheme (QT) is to be performed, proceed to the next stage. Alternatively, if a partitioning method using a specific tree scheme (BT and TT) is to be performed, confirm the partitioning direction information and proceed to the next stage.
[0119] The above examples illustrate cases where tree partitioning has a priority (examples 2 and 3) or does not (example 1), but various variations are possible. Also, the above examples illustrate cases where the current partitioning is independent of the results of previous partitioning stages, but it is also possible to set the current partitioning to depend on the results of previous partitioning stages.
[0120] For example, in examples 1 to 3, if some tree-like partitioning (QT) was performed in a previous stage before moving to the current stage, then it is possible to support the same tree-like partitioning (QT) in the current stage as well.
[0121] On the other hand, if the system proceeded to the current stage by performing other tree-type partitioning (BT or TT) without performing some tree-type partitioning (QT) in a previous stage, it is possible to configure the system to support some tree-type partitioning (BT and TT) in subsequent stages, including the current stage, while excluding some tree-type partitioning (QT).
[0122] In such cases, the tree structure supporting block partitioning is adaptive, meaning that the partitioning information structure described above can also be configured differently (assuming the example described later is the third example). That is, if partitioning of some tree schemes (QT) was not performed in a previous stage in the above example, the partitioning process can be carried out at the current stage without considering some tree schemes (QT). Also, partitioning information about the relevant tree schemes (e.g., information on whether partitioning is possible, partitioning direction information, etc.) is also relevant. <qt>Then, information regarding whether or not it can be divided can be removed and the structure can be configured.
[0123] The above example is for an adaptive partitioning information configuration when block partitioning is permitted (for example, when the block size is within the range between the maximum and minimum values, and the partitioning depth of each tree scheme does not reach the maximum depth <allowable depth>), but an adaptive partitioning information configuration is also possible when block partitioning is restricted (for example, when the block size is not within the range between the maximum and minimum values, and the partitioning depth of each tree scheme reaches the maximum depth).
[0124] As mentioned above, in the present invention, tree-based partitioning can be performed in a recursive manner. For example, if the partitioning flag of a coding block with a partitioning depth of k is 0, the coding of the coding block is performed by coding blocks with a partitioning depth of k, and if the partitioning flag of a coding block with a partitioning depth of k is 1, the coding of the coding block is performed by N sub-coding blocks with a partitioning depth of k+1 (where N is an integer of 2 or more, such as 2, 3, 4) according to the partitioning scheme.
[0125] The aforementioned sub-encoded block can be further set as an encoded block (k+1) and divided into a sub-encoded block (k+2) through the process described above. Such a hierarchical division method can be determined by division settings such as the division range and the allowable division depth.
[0126] Here, the bitstream structure for representing the partitioning information can be selected from one or more scanning methods. For example, the bitstream of partitioning information can be constructed based on the order of partitioning depth, or based on whether or not partitioning is possible.
[0127] For example, if the sorting order is based on the depth of the division, the method involves first obtaining the division information at the current depth level based on the first block, and then obtaining the division information at the next depth level. If the sorting feasibility is used as the criterion, it means prioritizing the acquisition of additional division information for the blocks that have been divided based on the first block, and other additional scanning methods can be considered.
[0128] (Splitting of the index infrastructure)
[0129] The index base partitioning method of the present invention can support methods such as the CSI (Constant Split Index) method and the VSI (Variable Split Index) method.
[0130] The CSI method can be a method for obtaining k subblocks by dividing in a predetermined direction, where k can be an integer of 2 or more, such as 2, 3, or 4. More specifically, it can be a partitioning method that determines the size and shape of the subblocks based on the k value, regardless of the size and shape of the block. Here, the predetermined direction can be a combination of one or more directions from horizontal, vertical, and diagonal directions (e.g., upper left to lower right, or lower left to upper right).
[0131] The CSI partitioning method for the index base of the present invention may include candidates to be divided into z parts in one direction, either horizontally or vertically. Here, z can be an integer of 2 or more, such as 2, 3, 4, and one of the horizontal and vertical lengths of each subblock may be the same, while the other may be the same or different. The ratio of the horizontal to vertical lengths of the subblocks is A1:A2:...:A Z A1~A Z can be an integer greater than or equal to 1, such as 1, 2, or 3.
[0132] Furthermore, the list can include candidates that are divided into x and y subblocks horizontally and vertically, respectively. Here, x and y can be integers greater than or equal to 1, such as 1, 2, 3, and 4, but the case where x and y are both 1 (since a already exists) can be restricted. In Figure 4, the case where the horizontal-to-vertical ratio of each subblock is the same is shown, but the list can also include candidates where the ratios are different.
[0133] Furthermore, it can include candidates that are divided into w parts in one of the diagonal directions (upper left to lower right) or in one of the diagonal directions (lower left to upper right), where w can be an integer of 2 or more, such as 2 or 3.
[0134] Referring to Figure 4, the divisions can be categorized into symmetrical (b) and asymmetrical (d, e) divisions based on the length ratio of each subblock, and further categorized into divisions biased in a specific direction (k, m) and divisions centrally positioned (k). The divisions can be defined by various encoding / decoding elements, including not only the length ratio of the subblocks but also the shape of the subblocks, and the supported divisions can be implicitly or explicitly determined by the encoding / decoding settings. Therefore, a group of candidate division schemes for the index base can be determined based on the supported divisions.
[0135] On the other hand, the VSI method is a method in which one or more subblocks are obtained by dividing in a predetermined direction while the width (w) or height (h) of the subblocks remains fixed, and w and h can be integers of 1 or more, such as 1, 2, 4, or 8. More specifically, it can be a division method in which the number of subblocks is determined based on the size and shape of the block and the w or n value.
[0136] The VSI partitioning method for the index base of the present invention can partition candidates while fixing either the horizontal or vertical dimension of the subblock. Alternatively, it can partition candidates while fixing both the horizontal and vertical dimensions of the subblock. Because the horizontal and vertical dimensions of the subblock are fixed, it can have the characteristic of allowing equal partitioning in the horizontal or vertical direction, but is not limited to this.
[0137] If the block before division is M x N, and the width of the subblock is fixed (w), or the height is fixed (h), or both the width and height are fixed (w, h), then the number of subblocks obtained can be (M*N) / w, (M*N) / h, or (M*N) / w / h, respectively.
[0138] The encoding / decoding settings can determine whether only the CSI scheme is supported, only the VSI scheme, or both, and information about the supported schemes can be implicitly or explicitly determined.
[0139] This invention will be explained assuming the case where the CSI method is supported.
[0140] Depending on the encoding / decoding settings, the index partitioning can be configured to include two or more candidates, forming a group of candidates.
[0141] For example, candidate groups such as {a, b, c}, {a, b, c, n}, and {a~g, n} can be constructed. These candidate groups may include block configurations that are predicted to occur frequently based on general statistical characteristics, such as block configurations that are divided horizontally or vertically, or both horizontally and vertically.
[0142] Alternatively, candidate groups such as {a, b}, {a, o}, {a, b, o} or {a, c}, {a, p}, {a, c, p} can be constructed. Each candidate group includes candidates that are divided into 2 and 4 sections horizontally and vertically, respectively. A candidate group can be an example that includes a block shape in which divisions in a particular direction are expected to occur frequently.
[0143] Alternatively, a candidate group such as {a, o, p} or {a, n, q} can be constructed. The candidate group may be an example that includes block shapes in which many partitions smaller than the original block are predicted to occur.
[0144] Alternatively, a candidate group such as {a, r, s} can be constructed. This candidate group may include examples of non-rectangular partitioning configurations, based on the assumption that the optimal partitioning result can be obtained from the rectangle using another method (tree method) from the block before partitioning.
[0145] As shown in the example above, a variety of candidate group configurations are possible, and one or more candidate group configurations can be supported by considering a variety of encoding / decoding elements.
[0146] Once the candidate group configuration is complete, various partitioned information configurations become possible.
[0147] For example, index selection information can be generated from a group of candidates that includes candidates that are not divided (a) and candidates that are divided (b~s).
[0148] Alternatively, information indicating whether or not partitioning is possible (whether or not the partitioning form is a) can be generated, and if partitioning is to be performed (if it is not a), index selection information can be generated from a group of candidate partitions (b~s) that make up the candidate partition.
[0149] In addition to the above description, various other methods of constructing the partitioned information are possible. Aside from the information indicating whether partitioning is possible, binary bits can be assigned to the index of each candidate in the candidate group using various methods such as fixed-length binary code and variable-length binary code. For example, if there are two candidates in the candidate group, one bit can be assigned to the index selection information, and if there are three or more candidates, one or more bits can be assigned to the index selection information.
[0150] Unlike tree-based partitioning methods, index-based partitioning methods can selectively construct candidate groups based on partitioning patterns that are predicted to occur frequently.
[0151] Furthermore, since the number of bits required to represent the index information can be increased by the number of supported candidate groups, this method can be suitable for single-level partitioning (for example, partitioning depth limited to 0) rather than hierarchical partitioning (recursive partitioning) in tree-based methods. In other words, it can be a method that supports a single partitioning operation, and the subblocks obtained by the index-based partitioning cannot be further partitioned.
[0152] Here, we can mean a case where it is not possible to further partition into smaller blocks of the same type (for example, an encoded block obtained by an index partitioning method cannot be further partitioned into other encoded blocks), but it is also possible to set it so that further partitioning into other types of blocks is not possible (for example, partitioning from an encoded block into not only other encoded blocks but also prediction blocks is not possible). Of course, we are not limited to the above example, and other modifications are possible.
[0153] Next, we will explain the case where the block division settings are determined primarily by the type of block within the encoding / decoding elements.
[0154] First, an encoded block can be obtained through the partitioning process. Here, a tree-based partitioning method can be used for the partitioning process, and depending on the type of tree, partition configurations such as a (no split), n (QT), b, c (BT), i, and l (TT) can be obtained as shown in Figure 4. Various combinations of each tree type, such as QT / QT+BT / QT+BT+TT, are possible depending on the encoding / decoding settings.
[0155] The example described later illustrates the process of finally partitioning the prediction block and the transformation block based on the encoded block obtained through the above process, and assumes that the prediction, transformation, and inverse transformation processes are performed based on the size of each partitioned block.
[0156] In one example (1), the prediction process can be performed by setting the prediction block while keeping the same size as the encoded block, and the transformation and inverse transformation processes can be performed by setting the transformation block while keeping the same size as the encoded block (or prediction block). In the case of prediction blocks and transformation blocks, no separate partitioning information is generated because they are set based on the encoded block.
[0157] In one example (2), the prediction process can be performed by setting the prediction block while keeping the size of the encoded block. In the case of a transform block, the transform block can be obtained by a partitioning process based on the encoded block (or prediction block), and the transform and inverse transform processes can be performed based on the obtained size.
[0158] Here, the splitting process can use a tree-based splitting method, and depending on the type of tree, the resulting splitting configurations can be as shown in Figure 4: a (no split), b, c (BT), i, l (TT), n (QT), etc. Various combinations of each tree type, such as QT / BT / QT+BT / QT+BT+TT, are possible depending on the encoding / decoding settings.
[0159] Here, the partitioning process can use an index-based partitioning method, and depending on the index type, partition configurations such as a (no split), b, c, and d shown in Figure 4 can be obtained. Depending on the encoding / decoding settings, various candidate group configurations such as {a, b, c} and {a, b, c, d} are possible.
[0160] In example (3), in the case of a prediction block, the division process can be performed based on the encoded block to obtain the prediction block, and the prediction process can be performed based on the obtained size. In the case of a transformation block, the transformation and inverse transformation processes can be performed by setting the size to that of the encoded block. This example can be considered as a case where the prediction block and the transformation block have an independent relationship with each other.
[0161] Here, the partitioning process can use an index-based partitioning method, and depending on the index type, partitioning results such as a (no split), b~g, n, r, s shown in Figure 4 can be obtained. Depending on the encoding / decoding settings, various candidate group configurations such as {a, b, c, n}, {a~g, n}, and {a, r, s} are possible.
[0162] In example (4), in the case of a prediction block, the division process can be performed based on the encoded block to obtain a prediction block, and the prediction process can be performed based on the obtained size. In the case of a transformation block, the transformation and inverse transformation processes can be performed by setting it at the same size as the prediction block. In this example, the transformation block can be set at the same size as the obtained prediction block, or vice versa (the prediction block can be set at the same size as the transformation block).
[0163] Here, the partitioning process can use a tree-based partitioning method, and depending on the tree type, partitioning forms such as a (no split), b, c (BT), and n (QT) shown in Figure 4 can be obtained. Various combinations of each tree type, such as QT / BT / QT+BT, are possible depending on the encoding / decoding settings.
[0164] Here, the partitioning process can use the index-based partitioning method, and depending on the index type, partition configurations such as a (no split), b, c, n, o, and p shown in Figure 4 can be obtained. Depending on the encoding / decoding settings, a variety of candidate group configurations such as {a, b}, {a, c}, {a, n}, {a, o}, {a, p}, {a, b, c}, {a, o, p}, {a, b, c, n}, and {a, b, c, n, p} are possible. Furthermore, among the index-based partitioning methods, the VSI method can be used alone or in combination with the CSI method to construct the candidate group.
[0165] In example (5), in the case of a prediction block, the division process can be performed based on the encoded block to obtain the prediction block, and the prediction process can be performed based on the obtained size. Similarly, in the case of a transformation block, the division process can be performed based on the encoded block to obtain the prediction block, and the transformation and inverse transformation processes can be performed based on the obtained size. This example can be one in which the division of the prediction block and the transformation block is performed based on the encoded block.
[0166] Here, the partitioning process can use either a tree-based or index-based partitioning method, and the candidate group can be constructed identically or similarly to that in Example 4.
[0167] The above example illustrates some possible cases, such as sharing the block partitioning process for each type of block, but it is not limited to this, and various modifications are possible. Furthermore, the block partitioning settings can be determined by considering not only the type of block but also various encoding / decoding elements.
[0168] Here, the encoding / decoding elements may include video type (I / P / B), color components (YCbCr), block size / shape / position, block aspect ratio, block type (encoded block, prediction block, transformation block, quantization block, etc.), division state, encoding mode (Intra / Inter), prediction-related information (in-screen prediction mode, inter-screen prediction mode, etc.), transformation-related information (transformation technique selection information, etc.), and quantization-related information (quantization region selection information, quantized transformation coefficient encoding information, etc.).
[0169] In a video encoding method according to one embodiment of the present invention, the in-screen prediction can be configured as follows. The in-screen prediction of the prediction unit may include a reference pixel configuration step, a prediction block generation step, a prediction mode determination step, and a prediction mode encoding step. The video encoding device may also be configured to include a reference pixel configuration unit, a prediction block generation unit, and a prediction mode encoding unit that embody the reference pixel configuration step, the prediction block generation step, the prediction mode determination step, and the prediction mode encoding step. Some of the processes described above may be omitted or other processes may be added, and the order may be changed to one other than that described above.
[0170] Figure 5 is an illustrative diagram showing an in-screen prediction mode according to one embodiment of the present invention.
[0171] Referring to Figure 5, 67 prediction modes are configured as a group of candidate prediction modes for in-screen prediction. Of these, 65 are directional modes and 2 are non-directional modes (DC, Planar), as will be assumed in the explanation, but this is not limited to this configuration, and various configurations are possible. Here, directional modes can be distinguished by gradient (e.g., dy / dx) or angular information (Degree). Furthermore, all or part of the prediction modes may be included in the group of candidate prediction modes for luminance components or chrominance components, and other additional modes may be included in the group of candidate prediction modes.
[0172] In this invention, the direction of the directional mode can mean a straight line, and a curved directional mode can also be additionally configured as a prediction mode. Furthermore, in the case of a non-directional mode, it is possible to include a DC mode in which a prediction block is obtained from the average (or weighted average, etc.) of pixels of surrounding blocks adjacent to the current block (e.g., blocks to the left, above, upper left, upper right, lower left, etc.), and a Planar mode in which a prediction block is obtained by linear interpolation or the like using pixels of surrounding blocks.
[0173] In DC mode, the reference pixels used to generate prediction blocks can be obtained from blocks consisting of various combinations such as left, top, left + top, left + bottom left, top + top right, left + top + bottom left + top right, and the reference pixel acquisition block position can be determined by encoding / decoding settings defined by the image type, color components, block size / shape / position, etc.
[0174] In Planar mode, the pixels used to generate prediction blocks can be obtained from regions composed of reference pixels (e.g., left, top, upper left, upper right, lower left, etc.) and regions not composed of reference pixels (e.g., right, bottom, lower right, etc.). In the case of regions not composed of reference pixels (i.e., not encoded), they can be implicitly obtained by using one or more pixels from the region composed of reference pixels (e.g., direct copy, weighted average, etc.), or information about at least one pixel in the region not composed of reference pixels can be explicitly generated. Thus, prediction blocks can be generated using regions composed of reference pixels and regions not composed of reference pixels.
[0175] The invention may include additional non-directional modes other than those described above. While the present invention primarily describes linear directional modes and DC and Planar non-directional modes, modifications and applications to other cases are also possible.
[0176] Figure 5 can represent a prediction mode that is fixedly supported regardless of block size. Furthermore, a prediction mode supported by block size can differ from that shown in Figure 4.
[0177] For example, the number of candidate prediction modes may be adaptive (e.g., the angles between prediction modes are equally spaced, but the angles are set differently, such as 9, 17, 33, 65, or 129 based on the directional mode), or the number of candidate prediction modes may be fixed but consist of other configurations (e.g., directional mode angles, non-directional types, etc.).
[0178] Furthermore, Figure 5 can represent a prediction mode that is fixedly supported regardless of the block configuration. Also, the prediction mode supported by the block configuration can differ from that shown in Figure 4.
[0179] For example, the number of candidate prediction modes may be adaptive (e.g., the number of prediction modes derived horizontally or vertically based on the ratio of the width to height of the block may be set to be fewer or more), or the number of candidate prediction modes may be fixed but configured in other ways (e.g., the prediction modes derived horizontally or vertically based on the ratio of the width to height of the block may be set more finely).
[0180] Alternatively, the prediction mode on the longer side of the block can support more prediction modes, and the prediction mode on the shorter side can support fewer prediction modes. When the block is long, the prediction mode interval can also support modes located to the right of mode 66 in Figure 5 (for example, modes with an angle of +45 degrees or more relative to mode 50, i.e., modes numbered 67 to 80) or modes located to the left of mode 2 (for example, modes with an angle of -45 degrees or more relative to mode 18, i.e., modes numbered -1 to -14). This can be determined by the ratio of the width to the height of the block, and the opposite situation is also possible.
[0181] This invention primarily describes the case where the prediction mode is fixedly supported (regardless of any encoding / decoding element), as shown in Figure 5. However, it is also possible to set a prediction mode that is adaptively supported by the encoding settings.
[0182] Furthermore, when classifying prediction modes, horizontal and vertical modes (modes 18 and 50), and some diagonal modes (Diagonal up right <mode 2>, Diagonal down right <mode 34>, Diagonal down left <mode 66>, etc.) can be used as criteria, and this may be a classification method based on certain directions (or angles such as 45 degrees, 90 degrees, etc.).
[0183] Furthermore, some modes located at both ends of the directional mode (modes 2 and 66) can serve as the criterion modes for predictive mode classification, which is possible when the in-screen predictive mode configuration is as shown in Figure 5. In other words, if the predictive mode configuration is adaptive, it is also possible for the aforementioned criterion modes to be changed. For example, mode 2 may be replaced by modes with numbers less than or greater than 2 (-2, -1, 3, 4, etc.), or mode 66 may be replaced by modes with numbers less than or greater than 66 (64, 66, 67, 68, etc.).
[0184] Furthermore, additional prediction modes related to color components (color copy mode, color mode) can be included in the group of prediction mode candidates. Here, the color copy mode may be a prediction mode related to a method for acquiring data for generating prediction blocks from regions located in other color spaces, and the color mode may be a prediction mode related to a method for acquiring a prediction mode from regions located in other color spaces.
[0185] Figure 6 is an illustrative diagram illustrating the reference pixel configuration used for in-screen prediction according to one embodiment of the present invention. The size and shape (M×N) of the prediction block can be obtained by the block division section.
[0186] While on-screen prediction is generally performed in units of prediction blocks, it can also be performed in units of encoding blocks, transformation blocks, etc., depending on the block division settings. After checking the block information, the reference pixel constructor can construct the reference pixels to be used for predicting the current block. Here, the reference pixels are stored in temporary memory (for example, an array). <array>It can be managed by (primary, secondary, etc.) and is generated and removed for each in-screen prediction process of a block, and the size of the temporary memory can be determined by the configuration of the reference pixels.
[0187] In this example, we assume that the blocks to the left, top, upper left, upper right, and lower left of the current block are used to predict the current block, but we are not limited to this, and other configurations of block candidate groups can also be used to predict the current block. For example, the candidate group of adjacent blocks for the reference pixel can be an example of a raster or Z-scan, and depending on the scan order, some of the candidate group may be removed or may include other block candidate groups (e.g., the right, bottom, and lower right blocks may be added as additional configurations).
[0188] Furthermore, if certain prediction modes (such as color copy mode) are supported, a portion of another color space can be used to predict the current block, and this can also be considered as a reference pixel.
[0189] Figure 7 is a conceptual diagram showing blocks adjacent to the target block for in-screen prediction according to one embodiment of the present invention. More specifically, the left side of Figure 10 shows blocks adjacent to the current block in the current color space, and the right side shows corresponding blocks in other color spaces. For the sake of explanation, the following description will assume that the blocks adjacent to the current block in the current color space consist of basic reference pixels.
[0190] As shown in Figure 6, the reference pixels used to predict the current block can be composed of adjacent pixels in the left, top, upper left, upper right, and lower left blocks (Ref_L, Ref_T, Ref_TL, Ref_TR, Ref_BL in Figure 6). Here, the reference pixels are generally composed of pixels in the block closest to the current block (a in Figure 6, which is referred to as the reference pixel line), but other pixels (b in Figure 6 and the pixels of the other outer lines) can also be composed of reference pixels.
[0191] Pixels adjacent to the current block can be classified into at least one reference pixel line. The pixel closest to the current block is ref_0 {for example, a pixel whose distance from the current block's boundary pixel is 1: p(-1,-1)~p(2m-1,-1), p(-1,0)~p(-1,2n-1)}, the next adjacent pixel {for example, a pixel whose distance from the current block's boundary pixel is 2: p(-2,-2)~p(2m,-2), p(-2,-1)~p(-2,2n)} is ref_1, and the next adjacent pixel {for example, a pixel whose distance from the current block's boundary pixel is 3: p(-3,-3)~p(2m+1,-3), p(-3,-2)~p(-3,2n+1)} is ref_2, and so on. In other words, pixels can be classified into reference pixel lines based on the distance between the current block's boundary pixel and adjacent pixels.
[0192] Here, there can be N or more supported reference pixel lines, and N can be an integer greater than or equal to 1, such as 1 to 5. Here, it is common, but not limited to, that the reference pixel lines included in the candidate group are those closest to the current block, starting with the most adjacent reference pixel lines. For example, when N is 3,<ref_0、ref_1、ref_2> The candidate group can be constructed sequentially as described above, or<ref_0、ref_1、ref_3> ,<ref_0、ref_2、ref_3> ,<ref_1、ref_2、ref_3> The candidate group can also be configured in a way that excludes the most adjacent reference pixel lines, rather than sequentially.
[0193] The prediction can be performed using all of the reference pixel lines within the candidate group, or using some of the reference pixel lines (one or more).
[0194] For example, you can select one of several reference pixel lines through encoding / decoding settings and perform in-screen prediction using that reference pixel line. Alternatively, you can select two or more reference pixel lines from among several and perform in-screen prediction using those reference pixel lines (for example, by applying a weighted average to the data of each reference pixel line).
[0195] Here, the selection of the reference pixel line can be determined implicitly or explicitly. For example, implicit means that it is determined by the encoding / decoding settings defined by one or more combinations of elements such as the image type, color components, and block size / shape / position. Explicit means that reference pixel line selection information can occur in units such as blocks.
[0196] This invention primarily describes the case where in-screen prediction is performed using the most adjacent reference pixel line, but it should be understood that the various embodiments described later can be applied in the same or similar manner to cases where multiple reference pixel lines are used.
[0197] For example, the system can support a setting that implicitly determines the information when using the nearest adjacent reference pixel line for in-screen prediction at the subblock level, as described later. That is, in-screen prediction at the subblock level can be performed using a pre-configured reference pixel line, and the reference pixel line can be selected through implicit processing. Here, the reference pixel line can mean, but is not limited to, the nearest adjacent reference pixel line.
[0198] Alternatively, a reference pixel line can be adaptively selected for subblock-level in-screen prediction, and a variety of reference pixel lines including the most adjacent reference pixels can be selected to perform subblock-level in-screen prediction. In other words, a reference pixel line determined by considering various encoding / decoding elements can be used to perform subblock-level in-screen prediction, and the reference pixel line can be selected through implicit or explicit processing.
[0199] The reference pixel configuration for screen prediction of the present invention may include a reference pixel generation unit, a reference pixel interpolation unit, a reference pixel filter unit, etc., and may include all or part of the above configuration.
[0200] The reference pixel configuration unit can check the availability of reference pixels and classify them into usable and unusable reference pixels. Here, a reference pixel is determined to be unusable if it satisfies at least one of the following conditions.
[0201] For example, if any of the following conditions are met, the data can be deemed unusable: if it is located outside the picture boundary; if it does not belong to the same division unit as the current block (e.g., a slice, tile, etc., which are not mutually referential units; however, if a unit like a slice or tile has the characteristic of being able to refer to each other, an exception is made even if it is not the same division unit); or if encoding / decoding has not been completed. In other words, if any of the above conditions are not met, the data can be deemed usable.
[0202] Furthermore, the use of reference pixels can be restricted by encoding / decoding settings. For example, even if it is determined that use is permitted under the above conditions, the use of reference pixels can be restricted by whether or not restricted in-screen prediction (e.g., constrained_intra_pred_flag) can be performed. Restricted in-screen prediction can be implemented when attempting to perform error-robust encoding / decoding for external parties such as communication environments, and it is necessary to prohibit the use of blocks that have been referenced and restored from other video as reference pixels.
[0203] If restricted in-screen prediction is disabled (for example, for I-image type, or P or B-image type with constrained_intra_pred_flag=0), all candidate reference pixel blocks can be made available.
[0204] Alternatively, if restricted in-screen prediction is activated (for example, constrained_intra_pred_flag=1 for P or B video types), the availability of a reference pixel candidate block is assumed to be determined by the encoding mode (Intra or Inter), but the condition can also be determined by a variety of other encoding / decoding elements.
[0205] Since a reference pixel is composed of one or more blocks, after checking the possibility of the reference pixel, it can be classified into three cases: <all usable>, <partially usable>, and <all unusable>. In the cases other than the all-usable case, the reference pixels at the candidate block positions of the unusable category can be satisfied or generated.
[0206] If a candidate block of reference pixels is available, the pixel at that location can be included in the reference pixel memory of the current block. Here, the pixel data can be copied directly, or it can be included in the reference pixel memory through processes such as reference pixel filtering or reference pixel interpolation. Furthermore, if a candidate block of reference pixels is unavailable, pixels acquired through the reference pixel generation process can be included in the reference pixel memory of the current block.
[0207] The following examples demonstrate how to generate reference pixels for unusable block locations using various methods.
[0208] For example, a reference pixel can be generated using any pixel value. Here, any pixel value can be a single pixel value belonging to a pixel value range (for example, a pixel value range based on bit depth or a pixel value range based on the pixel distribution in the relevant image) (e.g., the minimum, maximum, or median of the pixel value range). More specifically, this can be an example applicable when all candidate blocks of reference pixels are unavailable.
[0209] Alternatively, a reference pixel can be generated from a region where the encoding / decoding of the video is complete. More specifically, a reference pixel can be generated from at least one usable block adjacent to an unusable block. Here, at least one of the methods such as extrapolation, interpolation, or copying can be used.
[0210] After the reference pixel interpolation unit completes the configuration of the reference pixels, a small number of reference pixels can be generated by linear interpolation of the reference pixels. Alternatively, the reference pixel interpolation process can be performed after the reference pixel filtering process described later.
[0211] Here, interpolation is not performed in the case of horizontal, vertical, some diagonal modes (for example, modes with a 45-degree difference between vertical and horizontal, such as Diagonal up right, Diagonal down right, and Diagonal down left; corresponding to modes 2, 34, and 66 in Figure 5), non-directional modes, and color copy mode, but it can be performed in other modes (other diagonal modes).
[0212] The pixel positions to be interpolated (i.e., which decimal units are interpolated, determined from 1 / 2 to 1 / 64, etc.) can be determined by the prediction mode (e.g., the direction of the prediction mode, such as dy / dx) and the positions of the reference and predicted pixels. Here, regardless of the precision of the decimal units, one filter can be applied (for example, assuming the same filter for the mathematical formula used to determine the filter coefficients or the length of the filter taps, but with the coefficients adjusted only by the precision of the decimal units <e.g., 1 / 32, 7 / 32, 19 / 32>), or one filter can be selected and applied from among multiple filters (for example, assuming filters for which the mathematical formula used to determine the filter coefficients or the length of the filter taps are separated) based on the decimal units.
[0213] In the former case, it may be an example where integer-unit pixels are used as input for interpolation of decimal-unit pixels, and in the latter case, it may be an example where different input pixels are used at each stage (for example, integer pixels are used for 1 / 2 units, and integer and 1 / 2 unit pixels are used for 1 / 4 units, etc.), but it is not limited to these. This invention will mainly describe the former case.
[0214] For reference pixel interpolation, either fixed or adaptive filtering can be performed, which can be determined by the encoding / decoding settings (e.g., one or more combinations of image type, color components, block position / size / shape, block width-to-height ratio, prediction mode, etc.).
[0215] Fixed filtering can perform reference pixel interpolation using a single filter, while adaptive filtering can perform reference pixel interpolation using one of several filters.
[0216] Here, in the case of adaptive filtering, the encoding / decoding settings can implicitly or explicitly determine which of several filters to select. The filter types can consist of 4-tap DCT-IF filters, 4-tap cubic filters, 4-tap Gaussian filters, 6-tap Wiener filters, 8-tap Kalman filters, etc., and the sets of filter candidates supported by the color components can be defined differently (for example, some filter types may be the same or different, or the filter tap lengths may be short or long).
[0217] The reference pixel filter section can perform filtering on the reference pixels to improve prediction accuracy by reducing the degradation remaining after the encoding / decoding process. The filter used at this time can be, but is not limited to, a low-pass filter. Whether or not to apply filtering can be determined by the encoding / decoding settings (which can be deduced from the explanation above). Furthermore, when filtering is applied, either fixed filtering or adaptive filtering can be applied.
[0218] Fixed filtering means that reference pixel filtering is not performed or that reference pixel filtering is applied using one filter. Adaptive filtering means that the applicability of filtering is determined by the encoding / decoding settings, and if there are two or more types of supported filters, one of them can be selected.
[0219] Here, the filter type can support multiple filters, distinguished by various filter coefficients such as [1, 2, 1] / 4 (3-tap filter) and [2, 3, 6, 3, 2] / 16 (5-tap filter), as well as filter tap lengths.
[0220] The reference pixel interpolation unit and reference pixel filter unit introduced in the reference pixel configuration stage can be configurations necessary to improve the accuracy of predictions. The two processes can be performed independently, but a configuration that combines the two processes (i.e., processing with a single filter) is also possible.
[0221] The prediction block generation unit can generate prediction blocks using at least one prediction mode, and can use reference pixels based on the prediction mode. Here, the reference pixels can be used for methods such as extrapolation (directional mode) or for methods such as interpolation, averaging (DC), or copying (non-directional mode), depending on the prediction mode.
[0222] The prediction mode determination unit performs a process to select the optimal mode from a group of multiple prediction mode candidates. Generally, the optimal mode in terms of encoding cost can be determined using a rate-distortion technique that considers block distortion {e.g., distortion between the current block and the restored block, such as SAD (Sum of Absolute Difference) or SSD (Sum of Square Difference)} and the amount of bits generated by the relevant mode. The prediction blocks generated based on the prediction mode determined by the above process can be transmitted to the subtraction and addition units.
[0223] To determine the optimal prediction mode, all prediction modes in the candidate group can be explored, or the optimal prediction mode can be selected through other decision processes for the purpose of reducing computational complexity. For example, in the first stage, some modes that show good performance in terms of image quality degradation can be selected from the entire set of on-screen prediction mode candidates, and in the second stage, the optimal prediction mode can be selected from the modes selected in the first stage, taking into account not only image quality degradation but also the amount of bits generated. In addition to the above method, various methods for reducing computational complexity can be applied.
[0224] Furthermore, while the prediction mode determination unit can generally be configured to be included only in the encoder, it can also be included in the decoder depending on the encoding / decoding settings. For example, if the prediction method includes template matching or a method of inducing the in-screen prediction mode in an adjacent region of the current block, in the latter case, it can be understood that a method of implicitly acquiring the prediction mode in the decoder is used.
[0225] The prediction mode coding unit can encode the prediction mode selected by the prediction mode determination unit. It can encode index information corresponding to the prediction mode from the candidate prediction mode group, or it can predict the prediction mode and encode information about it. The former method may be applied to the luminance component, and the latter method may be applied to the chrominance component, but is not limited to these two cases.
[0226] When predicting and encoding a prediction mode, the predicted value (or prediction information) of the prediction mode can be called the Most Probable Mode (MPM). An MPM can consist of one prediction mode or multiple prediction modes, and the number of MPMs (k, where k is an integer greater than or equal to 1, such as 1, 2, 3, or 6) can be determined by the number of candidate prediction mode groups. When an MPM consists of multiple prediction modes, it can be called an MPM candidate group.
[0227] The MPM candidate set can be supported under fixed settings or adaptively through a variety of encoding / decoding elements. Examples of adaptive settings include determining the candidate set configuration by which of multiple reference pixel hierarchies to use, or determining the candidate set configuration by whether in-screen prediction is performed at the block level or the sub-block level. For the sake of explanation, we will assume the configuration of the MPM candidate set under a single setting, and it should be understood that candidate set configurations for other in-screen prediction modes, not just the MPM candidate set, can also be adaptive.
[0228] MPM is a concept that helps to efficiently encode prediction modes, and in practice, it can construct a group of candidate modes with high probability of occurrence in the current block's prediction mode.
[0229] For example, the MPM candidate group can consist of pre-defined prediction modes (or statistically frequently occurring prediction modes such as DC, Plaanr, vertical, horizontal, and some diagonal modes), prediction modes of adjacent blocks (such as left, upper, upper-left, upper-right, and lower-left blocks), etc. Here, the prediction modes of adjacent blocks can be obtained from L0-L3 (left block), T0-T3 (upper block), TL (upper-left block), R0-R3 (upper-right block), and B0-B3 (lower-left block) in Figure 10.
[0230] If a group of MPM candidates can be formed from two or more subblock positions (e.g., L0, L2, etc.) in an adjacent block (e.g., the left block), the prediction modes of the block can be included in the candidate group according to a predefined priority order (e.g., L0-L1-L2, etc.). Alternatively, if a group of MPM candidates cannot be formed from two or more subblock positions, the prediction modes of the subblocks corresponding to predefined positions (e.g., L0, etc.) can be included in the candidate group. In detail, the prediction modes of positions L3, T3, TL, R0, and B0 within an adjacent block can be selected as the prediction modes of the adjacent block and included in the MPM candidate group. The above description is only one of the cases in which the prediction modes of adjacent blocks are included in the candidate group, and is not limited thereto. The example described later assumes a case in which the prediction modes of predefined positions are included in the candidate group.
[0231] Furthermore, if one or more prediction modes are included in the MPM candidate group, modes derived from the one or more prediction modes already included can also be added to the MPM candidate group. Specifically, if the k-th mode (directional mode) is included in the MPM candidate group, modes that can be derived from that mode (modes with intervals of +a and -b relative to k, where a and b are integers of 1 or greater such as 1, 2, and 3) can be added to the MPM candidate group.
[0232] Priorities can exist for constructing the MPM candidate group, and the MPM candidate group can be constructed in the order of adjacent block prediction modes - pre-configured prediction modes - induced prediction modes, etc. The process of constructing the MPM candidate group can be completed by satisfying the maximum number of MPM candidates according to the priority. If a prediction mode already included in the process matches, the relevant prediction mode is not constructed in the candidate group, and the procedure can be included in a redundancy check process to move on to the next priority candidate.
[0233] Next, we will assume that the MPM candidate group consists of six prediction modes.
[0234] For example, candidate modes can be configured in the order of LT-TL-TR-BL-Planar-DC-Vertical-Horizontal-Diagonal modes. This can be done when you want to prioritize configuring the prediction modes of adjacent blocks in the candidate group and then additionally configure already set prediction modes.
[0235] Alternatively, LT-Planar-DC-<L+1> - <l-1> -<T+1>- <t-1>Candidate modes can be configured in the order of vertical, horizontal, diagonal, etc. This may involve prioritizing the configuration of some adjacent block prediction modes and some of the pre-configured prediction modes, and additionally configuring modes that are induced under the assumption that prediction modes in directions similar to the prediction modes of adjacent blocks occur, as well as some of the pre-configured prediction modes.
[0236] The above example is only one case regarding the composition of the MPM candidate group, and is not limited to this; various modifications are possible.
[0237] The MPM candidate group can use binary representations such as unary binarization and truncated rice binarization based on the index within the candidate group. That is, shorter bits can be assigned to candidates with smaller indices, and longer bits can be assigned to candidates with larger indices to represent the mode bits.
[0238] Modes that could not be included in the MPM candidate group can be classified as non-MPM candidate groups. Furthermore, depending on the encoding / decoding settings, non-MPM candidate groups can be further classified into two or more candidate groups.
[0239] Next, we assume that the prediction mode candidate group consists of 67 modes, including directional and non-directional modes, and that the MPM candidate supports 6 modes, resulting in a non-MPM candidate group of 61 prediction modes.
[0240] If the non-MPM candidate group consists of only one prediction mode, it represents a mode that could not be included in the MPM candidate group construction process, and therefore no additional candidate group construction process is required. Thus, based on the index within the non-MPM candidate group, binary methods such as Fixed Length Binarization and Truncated Unary Binarization can be used.
[0241] Assuming that a non-MPM candidate group consists of two or more candidate groups, in this example, the non-MPM candidate group is classified into non-MPM_A (hereinafter referred to as candidate group A) and non-MPM_B (hereinafter referred to as candidate group B). Assume that candidate group A (p members, equal to or greater than the number of MPM candidate groups) is composed of candidate modes that are more likely to occur in the prediction modes of the current block than candidate group B (q members, equal to or greater than the number of candidate groups A). Here, the process for constructing candidate group A can be added.
[0242] For example, some prediction modes with equal intervals within the directional modes (e.g., modes 2, 4, and 6) can be configured in candidate group A, or pre-configured prediction modes (e.g., modes derived from prediction modes included in the MPM candidate group) can be configured. The prediction modes remaining from the MPM candidate group configuration and candidate group A configuration can be configured in candidate group B, and no additional candidate group configuration process is required. Based on the indices within candidate group A and candidate group B, binary evolutions such as fixed-length binary evolution and truncated monometric binary evolution can be used.
[0243] The above example is a partial case where the non-MPM candidate group consists of two or more candidates, but is not limited to this, and various modifications are possible.
[0244] Next, we will show the process for predicting and encoding the prediction mode.
[0245] You can now check information (mpm_flag) about whether the current block's prediction mode matches the MPM (or some of the modes within the MPM candidate group).
[0246] If it matches the MPM, the MPM index information (mpm_idx) can be further verified by the MPM configuration (one or more). Then, the encoding process for the current block is completed.
[0247] If it does not match the MPM, and the non-MPM candidate group consists of only one, the non-MPM index information (remaining_idx) can be checked. Then, the encoding process for the current block is completed.
[0248] If there are multiple non-MPM candidate groups (two in this example), then information (non_mpm_flag) can be checked to see if the prediction mode of the current block matches some of the prediction modes within candidate group A.
[0249] If it matches candidate group A, the candidate group A index information (non_mpm_A_idx) can be checked; if it does not match candidate group A, the candidate group B index information (remaining_idx) can be checked. After that, the encoding process for the current block is completed.
[0250] If the candidate prediction mode group configuration is fixed, the prediction mode supported by the current block, the prediction mode supported by the adjacent block, and the previously set prediction mode can use the same prediction number index.
[0251] On the other hand, if the candidate prediction mode configuration is adaptive, the prediction mode supported by the current block, the prediction mode supported by the adjacent block, and the previously configured prediction mode can use the same prediction number index or different prediction number indexes. Refer to Figure 4 for the following explanation.
[0252] The prediction mode coding process can perform a process of unifying (or adjusting) the prediction mode candidate group for the construction of an MPM candidate group, etc. For example, the prediction mode of the current block may be one of the prediction mode candidate group from mode -5 to 61, and the prediction mode of the adjacent block may be one of the prediction mode candidate group from mode 2 to 66. In this case, a part of the prediction mode of the adjacent block (mode 66) may be a mode not supported by the prediction mode of the current block, so a process of unifying it can be performed during the prediction mode coding process. That is, this process may not be required when supporting a fixed in-screen prediction mode candidate group configuration, but it may be required when supporting an adaptive in-screen prediction mode candidate group configuration, and a detailed explanation of this is omitted.
[0253] Unlike the method using the aforementioned MPM, this method can perform coding by assigning an index to the prediction mode belonging to the group of prediction mode candidates.
[0254] For example, this could involve assigning an index to the prediction mode according to a predefined priority order, and then encoding the corresponding index once the prediction mode for the current block is selected. This implies a case where a fixed set of prediction mode candidates is constructed, and a fixed index is assigned to each prediction mode.
[0255] Alternatively, the aforementioned fixed index assignment method may not be suitable if the group of prediction mode candidates is adaptively configured. For this reason, a method can be applied in which an index is assigned to the prediction mode according to an adaptive priority, and when the prediction mode for the current block is selected, the corresponding index is encoded. This allows for efficient encoding of prediction modes by assigning different indices to the prediction modes depending on the adaptive configuration of the group of prediction mode candidates. In other words, the adaptive priority is used to assign candidates that are highly likely to be selected as the prediction mode for the current block to indices where short mode bits occur.
[0256] Next, we will assume a scenario where the prediction mode candidate group supports eight prediction modes, including the already set prediction modes (directional mode and non-directional mode), color copy mode, and color mode (color difference component).
[0257] For example, let's assume that the pre-configured prediction modes support four of the following: Planar, DC, horizontal, vertical, and diagonal modes (Diagonal down left in this example), and that it also supports one color mode (C) and three color copy modes (CP1, CP2, CP3). The basic order of the index assigned to the prediction mode can be given as a pre-configured order such as prediction mode - color copy mode - color mode.
[0258] Here, the pre-configured prediction modes—directional mode, non-directional mode, and color copy mode—can be easily categorized into prediction modes based on their prediction methods. However, in the case of color mode, it can be either directional mode or non-directional mode, which may overlap with pre-configured prediction modes. For example, if the color mode is vertical mode, it may overlap with one of the pre-configured prediction modes, which is vertical mode.
[0259] When the number of prediction mode candidate groups is adaptively adjusted based on the encoding / decoding settings, if the aforementioned overlap occurs, the number of candidate groups can be adjusted (e.g., 8 to 7). Alternatively, when the number of prediction mode candidate groups is fixed, if the aforementioned overlap occurs, other candidates can be added and considered when assigning the index; this setting will be assumed and described later. Furthermore, the adaptive prediction mode candidate group can also be configured to support variable modes such as color modes. Therefore, performing adaptive index assignment can be considered an example of an adaptive prediction mode candidate group configuration.
[0260] Next, we will explain the case where adaptive index assignment is performed based on the color mode. We assume that the basic index is assigned in the order of Planar(0)-Vertical(1)-Horizontal(2)-DC(3)-CP1(4)-CP2(5)-CP3(6)-C(7). Furthermore, we assume that if the color mode does not match the previously set prediction mode, the index assignment will be performed in the above order.
[0261] For example, if the color mode matches one of the pre-configured prediction modes (Planar, vertical, horizontal, DC mode), the prediction mode that matches the color mode index (7) will be satisfied. The index of the matching prediction mode (one of 0-3) will be satisfied with the pre-configured prediction mode (Diagoanal down left). In detail, if the color mode is horizontal mode, an index assignment such as Planar(0)-vertical(1)-Diagoanal down left(2)-DC(3)-CP1(4)-CP2(5)-CP3(6)-horizontal(7) can be performed.
[0262] Alternatively, if the color mode matches one of the pre-configured prediction modes, the prediction mode matching index 0 is satisfied. Then, the pre-configured prediction mode (Diagonal down left) is satisfied for index (7) of the color mode. Here, if the satisfied prediction mode is not the existing index 0 (i.e., not Planar mode), the existing index configuration can be adjusted. In detail, if the color mode is DC mode, an index assignment such as DC(0)-Planar(1)-Vertical(2)-Horizontal(3)-CP1(4)-CP2(5)-CP3(6)-Diagonal down left(7) can be performed.
[0263] The above examples represent only some cases of adaptive index assignment, and are not limited to them; various modifications are possible. Furthermore, based on the indexes within the candidate group, binary representations such as fixed-length binary, monomyal binary, truncated monomyal binary, and truncated Rice binary can be used.
[0264] Next, we will describe another example of assigning an index to a prediction mode belonging to a group of candidate prediction modes and performing coding.
[0265] For example, one such method involves classifying prediction modes into multiple candidate groups based on prediction modes, prediction methods, etc., assigning an index to the prediction modes belonging to the relevant candidate group, and then encoding them. In this case, candidate group selection information encoding can be performed prior to the index encoding. As an example, directional mode, non-directional mode, and color mode, which are prediction modes that perform predictions in the same color space, can belong to one candidate group (hereinafter referred to as S candidate group), and color copy mode, which is a prediction mode that performs predictions in a different color space, can belong to one candidate group (hereinafter referred to as D candidate group).
[0266] Next, we will assume a scenario where the prediction mode candidate group supports nine prediction modes, including the already set prediction mode, color copy mode, and color mode (color difference component).
[0267] For example, let's assume that the pre-configured prediction modes support four of the following: Planar, DC, horizontal, vertical, and diagonal modes, and one color mode (C) and four color copy modes (CP1, CP2, CP3, CP4). Candidate group S can have five candidates, each consisting of the pre-configured prediction modes and color modes, and candidate group D can have four candidates, each consisting of the color copy modes.
[0268] The S candidate group is an example of an adaptively constructed prediction mode candidate group, and since examples of adaptive index assignment have been mentioned above, a detailed explanation will be omitted. The D candidate group is an example of a fixedly constructed prediction mode candidate group, so a fixed index assignment method can be used. For example, an index assignment such as CP1(0)-CP2(1)-CP3(2)-CP4(3) can be performed.
[0269] Based on the index within the candidate group, binary evolutions such as fixed-length binary, monomyal binary, truncated monomyal binary, and truncated rice binary can be used. Furthermore, the examples are not limited to those described above, and various modifications are possible.
[0270] Candidate group construction for predictive mode coding, such as MPM, can be performed in block units. Alternatively, the candidate group construction process can be omitted, and a predetermined candidate group or a candidate group obtained by various methods can be used. This configuration can be supported for purposes such as reducing complexity.
[0271] In one example (1), you can use a predefined candidate group or one of several predefined candidate groups depending on the encoding / decoding settings. For example, in the case of an MPM candidate group, you can use a predefined candidate group such as {Planar-DC-Vertical-Horizontal-Diagonal down left<No. 66 in Figure 5>-Diagonal down right<No. 34 in Figure 5>}. Alternatively, you can apply to this example a candidate group that is formed when all adjacent blocks in the MPM candidate group configuration are unavailable.
[0272] In example (2), a group of candidate blocks whose encoding is complete can be used. Here, the encoded blocks can be selected based on the encoding order (a predetermined scanning method, e.g., z-scan, vertical scan, horizontal scan, etc.) or from adjacent blocks to the left, above, upper left, upper right, lower left, etc., of the current block. However, adjacent blocks can be limited to positions belonging to division units that can reference each other with the current block (for example, when each block has attributes that can reference each other even if they belong to different slices or tiles; for example, when they belong to the same tile group but are on different tiles), and if they belong to division units that cannot reference each other (for example, when each block belongs to different slices or tiles but has attributes that cannot reference each other; for example, when they belong to different tile groups), the block at that position can be excluded from the candidates.
[0273] Here, adjacent blocks can be determined by the current state of the block. For example, if the current block is a square, it can borrow (or share) a group of available candidate blocks from the block located in the area according to a predetermined first priority. Alternatively, if the current block is a rectangle, it can borrow a group of available candidate blocks from the block located in the area according to a second priority. Here, the ratio of the width to the height of the block can support the second or third priority. The priority order for selecting candidate blocks to borrow can have various configurations, such as left-top-upper-right-upper-left-upper, or top-left-upper-right-upper-left. Here, the first to third priorities may all have the same configuration, all have different configurations, or some of the configurations may be the same.
[0274] Currently, a group of candidate blocks can borrow from an adjacent block only if it is greater than or equal to a predetermined boundary value. Alternatively, it can borrow from an adjacent block only if it is less than or equal to a predetermined boundary value. Here, the boundary value can be defined as the minimum or maximum size of a block that is allowed to borrow from the candidate group. The boundary value can be expressed as the width (W), height (H), W×H, W*H, etc., of the block, and W and H can be integers of 4, 8, 16, 32, or more.
[0275] In one example (3), a common group of candidates can be formed by a higher-level block composed of a predetermined bundle of blocks, and lower-level blocks belonging to the higher-level block can use this group of candidates. Here, the number of lower-level blocks can be an integer of 1 or more, such as 1, 2, 3, or 4.
[0276] Here, the upper block can be the ancestor block (including the parent block) of the lower block or a block composed of a bundle of any blocks. Here, the ancestor block can mean the block before the previous stage (the difference in the division depth is 1 or more) in the division process for obtaining the lower block. As an example, the parent blocks of the 0th and 1st sub-blocks of 4N×2N in b of FIG. 4 can be shown as 4N×4N in a of FIG. 4.
[0277] The candidate group of the upper block can perform borrowing (or sharing) from the lower block only when it is above / beyond a predetermined first boundary value. Or, it can perform borrowing from the lower block only when it is below / less than a predetermined second boundary value.
[0278] Here, the boundary value can be defined as the minimum size or the maximum size of the block that allows candidate group borrowing. Only one of the boundary values can be supported or both can be supported. The boundary value can be expressed as the horizontal length (W), vertical length (H), W×H, W*H, etc. of the block, and the W and H can be integers such as 8, 16, 32, 64 or more.
[0279] On the other hand, the candidate group of the lower block can perform borrowing from the upper block only when it is above / beyond a predetermined third boundary value. Or, it can perform borrowing from the upper block only when it is below / less than a predetermined fourth boundary value.
[0280] Here, the boundary value can be defined as the minimum size or the maximum size of the block that allows candidate group borrowing. Only one of the boundary values can be supported or both can be supported. The boundary value can be expressed as the horizontal length (W), vertical length (H), W×H, W*H, etc. of the block, and the W and H can be integers such as 4, 8, 16, 32 or more.
[0281] Here, the first boundary value (or the second boundary value) can be greater than or the same as the third boundary value (or the fourth boundary value).
[0282] Candidate group borrowing (or sharing) can be selectively used based on any one of the foregoing embodiments, and candidate group borrowing can also be selectively used based on at least two combinations of Embodiments 1 to 3. Also, candidate group borrowing can be selectively used based on any one of the detailed configurations of each embodiment, and can be selectively used by combining one or more detailed configurations, etc.
[0283] Also, information such as whether candidate group borrowing is possible, the attributes (size / shape / position / ratio of horizontal length to vertical length) of the blocks involved in candidate group borrowing, and the division state (division method, division type, division depth, etc.) can be explicitly processed. Also, encoding elements such as video type and color components can act as input variables for candidate group borrowing settings. Candidate group borrowing can be performed based on the information and encoding / decoding settings.
[0284] The prediction-related information generated by the prediction mode encoding unit can be transmitted to the encoding unit and recorded in the bitstream.
[0285] Intra prediction in the video decoding method according to an embodiment of the present invention can be configured as follows. Intra prediction of the prediction unit can include a prediction mode decoding stage, a reference pixel configuration stage, and a prediction block generation stage. Also, the video decoding apparatus can be configured to include a prediction mode decoding unit, a reference pixel configuration unit, and a prediction block generation unit that implement the prediction mode decoding stage, the reference pixel configuration stage, and the prediction block generation stage. A part of the foregoing process can be omitted or other processes can be added, and it can be changed to other orders that are not the described order.
[0286] The reference pixel configuration unit and the prediction block generation unit of the video decoding apparatus perform the same role as the corresponding configurations of the video encoding apparatus, so detailed description is omitted, and the prediction mode decoding unit can be performed by using the method used in the prediction mode encoding unit in reverse.
[0287] (Intra prediction in sub-block units)
[0288] Figure 8 shows various sub-block divisions that can be obtained based on the encoded block. Here, the encoded block is called the parent block, and the sub-blocks can be child blocks. Here, the sub-blocks can be units in which prediction is performed or units in which transformation is performed. In this example, we assume that one prediction information is shared among the sub-blocks. That is, one prediction mode is generated and used in each sub-block.
[0289] Referring to Figure 8, the coding order of the subblocks can be determined by various combinations of the sequences a to p in Figure 8. For example, it can be one of the following sequences: z-scan (left to right, top to bottom), vertical scan (top to bottom), horizontal scan (left to right), inverse vertical scan (bottom to top), inverse horizontal scan (right to left).
[0290] The encoding order may be an order already promised to the video unit / decoder. Alternatively, the encoding order of the subblocks may be determined by considering the division direction of the parent block. For example, if the parent block is divided horizontally, the encoding order of the subblocks may be determined to be a vertical scan. If the parent block is divided vertically, the encoding order of the subblocks may be determined to be a horizontal scan.
[0291] When encoding is performed on a subblock basis, the reference data used for prediction can be obtained at a closer location, and since only one prediction mode is generated and shared among the subblocks, it can be efficient.
[0292] For example, referring to Figure 8(b), when encoding at the parent block level, the lower right subblock can perform predictions using pixels adjacent to the parent block. On the other hand, when encoding at the subblock level, the lower right subblock has upper left, upper right, and lower left subblocks that have been restored by a predetermined encoding order (z-scan in this example), so it can perform predictions using pixels adjacent to the parent block.
[0293] For in-screen prediction through optimal division considering video characteristics, a group of candidate divisions can be constructed based on one or more division patterns that are highly likely to occur.
[0294] In this example, we assume that the partitioning information is generated using the index-based partitioning method. A variety of partitioning configurations can be constructed as candidate groups, as follows:
[0295] Specifically, a candidate group can be constructed consisting of N partition configurations, where N can be an integer greater than or equal to 2. This candidate group can include at least two combinations of the seven partition configurations shown in Figure 8.
[0296] A parent block can be divided into predetermined subblocks using one of a plurality of division patterns belonging to the candidate group. The selection can be performed based on an index signaled by the video encoding device. The index can represent information that identifies the division pattern of the parent block. Alternatively, the selection can be performed by the video decoding device considering the attributes of the parent block. Here, the attributes can represent the block's position, size, shape, width, aspect ratio, length of either the width or height, division depth, video type (I / P / B), color components (e.g., luminance, chrominance), in-screen prediction mode value, whether the in-screen prediction mode is non-directional, the angle of the in-screen prediction mode, the position of a reference pixel, etc. The block can represent an encoded block or a prediction block and / or transform block corresponding to the encoded block. The block's position can represent whether it is located at the boundary of a predetermined video (or fragment video) of the parent block. Here, "image" (or "fragmented image") can mean at least one of the following: picture, slice group, tile group, slice, tile, CTU row, or CTU, to which the parent block belongs.
[0297] In one example (1), candidate groups such as {a~d} and {a~g} in Figure 8 can be constructed, and this can be a candidate group construction that takes into account various partitioning forms. Assuming the case of {a~d}, various binary bits can be assigned to each index (assuming the indices are assigned in alphabetical order; the same assumption is made in the next example).
[0298] [Table 1] In Table 1 above, bin type 1 can be an example of binary code that considers all possible partitioning configurations, and bin type 2 can be an example of binary code that first assigns a bit (first bit) indicating whether partitioning is possible, and if partitioning is possible (first bit is 1), it performs the operation by excluding only the candidates that were not partitioned from among the possible partitioning configurations.
[0299] In Example (2), candidate groups such as {a, c, d} and {a, f, g} in FIG. 8 can be formed, which can be candidate group formations considering division in a specific direction (horizontal or vertical direction). Assuming the case of {a, f, g}, various binary bits can be assigned to each index.
[0300]
Table 2
[0301] In the case of bin type 2, when the parent block is a horizontally long rectangle, 1 bit is assigned when it is divided horizontally, and the rest are 2 bits. Bin type 3, when the parent block is a vertically long rectangle, 1 bit is assigned when it is divided vertically, and the rest are 2 bits. This can be an example of assigning shorter bits when it is determined that further division is likely to occur, such as the form of the parent block, but is not limited to this and can include modified examples including the opposite case.
[0302] In Example (3), candidate groups such as {a, c, d, f, g} in FIG. 8 can be formed, which can be other examples of candidate group formations considering division in a specific direction.
[0303]
Table 3
[0304] The aforementioned partitioning information can be configured to support general situations, but it can be transformed into other configurations depending on the encoding / decoding environment, etc. That is, it can support exceptional configurations of the partitioning information or substitute the partitioning form represented by the partitioning information with other partitioning forms.
[0305] For example, there may be partitioning configurations that are not supported by block attributes. Here, block attributes were mentioned in a previous example, so a detailed explanation is omitted. Here, a block can mean at least one parent block or subblock.
[0306] For example, let's assume the supported division patterns are {a, f, g} in Figure 8, and the size of the parent block is 4M × 4N. If some division patterns are not supported due to the minimum size condition of the blocks in the video, or their position at the boundary of a predetermined video (or fragment video) of the parent block, then various processing options are possible. For the following example, let's assume the minimum width of the blocks in the video is 2M, and the minimum width of the blocks is 4*M*N.
[0307] First, the candidate set can be reconstructed by removing the unobtainable partitioning patterns. The existing candidate set can have 4M×4N, 4M×N, and M×4N, and the reconstructed candidate set, excluding the unobtainable candidate (4M×N), can be 4M×4N and 4M×N. In this case, binary evolution can be performed again on the candidates within the reconstructed candidate set. In this example, 4M×4N and 4M×N can be selected as a single flag (1 bit).
[0308] In another example, the candidate group can be reconstructed with alternative candidates for unobtainable partitioning configurations. In the above example, the unobtainable partitioning configuration could be a vertical partitioning configuration (4 partitions). However, the candidate group can be reconstructed by substituting other partitioning configurations that maintain the vertical partitioning configuration (e.g., 2M × 4N). This allows the existing flag configuration based on the partitioning information to be maintained.
[0309] As in the example above, the candidate group can be restructured by adjusting the number of candidates or by replacing existing candidates, and the example above only describes some cases, so various modifications are possible.
[0310] (Encoding order of subblocks)
[0311] As shown in Figure 8, various subblock coding orders can be set. The coding order can be implicitly determined by the coding / decoding settings. The coding / decoding elements can include the division form, video type, color components, size / form / position of the parent block, the ratio of the width to height of the block, prediction mode related information (e.g., in-screen prediction mode, reference pixel position used, etc.), and division state.
[0312] Alternatively, explicit processing of the coding order of subblocks is possible. Specifically, a candidate group can be constructed from candidates that are likely to occur depending on the partitioning configuration, and selection information can be generated for one of them. Therefore, the coding order candidates supported by the partitioning configuration can be constructed adaptively.
[0313] As in the example above, it is possible to perform the coding of subblocks using a single fixed coding order, but it is also possible to apply an adaptive coding order.
[0314] Referring to Figure 8, many coding sequences can be obtained through the various sequence assignments a to p shown in the diagram. Since the position and number of subblocks obtained can differ depending on the partitioning configuration, it is important to configure coding sequences that are specific to the partitioning configuration. Furthermore, for partitioning configurations like Figure 8(a) where partitioning into subblocks is not performed, no processing of coding sequences is necessary. Therefore, when explicitly processing information about coding sequences, it is possible to first check the partitioning information and then generate information about the coding sequences of the subblocks based on the selected partitioning information.
[0315] Referring to Figure 8(c), a vertical scan where 0 and 1 are applied to a and b, and an inverse scan where 1 and 0 are applied to a and b respectively can be supported as candidates, and a 1-bit flag can be generated to select one of them.
[0316] Alternatively, referring to Figure 8(b), it is possible to support z-scans where 0 to 3 are applied to a to d, z-scans where 1, 3, 0, and 2 are applied respectively and rotated 90 degrees to the left, inverse z-scans where 3, 2, 1, and 0 are applied respectively, and z-scans where 2, 0, 3, and 1 are applied respectively and rotated 90 degrees to the right, and a flag of one or more bits can be generated to select one of these.
[0317] Next, we will explain the case where the coding order of the subblocks is implicitly determined. For the sake of explanation, the coding order for obtaining the partition forms (c){or (f)} and (d){or (g)} in Figure 8 is shown (see Figure 5 for the prediction mode).
[0318] For example, when the on-screen prediction mode is vertical mode, horizontal mode, diagonal down right mode (modes 19 to 49), diagonal down left mode (mode 51 and above), or diagonal up right mode (mode 17 and below), it is possible to determine whether it is a vertical scan or a horizontal scan.
[0319] Alternatively, when the on-screen prediction mode is the Diagonal down left mode (number 51 or higher), it is possible to determine whether to perform a vertical scan or an inverse horizontal scan.
[0320] Alternatively, when the on-screen prediction mode is the Diagonal up-right mode (number 17 or lower), it is possible to determine whether to perform an inverse vertical scan or a horizontal scan.
[0321] The above example can be an encoding order based on a predetermined scan sequence. Here, the predetermined scan sequence can be one of z-scan, vertical scan, or horizontal scan. Alternatively, it can be a scan sequence determined by the position and distance of pixels referenced during in-screen prediction. For this purpose, an inverse scan can be additionally considered.
[0322] Figure 9 is an illustrative diagram of a reference pixel area used based on an in-screen prediction mode according to one embodiment of the present invention. Referring to Figure 9, it can be seen that the area referenced by the direction of the prediction mode is processed with shading.
[0323] Figure 9(a) shows an example where the region adjacent to the parent block is divided into left, upper, upper left, upper right, and lower left regions. Figure 9(b) shows the left and lower left regions as referenced in Diagonal up right mode, Figure 9(c) shows the left region as referenced in horizontal mode, Figure 9(d) shows the left, upper, and upper left regions as referenced in Diagonal down right mode, Figure 9(e) shows the upper region as referenced in vertical mode, and Figure 9(f) shows the upper and upper right regions as referenced in Diagonal down left mode.
[0324] When performing predictions using adjacent reference pixels (or restored pixels of subblocks), pre-defining the coding order of subblocks can have the advantage of eliminating the need to signal related information separately. A variety of partitioning configurations are possible, and based on the referenced region (or prediction mode), examples such as the following can be given:
[0325] Figure 10 shows an example of the coding order that can be had in a diagonal up-right prediction mode according to one embodiment of the present invention. An example in which priority is assigned to adjacent subblocks in the lower-left direction can be seen in Figures 10(a) to (g).
[0326] Figure 11 shows an example of an encoding order that can be had in horizontal mode according to one embodiment of the present invention. An example in which priority is assigned to adjacent subblocks in the leftward direction can be seen in Figures 11(a) to (g).
[0327] Figure 12 shows an example of the coding order that can be had in a diagonal down-right prediction mode according to one embodiment of the present invention. An example in which a priority is assigned to adjacent subblocks in the upper-left direction can be seen in Figures 12(a) to (g).
[0328] Figure 13 shows an example of an encoding order that can be had in vertical mode according to one embodiment of the present invention. An example in which priority is assigned to adjacent subblocks in the upward direction can be seen in Figures 13(a) to (g).
[0329] Figure 14 shows an example of the coding order that can be had in a diagonal down-left mode according to one embodiment of the present invention. An example in which a priority is assigned to adjacent subblocks in the upper right direction can be seen in Figures 14(a) to (g).
[0330] The above example is one instance of defining the coding order from adjacent regions where coding / decoding has been completed, and other variations are possible. Furthermore, various configurations are possible in which the coding order is defined by other coding / decoding elements.
[0331] Figure 15 is an illustrative diagram of the coding order considering the in-screen prediction mode and partition configuration according to one embodiment of the present invention. In detail, the coding order of subblocks can be implicitly determined for each partition configuration by the in-screen prediction mode. Furthermore, an example of candidate group reconstruction, in which an unobtainable partition configuration is replaced with another partition configuration, is described. For the sake of explanation, we assume that the parent block is 4M × 4N.
[0332] Referring to Figure 15(a), the image can be divided into a 4M×N format for in-screen prediction at the subblock level. In this example, this can be done using an inverse vertical scan order. If a 4M×N format is unavailable, the image can be divided into a 2M×2N format according to a predetermined priority. The encoding order for this case is shown in the figure, so a detailed explanation is omitted. If a 2M×2N format is unavailable, the image can be divided into a 4M×2N format, which has the next highest priority. The format can be supported by such predetermined priorities, and based on this, the subblock level division can be performed. If all predefined formatting options are unavailable, it can be seen that division of the parent block into subblocks is impossible, and encoding of the parent block is performed.
[0333] Referring to Figure 15(b), the image can be divided into an M×4N configuration for in-screen prediction at the subblock level. Assuming that the coding order is predetermined based on each division configuration, as shown in Figure 15(a), in this example, it can be understood that the order of substitution is 2M×2N and then 2M×4N.
[0334] Referring to Figure 15(c), the screen can be divided into a 2M×2N configuration for in-screen prediction at the sub-block level. In this example, it can be understood that the configurations are replaced in the order of 4M×2N and then 2M×4N.
[0335] Referring to Figure 15(d), the screen can be divided into a 4M×N configuration for in-screen prediction at the subblock level. In this example, it can be understood that the replacements are in the order of 2M×2N and then 4M×2N.
[0336] Referring to Figure 15(e), the screen can be divided into an M×4N configuration for in-screen prediction at the subblock level. In this example, it can be understood that the replacements are in the order of 2M×2N and 2M×4N.
[0337] The above example is one instance of supporting division into other forms in a predetermined order when an unobtainable division form exists, and various modifications are possible.
[0338] For example, when supporting a partition configuration of {4M×4N, M×4N, 4M×N}, if the M×4N or 4M×N configuration is unavailable, it may be replaced with a 2M×4N or 4M×2N configuration.
[0339] As shown in the example above, the subblock division settings can be determined by various encoding / decoding elements, and since these encoding / decoding elements can be derived from the explanation of subblock division already described above, a detailed explanation will be omitted.
[0340] Furthermore, if the division of subblocks is decided, prediction and transformation can be performed as is. As mentioned earlier, the in-screen prediction mode can be determined on a parent block basis, and prediction can be performed accordingly.
[0341] Furthermore, the conversion and inverse conversion can be performed on a subblock basis by determining the relevant settings based on the parent block and using the settings determined in the parent block. Alternatively, the relevant settings can be determined based on the subblock and used to perform the conversion and inverse conversion on a subblock basis. Conversion and inverse conversion can also be performed based on one of the aforementioned settings.
[0342] The methods according to the present invention can be embodied in a program instruction form that can be executed via various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the computer-readable medium may be specifically designed and configured for the present invention, or they may be publicly known and available to those skilled in the computer software art.
[0343] Examples of computer-readable media include hardware devices specifically configured to store and execute program instructions, such as ROM (Read Only Memory), RAM, and flash memory. Examples of program instructions include not only machine code, such as that produced by a compiler, but also high-level language code that can be executed by a computer using an interpreter or the like. The aforementioned hardware devices can be configured to operate as at least one software module to perform the operations of the present invention, and vice versa.
[0344] Furthermore, the methods or apparatus described above may be embodied either by combining all or part of their configuration or function, or by separating them.
[0345] While the above has been described with reference to preferred embodiments of the present invention, those skilled in the art will understand that the present invention can be modified and altered in various ways without departing from the spirit and scope of the invention as described in the following claims. [Industrial applicability]
[0346] This invention can be used to encode / decode video. < / l-1> < / array> < / qt>
Claims
1. An image decoding method performed by an image decoding device, Perform on-screen prediction to generate predicted blocks for the coded blocks, The residual blocks of the aforementioned encoded block are generated, Based on the predicted block and the residual block, the encoded block is restored. Equipped with, The encoded block is restored by adding the predicted block and the residual block. Performing the aforementioned in-screen prediction includes interpolating the reference sample, Depending on the in-screen prediction mode of the coding block for interpolating the reference sample, one of either fixed filtering or adaptive filtering is selectively performed. The fixed filtering uses one interpolation filter predefined in the image decoding device, while the adaptive filtering selectively uses one of a plurality of interpolation filters predefined in the image decoding device. Image decoding method.
2. The plurality of interpolation filters comprises at least one of a 4-tap cubic filter or a 4-tap Gaussian filter. The image decoding method according to claim 1.
3. An image encoding method performed by an image encoding device, Perform on-screen prediction to generate predicted blocks for the coded blocks, Based on the prediction block, residual blocks of the encoding block are generated. Encode the aforementioned residual block and then encode the encoded block. Equipped with, The residual block is generated by subtracting the prediction block from the coded block. Performing the aforementioned in-screen prediction includes interpolating the reference sample, Depending on the in-screen prediction mode of the coding block for interpolating the reference sample, one of either fixed filtering or adaptive filtering is selectively performed. The fixed filtering uses one interpolation filter predefined in the image encoding device, while the adaptive filtering selectively uses one of a plurality of interpolation filters predefined in the image encoding device. Image encoding method.
4. A non-temporary computer-readable medium for storing a bitstream generated by an image encoding method performed by an image encoding device, The aforementioned image encoding method is Perform on-screen prediction to generate predicted blocks for the coded blocks, Based on the prediction block, residual blocks of the encoding block are generated. The residual block is encoded, and the encoded block is encoded into the bitstream. Equipped with, The residual block is generated by subtracting the prediction block from the encoded block. Performing the aforementioned in-screen prediction includes interpolating the reference sample, Depending on the in-screen prediction mode of the coding block for interpolating the reference sample, one of either fixed filtering or adaptive filtering is selectively performed. The fixed filtering uses one interpolation filter predefined in the image encoding device, while the adaptive filtering selectively uses one of a plurality of interpolation filters predefined in the image encoding device. Non-temporary computer-readable media.
5. A method for transmitting a bitstream generated by an image encoding method performed by an image encoding device, The aforementioned image encoding method is Perform on-screen prediction to generate predicted blocks for the coded blocks, Based on the prediction block, residual blocks of the encoding block are generated. Encode the good-faith residual block and encode the encoded block into the bitstream. Equipped with, The residual block is generated by subtracting the prediction block from the encoded block. Performing the aforementioned in-screen prediction includes interpolating the reference sample, Depending on the in-screen prediction mode of the coding block for interpolating the reference sample, one of either fixed filtering or adaptive filtering is selectively performed. The fixed filtering uses one interpolation filter predefined in the image encoding device, while the adaptive filtering selectively uses one of a plurality of interpolation filters predefined in the image encoding device. method.