Intra-frame prediction method and device
Patent Information
- Authority / Receiving Office
- IL · IL
- Patent Type
- Patents
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2019-06-25
- Publication Date
- 2026-07-01
AI Technical Summary
Current video encoding and decoding technologies face inefficiencies in handling high-resolution images, particularly in intra prediction methods, where predicting pixel values within a block based on neighboring pixels is not optimized, leading to suboptimal compression and transmission of high-definition and ultra-high-definition images.
The proposed solution involves an intra-screen prediction method that filters reference pixels selectively, corrects predicted pixels based on their position, and performs intra prediction in sub-block units, using multiple pixel lines and deriving the intra prediction mode from default modes or multiple prediction modes (MPM) candidates, thereby enhancing encoding and decoding efficiency.
This approach improves encoding and decoding efficiency by selectively filtering and correcting reference pixels, allowing for more accurate prediction within sub-block units, leading to better compression and transmission of high-resolution images.
Smart Images

Figure 00000127_0000 
Figure 00000128_0000 
Figure 00000128_0001
Abstract
Description
In-screen prediction method and device
[0001] The present invention relates to image encoding and decoding technology, and more specifically, to an encoding / decoding method and apparatus in intra-frame prediction.
[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition) video, has been increasing across various application fields, and accordingly, high-efficiency video compression technologies are being discussed.
[0003] Various video compression technologies exist, such as inter-frame prediction technology that predicts pixel values in the current picture from previous or subsequent pictures, intra-frame prediction technology that predicts pixel values in the current picture using pixel information within the current picture, and entropy coding technology that assigns short codes to values with high frequency and long codes to values with low frequency; by utilizing these video compression technologies, video data can be effectively compressed for transmission or storage.
[0004] The present invention relates to image encoding and decoding technology, and more specifically, to an encoding / decoding method and apparatus in intra-frame prediction.
[0005] The in-frame prediction method and device according to the present invention can induce an in-frame prediction mode of a current block, determine a pixel line for in-frame prediction of the current block among a plurality of pixel lines, and perform in-frame prediction of the current block based on the in-frame prediction mode and the determined pixel line.
[0006] The in-frame prediction method and device according to the present invention can filter a first reference pixel of the determined pixel line.
[0007] In the in-frame prediction method and apparatus according to the present invention, the filtering step may be selectively performed based on a first flag indicating whether filtering is performed on a first reference pixel for in-frame prediction.
[0008] In the intra-frame prediction method and apparatus according to the present invention, the first flag is derived from a decoder based on the encoding parameter of the current block, and the encoding parameter may include at least one of a block size, a component type, an intra-frame prediction mode, or whether intra-frame prediction at the sub-block level is applied.
[0009] The in-frame prediction method and device according to the present invention can correct the predicted pixels of the current block according to the in-frame prediction.
[0010] In the in-frame prediction method and apparatus according to the present invention, the correction step may further include the step of determining at least one of a second reference pixel or a weight for the correction based on the position of the predicted pixel of the current block.
[0011] In the in-frame prediction method and apparatus according to the present invention, the correction step may be selectively performed by considering at least one of the position of the pixel line of the current block, the in-frame prediction mode of the current block, or whether to perform in-frame prediction at the sub-block unit level of the current block.
[0012] In the in-frame prediction method and apparatus according to the present invention, the in-frame prediction is performed in units of sub-blocks of the current block, and the sub-blocks may be determined based on at least one of a second flag regarding whether to divide, division direction information, or division count information.
[0013] In the in-frame prediction method and apparatus according to the present invention, the in-frame prediction mode of the current block may be derived based on a predetermined default mode or a plurality of MPM candidates.
[0014] According to the present invention, encoding / decoding efficiency can be improved through prediction at the sub-block level.
[0015] According to the present invention, the encoding / decoding efficiency of intra-frame prediction can be improved through multi-pixel line-based intra-frame prediction.
[0016] According to the present invention, the encoding / decoding efficiency of an in-frame prediction can be improved through filtering of a reference pixel.
[0017] According to the present invention, the encoding / decoding efficiency of intra-frame prediction can be improved through the correction of intra-frame prediction pixels.
[0018] According to the present invention, by inducing an intra-frame prediction mode based on a default mode or an MPM candidate, the encoding / decoding efficiency of the intra-frame prediction mode can be improved.
[0019] FIG. 1 is a block diagram of an image encoding device according to one embodiment of the present invention.
[0020] FIG. 2 is a block diagram of an image decoding device according to one embodiment of the present invention.
[0021] Figure 3 is an example diagram showing a tree-based block shape.
[0022] Figure 4 is an example diagram showing a type-based block shape.
[0023] FIG. 5 is an example diagram showing various block shapes that can be obtained from the block division part of the present invention.
[0024] FIG. 6 is an illustrative diagram for explaining tree-based partitioning according to one embodiment of the present invention.
[0025] FIG. 7 is an illustrative diagram for explaining tree-based partitioning according to one embodiment of the present invention.
[0026] FIG. 8 illustrates a block division process according to one embodiment of the present invention.
[0027] Figure 9 is an example diagram showing a pre-defined intra-frame prediction mode in a video encoding / decoding device.
[0028] Figure 10 shows an example of pixels compared between color spaces to obtain correlation information.
[0029] Figure 11 is an example diagram illustrating the reference pixel configuration used for in-frame prediction.
[0030] Figure 12 is an example diagram illustrating the range of reference pixels used for in-frame prediction.
[0031] Figure 13 is a figure showing the current block and adjacent blocks in relation to the generation of predicted blocks.
[0032] Figures 14 and 15 are examples of some blocks for verifying partition information.
[0033] Figure 16 is an example diagram showing various cases of block division.
[0034] FIG. 17 shows an example of block division according to one embodiment of the present invention.
[0035] Figure 18 shows various examples of setting up in-screen prediction mode candidates for a block where prediction information occurs (prediction block in this example, 2N x N).
[0036] Figure 19 shows various examples of setting up in-screen prediction mode candidates for a block where prediction information occurs (prediction block in this example, N x 2N).
[0037] FIG. 20 shows an example of block division according to one embodiment of the present invention.
[0038] Figures 21 and 22 show various examples of setting up in-screen prediction mode candidates for blocks where prediction information occurs.
[0039] FIGS. 23 to 25 show examples of generating a prediction block according to the prediction mode of a neighbor block.
[0040] Figure 26 is an example diagram of the relationship between the current block and neighboring blocks.
[0041] Figures 27 and 28 illustrate an in-screen prediction considering the directionality of the prediction mode.
[0042] FIG. 29 is an example diagram illustrating the reference pixel configuration used for in-frame prediction.
[0043] FIGS. 30 to 35 are example diagrams of reference pixel configurations.
[0044] The in-frame prediction method and device according to the present invention can induce an in-frame prediction mode of a current block, determine a pixel line for in-frame prediction of the current block among a plurality of pixel lines, and perform in-frame prediction of the current block based on the in-frame prediction mode and the determined pixel line.
[0045] The in-frame prediction method and device according to the present invention can filter a first reference pixel of the determined pixel line.
[0046] In the in-frame prediction method and apparatus according to the present invention, the filtering step may be selectively performed based on a first flag indicating whether filtering is performed on a first reference pixel for in-frame prediction.
[0047] In the intra-frame prediction method and apparatus according to the present invention, the first flag is derived from a decoder based on the encoding parameter of the current block, and the encoding parameter may include at least one of a block size, a component type, an intra-frame prediction mode, or whether intra-frame prediction at the sub-block level is applied.
[0048] The in-frame prediction method and device according to the present invention can correct the predicted pixels of the current block according to the in-frame prediction.
[0049] In the in-frame prediction method and apparatus according to the present invention, the correction step may further include the step of determining at least one of a second reference pixel or a weight for the correction based on the position of the predicted pixel of the current block.
[0050] In the in-frame prediction method and apparatus according to the present invention, the correction step may be selectively performed by considering at least one of the position of the pixel line of the current block, the in-frame prediction mode of the current block, or whether to perform in-frame prediction at the sub-block unit level of the current block.
[0051] In the in-frame prediction method and apparatus according to the present invention, the in-frame prediction is performed in units of sub-blocks of the current block, and the sub-blocks may be determined based on at least one of a second flag regarding whether to divide, division direction information, or division count information.
[0052] In the in-frame prediction method and apparatus according to the present invention, the in-frame prediction mode of the current block may be derived based on a predetermined default mode or a plurality of MPM candidates.
[0053] The present invention is capable of various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the invention to specific embodiments, and it should be understood that the invention includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the invention.
[0054] Terms such as first, second, A, B, etc., may be used to describe various components, but said components should not be limited by said terms. These terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.
[0055] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0056] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "having" are intended to specify the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0057] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having meanings consistent with the context of the relevant technology and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this application.
[0058] The video encoding and decoding devices may be user terminals such as a personal computer (PC), notebook computer, personal digital assistant (PDA), portable multimedia player (PMP), PlayStation portable (PSP), wireless communication terminal, smartphone, TV, virtual reality (VR), augmented reality (AR), mixed reality (MR), head-mounted display (HMD), smart glasses, etc., or server terminals such as application servers and service servers, and may include various devices equipped with communication devices such as communication modems for performing communication with various devices or wired / wireless communication networks, memory for storing various programs and data for encoding or decoding images or for predicting within or between frames for encoding or decoding, and processors for executing programs to perform calculations and control. In addition, the video encoded into a bitstream by the video encoding device can be transmitted to a video decoding device in real-time or non-real-time through wired or wireless communication networks such as the Internet, a local area wireless network, a wireless LAN network, a WiBro network, or a mobile communication network, or through various communication interfaces such as a cable or a Universal Serial Bus (USB), and then decoded by the video decoding device to be restored and played back as a video.
[0059] In addition, the image encoded into a bitstream by the image encoding device may be transmitted from the encoding device to the decoder through a computer-readable recording medium.
[0060] The aforementioned video encoding device and video decoding device may each be separate devices, but depending on the implementation, they may be made into a single video encoding / decoding device. In that case, some components of the video encoding device may be implemented to include at least the same structure or perform at least the same function as some components of the video decoding device, as they are substantially identical technical elements.
[0061] Therefore, in the detailed description of the technical elements and their operating principles below, redundant explanations of corresponding technical elements will be omitted.
[0062] Furthermore, since the image decoding device corresponds to a computing device that applies the image encoding method performed by the image encoding device to decoding, the following description will focus on the image encoding device.
[0063] A computing device may include a memory that stores a program or software module implementing an image encoding method and / or an image decoding method, and a processor connected to the memory that executes the program. Additionally, the image encoding device may be referred to as an encoder, and the image decoding device may be referred to as a decoder.
[0064]
[0065] Typically, an image can be composed of a series of still images, which can be classified into Group of Pictures (GOP) units, and each still image can be referred to as a Picture. In this case, a Picture can represent one of a Frame or Field in a progressive signal or an interlaced signal; if encoding / decoding is performed on a frame basis, the image can be represented as a 'Frame,' and if performed on a field basis, as a 'Field.' Although this invention assumes a progressive signal for explanation, it may also be applicable to interlaced signals. Higher-level concepts such as GOPs and Sequences may exist, and each Picture can be divided into specific regions such as slices, tiles, or blocks. Additionally, a single GOP may include units such as I-Picture, P-Picture, and B-Picture. I-Picture may refer to a picture that is encoded / decoded independently without using a reference picture, while P-Picture and B-Picture may refer to pictures that are encoded / decoded by performing processes such as motion estimation and motion compensation using a reference picture. Generally, for P-Picture, I-Picture and P-Picture can be used as reference pictures, and for B-Picture, I-Picture and P-Picture can be used as reference pictures; however, these definitions may change depending on the encoding / decoding settings.
[0066] Here, the picture referenced for encoding / decoding is called the Reference Picture, and the referenced block or pixel is called the Reference Block or Reference Pixel. Furthermore, the referenced data may include not only pixel values in the spatial domain but also coefficient values in the frequency domain, as well as various encoding / decoding information generated or determined during the encoding / decoding process. For example, this could include intra-frame prediction-related information or motion-related information in the prediction unit, transformation-related information in the transformation / inverse transformation unit, quantization-related information in the quantization / inverse quantization unit, encoding / decoding-related information (context information) in the encoding / decoding unit, and filter-related information in the in-loop filter unit.
[0067] The smallest unit constituting an image may be a pixel, and the number of bits used to represent a single pixel is called bit depth. Generally, bit depth can be 8 bits, but higher bit depths may be supported depending on the encoding settings. Bit depth may support at least one bit depth depending on the color space. Additionally, it may be composed of at least one color space depending on the image's color format. Depending on the color format, it may be composed of one or more pictures of a certain size or one or more pictures of different sizes. For example, in the case of YCbCr 4:2:0, it may be composed of one luminance component (Y in this example) and two chrominance components (Cb / Cr in this example), where the ratio of the chrominance component to the luminance component may be 1:2 horizontally and vertically. As another example, in the case of 4:4:4, the horizontal and vertical ratios may be the same. As in the example above, if the picture is composed of one or more color spaces, the picture can be divided into each color space.
[0068] In the present invention, the description will be based on a specific color space (Y in this example) of a specific color format (YCbCr in this example), and the same or similar application (setting dependent on a specific color space) may be applied to other color spaces (Cb, Cr in this example) according to the color format. However, it may also be possible to have partial differences in each color space (setting independent of a specific color space). That is, a setting dependent on each color space may mean having a setting proportional to or dependent on the composition ratio of each component (e.g., determined according to 4:2:0, 4:2:2, 4:4:4, etc.), and a setting independent of each color space may mean having a setting only for that color space, regardless of or independently of the composition ratio of each component. In the present invention, depending on the encoder / decoder, some components may have independent settings or dependent settings.
[0069] The configuration information or syntax elements required during the video encoding process can be determined at the unit level, such as video, sequence, picture, slice, tile, or block. These can be stored in a bitstream in units such as VPS (Video Parameter Set), SPS (Sequence Parameter Set), PPS (Picture Parameter Set), Slice Header, Tile Header, and Block Header, and transmitted to a decoder. The decoder can then parse these elements at the same level to recover the configuration information transmitted from the encoder and use it in the video decoding process. Additionally, related information can be transmitted via a bitstream in the form of SEI (Supplement Enhancement Information) or metadata, and parsed for use. Each parameter set has a unique ID value, and a lower parameter set may contain the ID value of a higher parameter set to be referenced. For example, a lower parameter set may reference information from a higher parameter set that has a matching ID value among one or more higher parameter sets. Among the various examples of units mentioned above, if one unit includes one or more other units, the corresponding unit may be called the superordinate unit, and the included unit may be called the subordinate unit.
[0070] In the case of setting information generated in the above unit, it may include content regarding independent settings for each unit, or content regarding settings dependent on previous, subsequent, or higher units. Here, dependent settings can be understood as representing the setting information of the unit by flag information indicating that it follows the settings of previous, subsequent, or higher units (for example, a 1-bit flag where 1 means follow and 0 means do not follow). Although the setting information in the present invention will be explained primarily based on examples of independent settings, examples of addition or replacement with content regarding a dependent relationship to the setting information of previous, subsequent, or higher units of the current unit may also be included.
[0071]
[0072] FIG. 1 is a block diagram of an image encoding device according to one embodiment of the present invention. FIG. 2 is a block diagram of an image decoding device according to one embodiment of the present invention.
[0073] Referring to FIG. 1, the image encoding device may be configured to include a prediction unit, a subtraction unit, a conversion unit, a quantization unit, an inverse quantization unit, an inverse conversion unit, an addition unit, an in-loop filter unit, a memory and / or an encoding unit, some of the above components may not necessarily be included, some or all of which may be optionally included depending on the implementation, and additional components not shown may be included.
[0074] Referring to FIG. 2, the image decoder may be configured to include a decoder, a prediction unit, an inverse quantization unit, an inverse transform unit, an adder, an in-loop filter unit and / or a memory, some of the above components may not necessarily be included, some or all of which may be optionally included depending on the implementation, and additional components not shown may be included.
[0075] The video encoding device and the video decoding device may each be separate devices, but depending on the implementation, they may be combined into a single video encoding / decoding device. In that case, some components of the video encoding device may be implemented to include at least the same structure or perform at least the same function as some components of the video decoding device, as they are substantially identical technical elements. Therefore, in the detailed description of the technical elements and their operating principles below, redundant descriptions of corresponding technical elements will be omitted. Since the video decoding device corresponds to a computing device that applies the video encoding method performed by the video encoding device to decoding, the following description will focus on the video encoding device. The video encoding device may be referred to as an encoder, and the video decoding device as a decoder.
[0076] The prediction unit may include an intra-frame prediction unit that performs intra-frame prediction and an inter-frame prediction unit that performs inter-frame prediction. Intra-frame prediction may determine an intra-frame prediction mode by configuring pixels of adjacent blocks of the current block as reference pixels and generate a prediction block using said intra-frame prediction mode, and inter-frame prediction may determine motion information of the current block using one or more reference images and generate a prediction block by performing motion compensation using said motion information. It may determine whether to use intra-frame prediction or inter-frame prediction for the current block (encoding unit or prediction unit) and determine specific information according to each prediction method (e.g., intra-frame prediction mode, motion vector, reference image, etc.). At this time, the processing unit in which the prediction is performed and the processing unit in which the prediction method and specific details are determined may be determined according to the encoding / decoding settings. For example, the prediction method, prediction mode, etc. are determined at the prediction unit (or encoding unit), and the prediction is performed at the prediction block unit (or encoding unit, conversion unit).
[0077] The subtraction unit generates a residual block by subtracting the prediction block from the current block. In other words, the subtraction unit calculates the difference between the pixel value of each pixel in the current block to be encoded and the predicted pixel value of each pixel in the prediction block generated by the prediction unit, thereby creating a residual block, which is a residual signal in the form of a block.
[0078] The converter can convert a signal belonging to the spatial domain into a signal belonging to the frequency domain, and the signal obtained through this conversion process is called the transformed coefficient. For example, a residual block containing a residual signal received from the subtraction unit can be converted to obtain a transformed block containing the transformed coefficient; the input signal is determined according to the encoding settings and is not limited to the residual signal.
[0079] The transformation unit can transform the residual block using transformation techniques such as the Hadamard Transform, Discrete Sine Transform (DST Based-Transform), and Discrete Cosine Transform (DCT Based-Transform), but is not limited to these and various transformation techniques that are improved or modified thereof may be used.
[0080] For example, at least one of the above transformations may be supported, and at least one detailed transformation technique may be supported for each transformation technique. In this case, the at least one detailed transformation technique may be a transformation technique in which a part of the basis vector is configured differently in each transformation technique. For example, DST-based transformations and DCT-based transformations may be supported as transformation techniques; in the case of DST, detailed transformation techniques such as DST-I, DST-II, DST-III, DST-V, DST-VI, DST-VII, and DST-VIII may be supported, and in the case of DCT, detailed transformation techniques such as DCT-I, DCT-II, DCT-III, DCT-V, DCT-VI, DCT-VII, and DCT-VIII may be supported.
[0081] One of the above transformations (e.g., one transformation technique && one detailed transformation technique) may be set as the default transformation technique, and additional transformation techniques (e.g., multiple transformation techniques || multiple detailed transformation techniques) may be supported. Whether additional transformation techniques are supported is determined at the level of sequences, pictures, slices, tiles, etc., and related information may be generated at the level of said unit; if additional transformation techniques are supported, transformation technique selection information is determined at the level of blocks, etc., and related information may be generated.
[0082] The transformation can be performed in the horizontal / vertical direction. For example, pixel values in the spatial domain can be transformed into the frequency domain by performing a one-dimensional transformation in the horizontal direction and a one-dimensional transformation in the vertical direction using basis vectors in the transformation, thereby performing a total of two-dimensional transformations.
[0083] Additionally, the transformation in the horizontal / vertical direction can be performed adaptively. Specifically, whether adaptive performance is performed can be determined based on at least one encoding setting. For example, in the case of intra-frame prediction, if the prediction mode is horizontal mode, DCT-I may be applied in the horizontal direction and DST-I in the vertical direction; if it is vertical mode, DST-VI may be applied in the horizontal direction and DCT-VI in the vertical direction; if it is diagonal down left, DCT-II may be applied in the horizontal direction and DCT-V in the vertical direction; and if it is diagonal down right, DST-I may be applied in the horizontal direction and DST-VI in the vertical direction.
[0084] The size and shape of each transformation block are determined according to the encoding cost per candidate for the size and shape of the transformation block, and information such as the image data of each determined transformation block and the size and shape of each determined transformation block can be encoded.
[0085] Among the above transformation forms, a square shape transformation may be set as the default transformation form, and additional transformation forms (e.g., a rectangular shape) may be supported. Whether additional transformation forms are supported is determined by units such as sequences, pictures, slices, and tiles, and related information may be generated by said units, while transformation form selection information is determined by units such as blocks, and related information may be generated.
[0086] Additionally, support for the form of a transformation block may be determined based on the encoding information. In this case, the encoding information may include slice type, encoding mode, block size and shape, block partitioning method, etc. That is, one transformation form may be supported based on at least one encoding information, and multiple transformation forms may be supported based on at least one encoding information. The former may be an implicit situation, and the latter may be an explicit situation. In the explicit case, adaptive selection information indicating the optimal candidate among multiple candidate groups may be generated and recorded in the bitstream. Including this example, in the present invention, when encoding information is explicitly generated, it can be understood that the information is recorded in the bitstream in various units, and the decoder parses the relevant information in various units to restore it as decoding information. Furthermore, when encoding / decoding information is processed implicitly, it can be understood that it is processed by the same process, rules, etc., in both the encoder and the decoder.
[0087] For example, support for rectangular transformations may be determined based on the slice type. In the case of an I-slice, the supported transformation may be a square shape, and in the case of a P / B slice, it may be a square or rectangular shape.
[0088] For example, support for rectangular conversions may be determined based on the encoding mode. In the case of Intra, the supported conversion type may be a square conversion, and in the case of Inter, the supported conversion type may be a square conversion or a rectangular conversion.
[0089] For example, support for rectangular transformations may be determined based on the size and shape of the block. For blocks larger than a certain size, the supported transformation may be a square transformation, and for blocks smaller than a certain size, the supported transformation may be a square or a rectangular transformation.
[0090] For example, support for rectangular transformations may be determined based on the block partitioning method. If the block on which the transformation is performed is obtained through a Quad Tree partitioning method, the supported transformation type may be a square transformation, and if the block is obtained through a Binary Tree partitioning method, the supported transformation type may be a square or a rectangular transformation.
[0091] The above example illustrates support for a conversion form based on a single piece of encoding information, and multiple pieces of information may be combined to influence additional conversion form support settings. The above example is merely one illustration of additional conversion form support based on various encoding settings and is not limited to the above; various variations are possible.
[0092] The conversion process may be omitted depending on the encoding settings or the characteristics of the image. For example, the conversion process (including the inverse process) may be omitted depending on the encoding settings (assuming a lossless compression environment in this example). As another example, the conversion process may be omitted if compression performance through conversion is not achieved due to the characteristics of the image. In this case, the omitted conversion may be the entire unit, or one of the horizontal or vertical units may be omitted; whether such omission is supported may be determined based on the block size and shape, etc.
[0093] For example, in a setting where the omission of horizontal and vertical transformations is combined, if the transformation omission flag is 1, the transformation in the horizontal and vertical directions is not performed, and if it is 0, the transformation in the horizontal and vertical directions can be performed. In a setting where the omission of horizontal and vertical transformations operates independently, if the first transformation omission flag is 1, the transformation in the horizontal direction is not performed, and if it is 0, the transformation in the horizontal direction is performed, and if the second transformation omission flag is 1, the transformation in the vertical direction is not performed, and if it is 0, the transformation in the vertical direction is performed.
[0094] If the block size falls within range A, the omission of transformation may be supported, and if it falls within range B, the omission of transformation may not be supported. For example, if the width of the block is greater than M or the height of the block is greater than N, the above-mentioned omission of transformation flag may not be supported, and if the width of the block is less than m or the height of the block is less than n, the above-mentioned omission of transformation flag may be supported. M(m) and N(n) may be the same or different. The above-mentioned transformation-related settings may be determined at the level of a sequence, picture, slice, etc.
[0095] If additional conversion techniques are supported, the conversion technique settings may be determined based on at least one piece of encoding information. In this case, the encoding information may include slice type, encoding mode, block size and shape, prediction mode, etc.
[0096] For example, the support for conversion techniques may be determined according to the encoding mode. In the case of Intra, the supported conversion techniques may be DCT-I, DCT-III, DCT-VI, DST-II, and DST-III, and in the case of Inter, the supported conversion techniques may be DCT-II, DCT-III, and DST-III.
[0097] For example, support for conversion techniques may be determined based on the slice type. In the case of an I slice, the supported conversion techniques may be DCT-I, DCT-II, and DCT-III; in the case of a P slice, the supported conversion techniques may be DCT-V, DST-V, and DST-VI; and in the case of a B slice, the supported conversion techniques may be DCT-I, DCT-II, and DST-III.
[0098] For example, support for transformation techniques may be determined based on the prediction mode. The transformation techniques supported in prediction mode A may be DCT-I or DCT-II, the transformation techniques supported in prediction mode B may be DCT-I or DST-I, and the transformation technique supported in prediction mode C may be DCT-I. In this case, prediction modes A and B may be directional modes, and prediction mode C may be a non-directional mode.
[0099] For example, support for transformation techniques may be determined based on the size and shape of the block. For blocks larger than a certain size, the transformation technique supported may be DCT-II; for blocks smaller than a certain size, the transformation techniques supported may be DCT-II and DST-V; and for blocks larger than and smaller than a certain size, the transformation techniques supported may be DCT-I, DCT-II, and DST-I. Additionally, for square shapes, the transformation techniques supported may be DCT-I and DCT-II, and for rectangular shapes, the transformation techniques supported may be DCT-I and DST-I.
[0100] The above example is an example of supporting a conversion technique based on a single encoding information, and multiple pieces of information may be combined to be involved in setting additional conversion technique support. It is not limited to the above example and variations to other examples may also be possible. In addition, the conversion unit may transmit information necessary to generate a conversion block to the encoding unit to encode it, record the corresponding information in a bitstream and transmit it to the decoder, and the decoder of the decoder may parse the information and use it in the inverse conversion process.
[0101] The quantization unit can quantize the input signal, and the signal obtained through this quantization process is called the quantized coefficient. For example, a residual block containing residual transform coefficients received from the conversion unit can be quantized to obtain a quantized block containing quantized coefficients; the input signal is determined by the encoding settings and is not limited to the residual transform coefficients.
[0102] The quantization unit can quantize the transformed residual block using quantization techniques such as Dead Zone Uniform Threshold Quantization and Quantization Weighted Matrix, but is not limited to these and various quantization techniques that improve and modify them may be used.
[0103] The quantization unit can transmit information necessary to generate a quantization block to the encoding unit to encode it, and the corresponding information can be stored in a bitstream and transmitted to the decoder, and the decoder of the decoder can parse the information and use it in the inverse quantization process.
[0104] Although the above example was described under the assumption that the residual block is transformed and quantized through the transformation and quantization units, it is possible to generate a residual block with transformation coefficients by transforming the residual signal of the residual block and not perform the quantization process; furthermore, it is possible to perform only the quantization process without transforming the residual signal of the residual block into transformation coefficients, or to perform neither the transformation nor the quantization process. This can be determined by the encoder settings.
[0105] The encoding unit can scan the quantization coefficients, transform coefficients, or residual signals of the generated residual block according to at least one scan order (e.g., zig-zag scan, vertical scan, horizontal scan, etc.) to generate a sequence of quantization coefficients, a sequence of transform coefficients, or a sequence of signals, and can encode them using at least one entropy coding technique. At this time, information regarding the scan order may be determined according to the encoding settings (e.g., encoding mode, prediction mode, etc.), and may be implicitly determined or explicitly generated. For example, one of a plurality of scan orders may be selected according to the in-frame prediction mode. At this time, the scan pattern may be set to one of various patterns such as zig-zag, diagonal, or raster.
[0106] In addition, encoded data containing encoding information transmitted from each component can be generated and output as a bitstream, which can be implemented as a multiplexer (MUX). At this time, encoding can be performed using methods such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), and Context Adaptive Binary Arithmetic Coding (CABAC) as encoding techniques, but is not limited to these and various encoding techniques that are improved and modified thereof may be used.
[0107] When performing entropy encoding (assumed to be CABAC in this example) on the above residual block data and syntax elements such as information generated during the encoding / decoding process, the entropy encoding device may include a binarizer, a context modeler, and a binary arithmetic coder. In this case, the binary arithmetic coder may include a regular coding engine and a bypass coding engine. In this case, the regular coding engine may be a process performed in relation to the context modeler, and the bypass coding engine may be a process performed independently of the context modeler.
[0108] Since the syntax elements input to the above entropy encoding device may not be binary values, if the syntax elements are not binary values, the binarization unit can binarize the syntax elements and output a bin string composed of 0s or 1s. At this time, a bin represents a bit composed of 0s or 1s and can be encoded through a binary arithmetic encoding unit. At this time, either a regular coding unit or a bypass coding unit may be selected based on the probability of occurrence of 0s and 1s, and this may be determined according to the encoding / decoding settings. If the syntax elements are data where the frequency of 0s and 1s is the same, the bypass coding unit may be used; otherwise, the regular coding unit may be used and can be referenced when performing the next regular coding unit through context modeling (or context information update).
[0109] In this case, context is information regarding the probability of occurrence of a bin, and context modeling is a process of estimating the probability of a bin required for binary arithmetic coding by taking the bin resulting from binarization as input. For probability estimation, information on the syntactic elements of the bin, an index which is the position of the bin in the bin string, and the probability of the bin included in the surrounding blocks may be used, and at least one context table may be used for this purpose. For example, for information regarding some flags, multiple context tables may be used depending on the combination of whether the surrounding blocks use flags.
[0110] Various methods may be used when performing binarization on the above-mentioned syntactic elements. For example, it can be classified into fixed-length binarization and variable-length binarization; in the case of variable-length binarization, unary binarization (Trunacted Unary Binarization), truncated rice binarization, K-th Exp-Golomb binarization, truncated binary binarization, etc., may be used. Additionally, signed or unsigned binarization may be performed depending on the range of values of the syntactic elements. The binarization process for syntactic elements occurring in the present invention may be performed by including not only the binarization mentioned in the above examples but also other additional binarization methods.
[0111] The inverse quantization unit and the inverse transform unit can be implemented by performing the processes in the transform unit and the quantization unit in reverse. For example, the inverse quantization unit can inversely quantize the quantized transform coefficients generated in the quantization unit, and the inverse transform unit can inversely transform the inversely quantized transform coefficients to generate a restored residual block.
[0112] The adder adds the prediction block and the restored residual block to restore the current block. The restored block is stored in memory and can be used as reference data (prediction unit and filter unit, etc.).
[0113] The in-loop filter section may include at least one post-processing filter process, such as a deblocking filter, a Sample Adaptive Offset (SAO), and an Adaptive Loop Filter (ALF). The deblocking filter can remove block distortion occurring at the boundaries between blocks in the reconstructed image. The ALF can perform filtering based on a value obtained by comparing the reconstructed image with the input image. Specifically, filtering can be performed based on a value obtained by comparing the reconstructed image with the input image after the blocks have been filtered through the deblocking filter. Alternatively, filtering can be performed based on a value obtained by comparing the reconstructed image with the input image after the blocks have been filtered through the SAO.
[0114] Memory can store restored blocks or pictures. The restored blocks or pictures stored in memory can be provided to a prediction unit that performs intra-frame prediction or inter-frame prediction. Specifically, a storage space in the form of a queue for the bitstream compressed in the encoder can be processed as a Coded Picture Buffer (CPB), and a space for storing the decoded image in picture units can be processed as a Decoded Picture Buffer (DPB). In the case of the CPB, decoding units are stored according to the decoding order and emulate the decoding operation within the encoder, and the compressed bitstream can be stored during the emulation process. The bitstream output from the CPB is restored through a decoding process, and the restored image is stored in the DPB, and the pictures stored in the DPB can be referenced during subsequent image encoding and decoding processes.
[0115] The decoding unit can be implemented by performing the process in reverse in the encoding unit. For example, it can receive a sequence of quantization coefficients, a sequence of transformation coefficients, or a signal sequence from a bitstream and decode them, and parse decoded data containing decoding information and transmit it to each component.
[0116]
[0117] Meanwhile, although not illustrated in the image encoding device and image decoding device of FIGS. 1 and 2, a block splitting unit may be further included. Information regarding the basic encoding unit can be obtained from the picture splitting unit, and the basic encoding unit may refer to the basic (or starting) unit for prediction, transformation, quantization, etc., during the image encoding / decoding process. In this case, the encoding unit may be composed of one luminance encoding block and two chrominance encoding blocks according to the color format (YCbCr in this example), and the size of each block may be determined according to the color format. In the example described below, the explanation will be based on the block (luminance component in this example). At this time, the explanation is based on the premise that the block is a unit that can be obtained after each unit is determined, and assumes that similar settings can be applied to other types of blocks.
[0118] The block division unit can be configured in relation to each component of the image encoding device and the decoder, and through this process, the size and shape of the block can be determined. At this time, the configured block may be defined differently depending on the component; for example, a prediction block for the prediction unit, a conversion block for the conversion unit, and a quantization block for the quantization unit may be included. This is not limited to these, and additional block units may be defined according to other components. The size and shape of the block may be defined by the width and height of the block.
[0119] In the block division section, blocks can be represented as M × N, and the maximum and minimum values of each block can be obtained within a range. For example, if the block shape supports a square and the maximum value is set to 256×256 and the minimum value to 8×8, then 2 m ×2 m You can obtain a block of size (in this example, m is an integer from 3 to 8; e.g., 8×8, 16×16, 32×32, 64×64, 128×128, 256×256), a block of size 2m × 2m (in this example, m is an integer from 4 to 128), or a block of size m × m (in this example, m is an integer from 8 to 256). Alternatively, if the block shape supports squares and rectangles and has the same range as the above example, 2 m × 2 n A block of size (in this example, m and n are integers from 3 to 8. Assuming the case where the aspect ratio is at most 2:1, for example, 8×8, 8×16, 16×8, 16×16, 16×32, 32×16, 32×32, 32×64, 64×32, 64×64, 64×128, 128×64, 128×128, 128×256, 256×128, 256×256. Depending on the decoding / encoding settings, there may be no limit on the aspect ratio or a maximum value of the aspect ratio may exist) can be obtained. Or, a block of size 2m × 2n (in this example, m and n are integers from 4 to 128) can be obtained. Alternatively, you can obtain a block of size m × n (in this example, m and n are integers from 8 to 256).
[0120] The obtainable blocks may be determined according to the encoding / decoding settings (e.g., block type, partitioning method, partitioning settings, etc.). For example, the coding block is 2 m × 2 n Blocks of size, Prediction Blocks are blocks of size 2m × 2n or m × n, Transform Blocks are 2m × 2 n A block of size can be obtained. Based on the above settings, information such as block size and range (e.g., information related to exponents, multiples, etc.) can be generated.
[0121] The above range (defined as the maximum and minimum values in this example) may be determined according to the type of block. Additionally, for some blocks, range information may be explicitly generated, while for some blocks, range information may be implicitly determined. For example, for encoding and conversion blocks, relevant information may be explicitly generated, while for prediction blocks, relevant information may be implicitly processed.
[0122] In explicit cases, at least one range information may be generated. For example, in the case of an encoding block, range information may be generated for a maximum value and a minimum value. Alternatively, it may be generated based on the difference between a maximum value and a preset minimum value (e.g., 8) (e.g., generated based on the above setting, such as information on the difference value of the exponent between the maximum value and the minimum value). Additionally, information on multiple ranges for the width and height of a rectangular block may be generated.
[0123] In the case of implicit, range information can be obtained based on the encoding / decoding settings (e.g., block type, partitioning method, partitioning settings, etc.). For example, in the case of a prediction block, maximum and minimum value information can be obtained from the candidate set (M × N and m / 2 × n / 2 in this example) obtainable from the upper unit encoding block (maximum size M × N and minimum size m × n in this example) with the partitioning settings of the prediction block (quad tree partitioning + partitioning depth 0 in this example).
[0124] The size and shape of the initial (or starting) block of the block division unit can be determined from the higher-level unit. In the case of an encoding block, the basic encoding block obtained from the picture division unit may be the initial block; in the case of a prediction block, the encoding block may be the initial block; and in the case of a conversion block, either the encoding block or the prediction block may be the initial block, which can be determined according to the encoding / decoding settings. For example, if the encoding mode is Intra, the prediction block may be the higher-level unit of the conversion block, and if it is Inter, the prediction block may be a unit independent of the conversion block. The initial block can be divided into small-sized blocks as the starting unit of the division. Once the optimal size and shape for the division of each block are determined, that block may be designated as the initial block of the lower-level unit. For example, in the former case, it may be the encoding block, and in the latter case (lower-level unit), it may be the prediction block or the conversion block. Once the initial block of the lower-level unit is determined as in the above example, a division process can be performed to find the block of the optimal size and shape, similar to the higher-level unit.
[0125] In summary, the block splitting unit can divide a basic coding unit (or maximum coding unit) into at least one coding unit (or sub-coding unit). Additionally, the coding unit can be divided into at least one prediction unit and can be divided into at least one transformation unit. The coding unit can be divided into at least one coding block, and the coding block can be divided into at least one prediction block and can be divided into at least one transformation block. The prediction unit can be divided into at least one prediction block, and the transformation unit can be divided into at least one transformation block.
[0126] As shown in the example above, when a block of optimal size and shape is found through a mode determination process, mode information (e.g., segmentation information, etc.) can be generated. The mode information can be stored in a bitstream along with information generated in the component to which the block belongs (e.g., prediction-related information, transformation-related information, etc.) and transmitted to a decoder, and can be parsed at the same level unit in the decoder and used in the image decoding process.
[0127] The following example will explain the division method, assuming that the initial block is square-shaped; however, the same or similar application may be possible if it is rectangular.
[0128] The block partitioning section can support various partitioning methods. For example, it can support tree-based or type-based partitioning, and other methods may be applied. In the case of tree-based partitioning, partitioning information can be generated using partitioning flags, and in the case of type-based partitioning, partitioning information can be generated using index information for block types included in a pre-configured candidate group.
[0129]
[0130] Figure 3 is an example diagram showing a tree-based block shape.
[0131] Referring to FIG. 3, a represents an example in which one 2N × 2N block is obtained without any splitting, b represents two 2N × N blocks obtained through some splitting flags (horizontal splitting of a binary tree in this example), c represents two N × 2N blocks obtained through some splitting flags (vertical splitting of a binary tree in this example), and d represents four N × N blocks obtained through some splitting flags (four-part splitting of a quad tree or horizontal and vertical splitting of a binary tree in this example). The shape of the blocks obtained may be determined by the type of tree used for splitting. For example, when performing quad tree splitting, the obtainable candidate blocks may be a and d. When performing binary tree splitting, the obtainable candidate blocks may be a, b, c, and d. In the case of a quad tree, one splitting flag is supported, and if the flag is '0', a can be obtained, and if it is '1', d can be obtained. In the case of a binary tree, multiple partition flags are supported, one of which may indicate whether to partition, one of which may indicate whether to partition horizontally or vertically, and one of which may indicate whether to allow duplicates of horizontal and vertical partitions. If duplicates are allowed, the obtainable candidate blocks may be a, b, c, and d, and if duplicates are not allowed, the obtainable candidate blocks may be a, b, and c. In the case of a quad tree, it may be a basic tree-based partitioning method, and additionally, a tree partitioning method (binary tree in this example) may be included in the tree-based partitioning method. Multiple tree partitions can be performed if a flag allowing additional tree partitioning is implicitly or explicitly enabled. The tree-based partitioning may be a method capable of recursive partitioning. That is, a partitioned block may be set back as the initial block to perform tree-based partitioning, which can be determined by partitioning settings such as the partition range and the allowed partition depth. This may be an example of a hierarchical partitioning method.
[0132]
[0133] Figure 4 is an example diagram showing a type-based block shape.
[0134] Referring to FIG. 4, depending on the type, the block after partitioning may have a 1-partition (a in this example), a 2-partition (b, c, d, e, f, g in this example), or a 4-partitioned form (h in this example). A candidate group can be formed through various configurations. For example, a candidate group can be formed as a, b, c, n in FIG. 5, or a, b ~ g, n, or a, n, q, etc., but is not limited thereto and various variations including the examples described below may be possible. When a flag allowing symmetric partitioning is activated, the supported blocks may be FIG. 4 a, b, c, and h, and when a flag allowing asymmetric partitioning is activated, the supported blocks may be FIG. 4 a ~ h. In the former case, the relevant information (the flag allowing symmetric partitioning in this example) may be implicitly activated, and in the latter case, the relevant information (the flag allowing asymmetric partitioning in this example) may be explicitly generated. Type-based partitioning may be a method that supports a single partition. Compared to tree-based partitioning, blocks obtained through type-based partitioning may not be able to be further partitioned. This may be an example where the partitioning depth allowed is 0 (e.g., single-layer partitioning).
[0135]
[0136] FIG. 5 is an example diagram showing various block shapes that can be obtained from the block division part of the present invention.
[0137] Referring to FIG. 5, blocks a to s can be obtained depending on the division setting and division method, and additional block shapes not shown may also be possible.
[0138] For example, asymmetric partitioning may be allowed in tree-based partitioning. For example, in the case of a binary tree, blocks such as FIG. 5 b and c (where the tree is partitioned into multiple blocks in this example) may be possible, or blocks such as FIG. 5 b through g (where the tree is partitioned into multiple blocks in this example) may be possible. If the flag allowing asymmetric partitioning is explicitly or implicitly disabled according to the decoding / encoding settings, the obtainable candidate blocks may be b or c (assuming that duplicate horizontal and vertical partitioning is not allowed in this example), and if the flag allowing asymmetric partitioning is enabled, the obtainable candidate blocks may be b, d, e (where the tree is partitioned horizontally in this example) or c, f, g (where the tree is partitioned vertically in this example). This example may correspond to a case where the partitioning direction is determined by the horizontal or vertical partitioning flag and the block shape is determined according to the asymmetric allowance flag, but is not limited thereto and variations to other examples may also be possible.
[0139] For example, additional tree splitting may be allowed in tree-based splitting. For instance, splitting such as triple trees, quad-type trees, and octa trees may be possible, thereby obtaining n split blocks (3, 4, 8 in this example, where n is an integer). In the case of a triple tree, the supported blocks (when split into multiple blocks in this example) may be h to m; in the case of a quad-type tree, the supported blocks may be n to p; and in the case of an octa tree, the supported blocks may be q. Whether the above tree-based splitting is supported may be implicitly determined by the decoding / encoding settings or may be explicitly generated with relevant information. Additionally, depending on the decoding / encoding settings, it may be used alone or a combination of binary tree and quad tree splitting may be used. For example, in the case of a binary tree, blocks such as Fig. 5 b and c may be possible, and in the case where a binary tree and a triple tree are used in combination (assuming in this example that the usage scope of the binary tree and the usage scope of the triple tree partially overlap), blocks such as b, c, i, and l may be possible. If a flag allowing additional splitting other than the existing tree is explicitly or implicitly disabled according to the decryption / encryption settings, the obtainable candidate block may be b or c; if enabled, the obtainable candidate block may be b, i or b, h, i, j (horizontal splitting in this example) or c, l or c, k, l, m (vertical splitting in this example). This example may correspond to a case where the splitting direction is determined by a horizontal or vertical splitting flag and the block shape is determined according to a flag allowing additional splitting, but is not limited thereto and variations to other examples may also be possible.
[0140] For example, non-rectangular partitions may be allowed in type-based blocks. For example, partitions of the form r, s may be possible. When combined with the aforementioned group of type-based block candidates, blocks a, b, c, h, r, s or a ~ h, r, s may be supported blocks. Additionally, blocks supporting n partitions such as h ~ m (e.g., n is an integer; in this example, 3 excluding 1, 2, and 4) may be included in the candidate group.
[0141] The splitting method can be determined based on the encoding / decoding settings.
[0142] For example, the partitioning method may be determined based on the type of block. For instance, encoding blocks and transform blocks may use tree-based partitioning, while prediction blocks may use type-based partitioning. Alternatively, a partitioning method combining both approaches may be used. For instance, prediction blocks may use a partitioning method that mixes tree-based and type-based partitioning, and the partitioning method applied may differ depending on at least one range of the block.
[0143] For example, the partitioning method can be determined based on the block size. For instance, between the maximum and minimum values of a block, tree-based partitioning may be possible for some ranges (e.g., a×b to c×d, where the latter is larger), while type-based partitioning may be possible for other ranges (e.g., e×f to g×h). In this case, range information based on the partitioning method may be explicitly generated or implicitly determined.
[0144] For example, the partitioning method may be determined based on the shape of the block (or the block before partitioning). For instance, if the block is square, tree-based partitioning and type-based partitioning may be possible. Alternatively, if the block is rectangular, tree-based partitioning may be possible.
[0145] The splitting settings can be determined based on the encoding / decoding settings.
[0146] For example, splitting settings can be determined based on the type of block. For instance, in tree-based splitting, quadtree splitting can be used for the encoding block and prediction block, and binary tree splitting can be used for the transformation block. Alternatively, the allowable splitting depth for the encoding block can be set to m, for the prediction block to n, and for the transformation block to o, where m, n, and o may or may not be the same.
[0147] For example, partitioning settings may be determined based on the size of the block. For instance, quad tree partitioning may be possible for a portion of the block range (e.g., a×b to c×d), and binary tree partitioning for another portion (e.g., e×f to g×h; in this example, c×d is assumed to be larger than g×h). In this case, the aforementioned ranges may include all ranges between the maximum and minimum values of the block, and the ranges may have non-overlapping or overlapping settings. For instance, the minimum value of a portion of the range may be equal to the maximum value of that portion, or the minimum value of a portion may be smaller than the maximum value of that portion. If overlapping ranges exist, the partitioning method with the higher maximum value may have priority. That is, the execution of a lower-priority partitioning method may be determined based on the partitioning result of the higher-priority partitioning method. In this case, range information based on the tree type may be explicitly generated or implicitly determined.
[0148] As another example, a type-based partitioning with some candidate groups may be possible in a portion of the block's range (same as the above example), and a type-based partitioning with some candidate groups (which differ from the former candidate group in at least one configuration in this example) may be possible in a portion of the block's range (same as the above example). In this case, the said range may include all ranges between the maximum and minimum values of the block, and the said ranges may have a setting where they do not overlap with each other.
[0149] For example, the partitioning settings can be determined based on the shape of the block. For instance, if the block is square, quad tree partitioning may be possible. Alternatively, if the block is rectangular, binary tree partitioning may be possible.
[0150] For example, splitting settings may be determined based on encoding / decoding information (e.g., slice type, color component, encoding mode, etc.). For instance, if the slice type is I, quad tree (or binary tree) splitting may be possible within a specific range (e.g., a×b ~ c×d); if it is P, within a specific range (e.g., e×f ~ g×h); and if it is B, within a specific range (e.g., i×j ~ k×l). Additionally, the allowable splitting depth for quad tree (or binary tree) splitting can be set to m for slice type I, n for slice type P, and o for slice type B, where m, n, and o may or may not be the same. Some slice types may have the same settings as other slices (e.g., slices P and B).
[0151] As another example, the allowable depth of the quad tree (or binary tree) partition can be set to m when the color component is a luminance component and to n when it is a chrominance component, and m and n may or may not be the same. Additionally, the range of the quad tree (or binary tree) partition when the color component is a luminance component (e.g., a×b ~ c×d) and the range of the quad tree (or binary tree) partition when the color component is a chrominance component (e.g., e×f ~ g×h) may or may not be the same.
[0152] As another example, when the encoding mode is Intra, the allowable depth for quad tree (or binary tree) partitioning may be m, and when it is Inter, it may be n (assuming n is greater than m in this example), and m and n may or may not be the same. Additionally, the range of quad tree (or binary tree) partitioning when the encoding mode is Intra and the range of quad tree (or binary tree) partitioning when the encoding mode is Inter may or may not be the same.
[0153] In the above example, information regarding whether to support adaptive partitioning candidate group configuration based on encoding / decoding information may be explicitly generated or implicitly determined.
[0154] The above example explains cases where the partitioning method and partitioning settings are determined according to the encoding / decoding settings. The above example represents some cases based on each element, and variations to other cases may also be possible. Furthermore, the partitioning method and partitioning settings may be determined by a combination of multiple elements. For example, the partitioning method and partitioning settings may be determined by the block type, size, shape, encoding / decoding information, etc.
[0155] In addition, the elements involved in the partitioning method, settings, etc. in the above example may be implicitly determined or explicitly generate information to determine whether adaptive cases such as the above example are permitted.
[0156] Among the above partition settings, the partition depth refers to the number of times the initial block is spatially partitioned (in this example, the partition depth of the initial block is 0), and as the partition depth increases, it can be partitioned into smaller blocks. The depth-related settings can be varied depending on the partitioning method. For example, among the methods for performing tree-based partitioning, the partition depth of a quad tree and the partition depth of a binary tree can use a single common depth, and individual depths can be used depending on the type of tree.
[0157] In the above example, if individual split depths are used depending on the type of tree, the split depth can be set to 0 at the starting position of the tree's split (the block before the split in this example). The split depth can be calculated based on the starting position of the split, rather than based on the split range of each tree (the maximum value in this example).
[0158] FIG. 6 is an illustrative diagram for explaining tree-based partitioning according to one embodiment of the present invention.
[0159] a represents an example of quad tree and binary tree partitioning. Specifically, the top-left block of a represents quad tree partitioning, the top-right and bottom-left blocks represent quad tree and binary tree partitioning, and the bottom-right block represents binary tree partitioning. In the figure, the solid line (Quad1 in this example) represents the boundary line partitioned into a quad tree, the dotted line (Binary1 in this example) represents the boundary line partitioned into a binary tree, and the thick solid line (Binary2 in this example) represents the boundary line partitioned into a binary tree. The difference between the dotted line and the thick solid line lies in the partitioning method.
[0160] For example, (the top-left block has a quad tree splitting allowable depth of 3. If the current block is N × N, splitting is performed until either the width or the height reaches (N >> 3), but splitting information is generated up to (N >> 2). This applies commonly to the examples described below. Assume the maximum and minimum values of the quad tree are N × N and (N >> 3) × (N >> 3).) When quad tree splitting is performed, the top-left block can be split into 4 blocks, each having a length of 1 / 2 of the width and height. The split flag can have a value of '1' when splitting is enabled and '0' when splitting is disabled. According to the above settings, the split flag of the top-left block can occur as in the top-left block of b.
[0161] For example, (assuming the upper right block has a quad tree splitting allowable depth of 0 and a binary tree splitting allowable depth of 4, the maximum and minimum values for quad tree splitting are N × N and (N >> 2) × (N >> 2), and the maximum and minimum values for binary tree splitting are (N >> 1) × (N >> 1) and (N >> 3) × (N >> 3)), the upper right block can be split into 4 blocks, each having a length of 1 / 2 of the width and height when performing quad tree splitting on the initial block. The size of the split blocks is (N >> 1) × (N >> 1), and depending on the settings of this example, binary tree splitting (where the size is larger than the minimum value for quad tree splitting in this example but the splitting allowable depth is limited) may be possible. That is, this example may be an example where the overlapping use of quad tree splitting and binary tree splitting is not possible. The binary tree splitting information in this example may consist of multiple splitting flags. Some flags may be horizontal split flags (corresponding to x in x / y in this example) and some flags may be vertical split flags (corresponding to y in x / y in this example), and the configuration of the split flags may have settings similar to quad tree splitting (e.g., whether it is enabled). In this example, the two flags may be enabled simultaneously. In the figure, if flag information is generated as '-', '-' may correspond to the implicit processing of a flag that occurs when additional splitting is impossible according to conditions such as the maximum value, minimum value, and allowed splitting depth resulting from tree splitting. According to the above settings, the split flag of the upper right block may occur as in the upper right block of b.
[0162] For example, (assuming the bottom-left block has a quad tree splitting allowable depth of 3 and a binary tree splitting allowable depth of 2. The maximum and minimum values for quad tree splitting are N × N, (N >> 3) × (N >> 3). The maximum and minimum values for binary tree splitting are (N >> 2) × (N >> 2) and (N >> 4) × (N >> 4). In the overlapping range, the splitting priority is given to quad tree splitting.) The bottom-left block can be split into 4 blocks, each having a length of 1 / 2 of the width and height when quad tree splitting is performed on the initial block. The size of the split blocks is (N >> 1) × (N >> ), and depending on the settings of this example, quad tree splitting and binary tree splitting may be possible. That is, this example may be an example where the overlapping use of quad tree splitting and binary tree splitting is possible. In this case, whether to perform binary tree splitting may be determined based on the result of quad tree splitting, which is given priority. If quad tree splitting is performed, binary tree splitting is not performed, and if quad tree splitting is not performed, binary tree splitting may be performed. If quad tree splitting is not performed, further quad tree splitting may not be possible even if conditions for splitting are met according to the above settings. The splitting information of the binary tree in this example may consist of multiple splitting flags. Some flags may be splitting flags (corresponding to x in x / y in this example), and some flags may be splitting direction flags (corresponding to y in x / y in this example; whether y information is generated may be determined based on x), and the splitting flags may have settings similar to quad tree splitting. In this example, horizontal splitting and vertical splitting cannot be activated simultaneously. If flag information is generated as '-' in the figure, '-' may have settings similar to the above example. According to the above settings, the splitting flag of the bottom-left block may occur as in the bottom-left block of b.
[0163] For example, (assuming the bottom-right block has a binary tree splitting allowance depth of 5, and the maximum and minimum values for binary tree splitting are N × N and (N >> 2) × (N >> 3)), the bottom-right block can be split into two blocks with a length of 1 / 2 of the horizontal or vertical length when binary tree splitting is performed on the initial block. The split flag setting in this example may be the same as that of the bottom-left block. In the figure, if flag information is generated as '-', '-' may have a setting similar to the above example. This example illustrates a case where the minimum values of the horizontal and vertical of the binary tree are set differently. Depending on the above setting, the split flag of the bottom-right block may occur as in the bottom-right block of b.
[0164] As shown in the example above, after checking block information (e.g., block type, size, shape, location, slice type, color component, etc.), a division method and division settings can be determined accordingly, and a division process can be performed accordingly.
[0165] FIG. 7 is an illustrative diagram for explaining tree-based partitioning according to one embodiment of the present invention.
[0166] Referring to blocks a and b, the thick solid line (L0) represents the maximum encoding block, and the blocks demarcated by the thick solid line and other lines (L1 to L5) represent the divided encoding blocks. The numbers inside the blocks represent the location of the divided sub-blocks (following the Raster Scan order in this example), the number of '-'s represents the depth of division of the corresponding block, and the number of boundary lines between blocks may represent the number of divisions. For example, in the case of 4 divisions (quad tree in this example), the order may be UL(0)-UR(1)-DL(2)-DR(3), and in the case of 2 divisions (binary tree in this example), the order may be L or U(0) - R or D(1), which can be defined at each division depth. The example described below illustrates a case where the obtainable encoding blocks are limited.
[0167] For example, assume that the maximum encoding block of a is 64×64 and the minimum encoding block is 16×16, and that quad tree partitioning is used. In this case, since the 2-0, 2-1, and 2-2 blocks (size 16×16 in this example) are equal to the minimum encoding block size, they may not be partitioned into smaller blocks such as 2-3-0, 2-3-1, 2-3-2, and 2-3-3 blocks (size 8×8 in this example). In this case, block partitioning information is not generated because the obtainable block in the 2-0, 2-1, 2-2, and 2-3 blocks is the 16×16 block, i.e., has only one candidate.
[0168] For example, assume that the maximum encoding block of b is 64×64, the minimum encoding block is 8 in width or height, and the allowable division depth is 3. In this case, the 1-0-1-1 block (size 16×16 in this example, with a division depth of 3) can be divided into smaller blocks because it satisfies the minimum encoding block condition. However, it may not be divided into blocks with a higher division depth (blocks 1-0-1-0-0, 1-0-1-0-1 in this example) because it is equal to the allowable division depth. In this case, block division information is not generated because the obtainable block in the 1-0-1-0, 1-0-1-1 block is 16×8 blocks, i.e., has only one candidate group.
[0169] As shown in the example above, quad tree splitting or binary tree splitting may be supported depending on the encoding / decoding settings. Alternatively, a combination of quad tree splitting and binary tree splitting may be supported. For example, one of the above methods or a combination thereof may be supported depending on the block size, splitting depth, etc. Quad tree splitting may be supported when the block falls within the first block range, and binary tree splitting may be supported when it falls within the second block range. If multiple splitting methods are supported, at least one setting such as a maximum encoding block size, a minimum encoding block size, and an allowable splitting depth may be provided for each method. The above ranges may be set to overlap with each other, or they may not. Alternatively, it may be possible to set one range to include another range. These settings may be determined based on individual or combined factors such as slice type, encoding mode, and color component.
[0170] For example, the splitting settings may be determined according to the slice type. In the case of an I slice, the supported splitting settings may support splitting in the range of 128×128 to 32×32 for a quad tree and splitting in the range of 32×32 to 8×8 for a binary tree. In the case of a P / B slice, the supported block splitting settings may support splitting in the range of 128×128 to 32×32 for a quad tree and splitting in the range of 64×64 to 8×8 for a binary tree.
[0171] For example, the splitting settings may be determined according to the encoding mode. When the encoding mode is Intra, the supported splitting settings may support splitting in the range of 64×64 to 8×8 for binary trees and an allowable splitting depth of 2. When Inter, the supported splitting settings may support splitting in the range of 32×32 to 8×8 for binary trees and an allowable splitting depth of 3.
[0172] For example, the division settings may be determined according to the color component. In the case of the luminance component, the quad tree may support division in the range of 256×256 to 64×64, and the binary tree may support division in the range of 64×64 to 16×16. In the case of the chrominance component, the quad tree may support the same settings as the luminance component (in this example, the length of each block is proportional to the chrominance format), and the binary tree may support division in the range of 64×64 to 4×4 (in this example, the range for the luminance component is 128×128 to 8×8, assuming 4:2:0).
[0173] The above example describes a case where the partitioning settings are varied depending on the type of block. Additionally, some blocks may be combined with other blocks to perform a single partitioning process. For instance, when an encoding block and a transform block are combined into a single unit, a partitioning process is performed to obtain the optimal block size and shape; this may be the optimal size and shape of the encoding block as well as the transform block. Alternatively, an encoding block and a transform block may be combined into a single unit, a prediction block and a transform block may be combined into a single unit, an encoding block, a prediction block, and a transform block may be combined into a single unit, and other blocks may be combined.
[0174] Although the present invention has been described primarily in the case where individual division settings are provided for each block, it is also possible for multiple units to be combined into one to have a single division setting.
[0175] In the above process, the generated information will be stored in a bitstream in the encoder as at least one unit among a sequence, picture, slice, tile, etc., and the decoder will parse the relevant information from the bitstream.
[0176] In the video encoding / decoding process, there may be cases where the input pixel value differs from the output pixel value, and a pixel value adjustment process may be performed to prevent distortion caused by computational errors. The pixel value adjustment method is a process of adjusting pixel values that exceed the range of pixel values to within the range of pixel values, and may also be referred to as clipping.
[0177] pixel_val' = Clip_x (pixel_val, minI, maxI)Clip_x (A, B, C){if (A < B) output = B;else if (A > C) output = C;else output = A;}
[0178] Table 1 is example code for a clipping function (Clip_x) that performs pixel value adjustment. Referring to Table 1, the input pixel value (pixel_val) and the minimum value (minI) and maximum value (maxI) of the allowed pixel value range can be input as parameters to the clipping function (Clip_x). In this case, based on bit depth (bit_depth), the minimum value (minI) can be 0 and the maximum value (maxI) can be 2bit_depth - 1. When the clipping function (Clip_x) is executed, input pixel values (pixel_val, parameter A) smaller than the minimum value (minI, parameter B) can be changed to the minimum value (minI), and input pixel values larger than the maximum value (maxI, parameter C) can be changed to the maximum value (maxI). Therefore, the output value (output) can be returned as the output pixel value (pixel_val') after pixel value adjustment is completed.
[0179] At this time, the range of pixel values is determined by the bit depth, but since the pixel values constituting an image (e.g., picture, slice, tile, block, etc.) vary depending on the type and characteristics of the image, they do not necessarily occur within the entire range of pixel values. According to one embodiment of the present invention, the range of pixel values constituting an actual image can be referenced and utilized in the image encoding / decoding process.
[0180] For example, in the pixel value adjustment method according to Table 1, the minimum value (minI) of the clipping function can be used as the smallest value among the pixel values constituting the actual image, and the maximum value (maxI) of the clipping function can be used as the largest value among the pixel values constituting the actual image.
[0181] In summary, the image encoding / decoding device may include a pixel value adjustment method based on bit depth and / or a pixel value adjustment method based on a range of pixel values constituting the image. The encoding / decoding device may support flag information that determines whether to support an adaptive pixel value adjustment method; if the flag information is '1', pixel value adjustment method selection information may be generated, and if it is '0', a pre-set pixel value adjustment method (a method based on bit depth in this example) may be used as the default pixel value adjustment method. If the pixel value adjustment method selection information indicates a pixel value adjustment method based on a range of pixel values constituting the image, information related to the pixel values of the image may be included. For example, this may be an example of information regarding the minimum and maximum values of each image according to color components, and the median value described later. Information generated regarding pixel value adjustment may be recorded and transmitted in units such as video, sequence, picture, slice, tile, or block of the encoder, and the related information may be restored in the same unit by parsing the recorded information in the decoder.
[0182] Meanwhile, through the above process, the range of pixel values including the minimum and maximum pixel values can be changed (determined or defined) by adjusting pixel values based on bit depth or by adjusting pixel values based on the range of pixel values constituting the image, and additional pixel value range information can also be changed (determined or defined). For example, the maximum and minimum pixel values constituting the actual image can be changed, and the median value of the pixel values constituting the actual image can also be changed.
[0183] That is, in the process of adjusting pixel values according to bit depth, minI may represent the minimum pixel value of the image, maxI may represent the maximum pixel value of the image, I may represent the color component, and medianI may represent the median pixel value of the image. minI may be 0, maxI may be (1 << bit_depth) - 1, and medianI may be 1 << (bit_depth - 1), and median may be obtained in other forms including the above example depending on the encoding / decoding settings. The median is merely a term for explanation in the present invention and may be information representing a pixel value range that can be changed (determined or defined) according to the pixel value adjustment process during the encoding / decoding process of the image.
[0184] For example, in the process of adjusting pixel values based on the range of pixel values constituting an image, minI may be the minimum pixel value of the image, maxI may be the maximum pixel value, and medianI may be the median pixel value. medianI may be the average of the pixel values within the image, the value located in the center when aligning the pixels of the image, or a value obtained based on the pixel value range information of the image. MedianI can be derived from at least one minI and one maxI. That is, medianI may be a single pixel value existing within the pixel value range of the image.
[0185] Specifically, medianI may be a value obtained according to the pixel value range information of the image (minI, maxI in this example), such as (minI + maxI) / 2 or (minI + maxI) >> 1, (minI + maxI + 1) / 2, (minI + maxI + 1) >> 1, etc., and median may be obtained in other forms including the above examples depending on the encoding / decoding settings.
[0186] The following describes an example of the pixel value adjustment process (the median in this example).
[0187] For example, the basic bit depth is 8 bits (0 to 255), and a pixel value adjustment process based on the range of pixel values constituting the image {in this example, a minimum value of 10 and a maximum value of 190. A median value of 100 is selected under the setting (average) derived from the minimum and maximum values}. When the current block position is the first block within the image (picture in this example), there are no neighboring blocks (left, bottom left, top left, top, top right in this example) to be used for encoding / decoding, so the reference pixel can be filled with the median value (100). Using the above reference pixel, an in-frame prediction process can be performed according to the prediction mode.
[0188] For example, the basic bit depth is 10 bits (0 to 1023), and a pixel value adjustment process based on the range of pixel values constituting the image (median value 600 in this example; related syntax elements exist) is selected, and when the current block position is the first block in the image (slice, tile in this example), since there are no neighbor blocks (left, bottom left, top left, top, top right in this example) to be used for encoding / decoding, the reference pixel can be filled with the median value (600). Using the above reference pixel, an in-frame prediction process can be performed according to the prediction mode.
[0189] For example, the default bit depth is 10 bits, and a pixel value adjustment process based on the range of pixel values constituting the image (in this example, the median is 112; a related syntax element exists) is selected, and a setting is enabled in which the availability of pixels in the current block for prediction is determined according to the encoding mode of neighboring blocks (intra-frame prediction / inter-frame prediction), etc. {In this example, if the encoding mode of the block is intra-frame prediction, it can be used as a reference pixel for the current block; if it is inter-frame prediction, it cannot be used. If this setting is disabled, it can be used as a reference pixel for the current block regardless of the encoding mode of the block. The relevant syntax element is constrained_intra_pred_flag, which can occur in P or B image types. If the current block is located on the left side of the image, there are no neighbor blocks (left, bottom-left, top-left in this example) to be used for encoding / decoding. Even if neighbor blocks (top, top-right in this example) to be used for encoding / decoding exist, their use is prohibited by the above setting because the encoding mode of the corresponding block is inter-frame prediction. Therefore, if no available reference pixel exists, the reference pixel can be filled with the median value (112 in this example). In other words, since no available reference pixel exists, it can be filled with the median value of the image pixel value range. Using the above reference pixel, the intra-frame prediction process can be performed according to the prediction mode.
[0190] Although the above embodiment illustrates various cases related to the median value in the prediction unit, this may also be configured by being included in other components of image encoding / decoding. Furthermore, the above embodiment is not limited to this specific case and can be modified and expanded into various other cases.
[0191] In the present invention, the pixel value adjustment process can be applied to encoding / decoding processes such as a prediction unit, a transformation unit, a quantization unit, an inverse quantization unit, an inverse transformation unit, a filter unit, and a memory. For example, the input pixel in the pixel value adjustment method may be a reference sample or a prediction sample in the prediction process, and may be a reconstructed sample in the transformation, quantization, inverse transformation, and inverse quantization processes. Additionally, it may be a reconstructed pixel in the in-loop filter process and a storage sample in memory. In this case, the reconstructed pixel in the transformation, quantization, and inverse processes may refer to the reconstructed pixel before the application of the in-loop filter. The reconstructed pixel in the in-loop filter may refer to the reconstructed pixel after the application of the in-loop filter. The reconstructed pixel in the deblocking filter process may refer to the reconstructed pixel after the application of the deblocking filter. The reconstructed pixel in the SAO process may refer to the reconstructed pixel after the application of the SAO. The restored pixels in the ALF process may refer to the restored pixels after the ALF application. Although examples for various cases have been explained above, this is not limited to these cases and can be applied to the input, intermediate, and output stages of all encoding and decoding processes where the pixel value adjustment process is called.
[0192] The following examples are described under the assumption that the clipping function Clip_Y for the luminance component (Y) and the clipping functions Clip_Cb and Clip_Cr for the chrominance components (Cb, Cr) are supported.
[0193]
[0194] In the present invention, the prediction unit can be classified into intra-frame prediction and inter-frame prediction, and intra-frame prediction and inter-frame prediction can be defined as follows.
[0195] Intra-frame prediction may be a technique for generating a prediction value from an area of the current image (e.g., picture, slice, tile, etc.) where encoding / decoding is completed, and inter-frame prediction may be a technique for generating a prediction value from at least one image (e.g., picture, slice, tile, etc.) where encoding / decoding is completed prior to the current image.
[0196] Alternatively, intra-frame prediction may be a technique for generating a prediction value from a region of the current image where encoding / decoding is complete, but may be a prediction that excludes some prediction methods {e.g., a method for generating a prediction value from a reference image, such as block matching, template matching, etc.}, and inter-frame prediction may be a technique for generating a prediction value from at least one image where encoding / decoding is complete, and said image where encoding / decoding is complete may be composed including the current image.
[0197] Depending on the encoding / decoding settings, one of the above definitions may be followed, and the examples described below assume that the first definition is followed. Additionally, the predicted value is explained under the assumption that it is a value obtained through prediction in the spatial domain, but is not limited thereto.
[0198] FIG. 8 illustrates a block division process according to an embodiment of the present invention. Specifically, it shows examples of the size and shape of blocks obtainable according to one or more division methods, starting with a basic encoding block.
[0199] In the figure, the thick solid line represents the basic encoding block, the thick dotted line represents the quad tree partition boundary, the double solid line represents the symmetric binary tree partition boundary, the solid line represents the binary tree partition boundary, and the thin dotted line represents the asymmetric binary tree partition boundary. Except for the thick solid line, the boundaries are defined according to each partitioning method. The partitioning settings described below (e.g., partitioning type, partitioning information, order of partitioning information configuration, etc.) are not limited to the specific example and various variations may be possible.
[0200] For the sake of convenience of explanation, the explanation assumes a case where individual block partitioning settings are applied to the upper-left, upper-right, lower-left, and lower-right blocks (N x N, 64 x 64) based on the basic encoding block (2N x 2N, 128 x 128). First, it is assumed that four sub-blocks are acquired due to a single partitioning operation (partition depth 0 -> 1, i.e., partition depth increased by 1) in the initial block, and that the partitioning settings regarding the quad tree are such that the maximum encoding block is 128 x 128, the minimum encoding block is 8 x 8, and the maximum partition depth is 4, and that these are settings commonly applied to each block.
[0201] (No. 1. Top-left Block. A1 ~ A6)
[0202] This example is a case where a single tree method of partitioning (a quad tree in this example) is supported, and the size and shape of the obtainable block can be determined through a single block partitioning setting such as a maximum code block, a minimum code block, and a partition depth. In this example, when there is only one obtainable block depending on the partitioning (divided horizontally and vertically into 2 each), the partitioning information required for a single partitioning operation (based on a block of 4M x 4N before partitioning, with the partition depth increased by 1) is a flag indicating whether to partition (in this example, 0 means no partitioning, 1 means partitioning), and the obtainable candidates may be 4M x 4N and 2M x 2N.
[0203] (No. 2. Upper Right Block. A7 ~ A11)
[0204] This example is a case where multi-tree partitioning (quad tree, binary tree in this example) is supported, and the size and shape of the obtainable blocks can be determined through multiple block partitioning settings. In this example, for the binary tree, it is assumed that the maximum encoding block is 64 x 64, the minimum encoding block has a side length of 4, and the maximum partition depth is 4.
[0205] In this example, when there are two or more blocks obtainable by partitioning (two or four in this example), the partitioning information required for one partitioning operation (increasing the quad tree partitioning depth by 1) is a flag indicating whether to partition, a flag indicating the type of partition, a flag indicating the partition shape, and a flag indicating the direction of partitioning, and the obtainable candidates may be 4M x 4N, 4M x 2N, 2M x 4N, 4M x N / 4M x 3N, 4M x 3N / 4M x N, M x 4N / 3M x 4N, and 3M x 4N / M x 4N.
[0206] If the quad tree and binary tree splitting ranges overlap (i.e., a range where both quad tree splitting and binary tree splitting are possible at the current stage) and the current block (a state before splitting) is a block obtained by quad tree splitting (a block obtained by quad tree splitting from a parent block <where the splitting depth is 1 less than the current one>), the splitting information can be configured by classifying it into the following cases. That is, if a block supported according to each splitting setting can be obtained by multiple splitting methods, splitting information can be generated by classifying it through the following process.
[0207] (1) Case where quad tree splitting and binary tree splitting overlap
[0208]
[0209] In the table above, 'a' represents a flag indicating whether to split a quad tree, and if it is 1, a quad tree split (QT) is performed. If the above flag is 0, check 'b', a flag indicating whether to split a binary tree. If 'b' is 0, no further splitting is performed in the block (No Split), and if it is 1, a binary tree split is performed.
[0210] c is a flag indicating the partitioning direction; 0 means horizontal partitioning (hor) and 1 means vertical partitioning (ver). d is a flag indicating the partitioning type; 0 means symmetric partitioning (SBT, Symmetric Binary Tree) and 1 means asymmetric partitioning (ABT, Asymmetric Binary Tree). Information regarding the detailed partitioning ratio (1 / 4 or 3 / 4) in asymmetric partitioning is checked only when d is 1. When d is 0, the left and top blocks have a 1 / 4 ratio and the right and bottom blocks have a 3 / 4 ratio in the left / right or top / bottom blocks, and when d is 1, the opposite ratios are used.
[0211] (2) Cases where only binary tree splitting is possible
[0212] In the table above, partition information can be expressed using flags b through e, excluding a.
[0213] In the case of block A7 in Fig. 8, since quad tree splitting was possible in the blocks before splitting (A7 to A11) (i.e., quad tree splitting was possible but binary tree splitting was performed instead of quad tree splitting), it corresponds to the case where splitting information in (1) is generated.
[0214] On the other hand, in the case of A8 to A11, if binary tree splitting has already been performed in the blocks prior to splitting (A8 ~ A11) without quad tree splitting being performed (i.e., no longer the corresponding blocks<A8 ~ A11> In this case, it corresponds to the case where quad tree splitting is impossible) and the splitting information in (2) is generated.
[0215]
[0216] (No. 3. Bottom-left block. A12 ~ A15)
[0217] This example is a case where multi-tree partitioning (quad tree, binary tree, and terminal tree in this example) is supported, and the size and shape of the obtainable block can be determined through multiple block partitioning settings. In this example, for binary tree / terminal, it is assumed that the maximum encoding block is 64 x 64, the minimum encoding block has a side length of 4, and the maximum partition depth is 4.
[0218] In this example, when there are two or more blocks obtainable by division (2, 3, or 4 in this example), the division information required for a single division operation is a flag indicating whether to divide, a flag indicating the type of division, and a flag indicating the direction of division, and the obtainable candidates may be 4M x 4N, 4M x 2N, 2M x 4N, 4M x N / 4M x 2N / 4M x N, M x 4N / 2M x 4N / M x 4N.
[0219] If the quad tree and binary tree / binary tree partitioning ranges overlap and the current block is a block obtained by quad tree partitioning, the partitioning information can be configured by classifying it into the following cases.
[0220] (1) Case where quad tree splitting and binary tree / binary tree splitting overlap
[0221]
[0222] In the table above, a is a flag indicating whether to split a quad tree, and if it is 1, a quad tree split is performed. If the above flag is 0, check b, a flag indicating whether to split a binary tree or a tertiary tree. If b is 0, no further splitting is performed in the block, and if it is 1, a binary tree or tertiary tree split is performed.
[0223] c is a flag indicating the direction of splitting, where 0 means horizontal splitting and 1 means vertical splitting, and d is a flag indicating the type of splitting, where 0 means binary tree splitting (BT) and 1 means binary tree splitting (TT).
[0224] (2) Cases where only binary tree / binary tree splitting is possible
[0225] In the table above, the split information can be expressed using flags b through d, excluding a.
[0226] In Fig. 8, for blocks A12 and A15, since quad tree partitioning was possible in the blocks before partitioning (A12 ~ A15), it corresponds to the case where partitioning information in (1) is generated.
[0227] On the other hand, A13 and A14 correspond to cases where the splitting information in (2) is generated because the splitting was not performed into a quad tree in the blocks (A13, A14) before splitting but into a ternary tree.
[0228]
[0229] (No. 4. Bottom-left block. A16 ~ A20)
[0230] This example is a case where multi-tree partitioning (quad tree, binary tree, and terminal tree in this example) is supported, and the size and shape of the obtainable block can be determined through multiple block partitioning settings. In this example, for binary tree / terminal tree, it is assumed that the maximum encoding block is 64 x 64, the minimum encoding block has a side length of 4, and the maximum partition depth is 4.
[0231] In this example, when there are two or more blocks obtainable by division (2, 3, or 4 in this example), the division information required for a single division operation is a flag indicating whether to divide, a flag indicating the type of division, a flag indicating the shape of the division, and a flag indicating the direction of the division, and the obtainable candidates may be 4M x 4N, 4M x 2N, 2M x 4N, 4M x N / 4M x 3N, 4M x 3N / 4M x N, M x 4N / 3M x 4N, 3M x 4N / M x 4N, 4M x N / 4M x 2N / 4M x N, M x 4N / 2M x 4N / M x 4N.
[0232] If the quad tree and binary tree / binary tree partitioning ranges overlap and the current block is a block obtained by quad tree partitioning, the partitioning information can be configured by classifying it into the following cases.
[0233] (1) Case where quad tree splitting and binary tree / binary splitting overlap
[0234]
[0235] In the table above, 'a' represents a flag indicating whether to split a quad tree, and if it is 1, a quad tree split is performed. If the above flag is 0, check 'b', which is a flag indicating whether to split a binary tree. If 'b' is 0, no further splitting is performed in the block, and if it is 1, a binary tree or binary tree split is performed.
[0236] c is a flag indicating the direction of partitioning, meaning horizontal partitioning if 0 and vertical partitioning if 1. d is a flag indicating the type of partitioning, meaning binary partitioning if 0 and binary tree partitioning if 1. When d is 1, the flag e for the partition type is checked; if e is 0, symmetric partitioning is performed, and if e is 1, asymmetric partitioning is performed. When e is 1, information regarding the detailed partition ratios in asymmetric partitioning is checked, which is the same as in the previous example.
[0237] (2) Cases where only binary tree / binary tree splitting is possible
[0238] In the table above, partition information can be expressed using flags b through f, excluding a.
[0239] In Fig. 8, block A20 corresponds to the case where quad tree partitioning is possible in blocks (A16 to A19) before partitioning, so it corresponds to the case where partitioning information in (1) is generated.
[0240] On the other hand, in the case of A16 to A19, the quad tree splitting was not performed in the blocks before splitting (A16 to A19) and the binary tree splitting was performed, so it corresponds to the case where the splitting information in (2) is generated.
[0241]
[0242] The following is an explanation of the in-screen prediction of the prediction unit in the present invention.
[0243] Figure 9 is an example diagram showing a pre-defined intra-frame prediction mode in a video encoding / decoding device.
[0244] Referring to FIG. 9, 67 prediction modes are configured as a candidate group of prediction modes for in-frame prediction, of which 65 are directional modes (Nos. 2 to 66) and 2 are non-directional modes (DC, Planar). At this time, the directional modes can be classified by tilt (e.g., dy / dx) or angle information (Degree). All or part of the prediction modes described in the above example may be included in the candidate group of prediction modes for luminance components or chrominance components, and other additional modes may be included in the candidate group of prediction modes.
[0245] In addition, by utilizing the correlation between color spaces, a reconstructed block of another color space, for which encoding and decoding have been completed, can be used for the prediction of the current block, and a prediction mode supporting this can be included. For example, in the case of a chrominance component, a prediction block of the current block can be generated using a reconstructed block of the luminance component corresponding to the current block. That is, a prediction block can be generated based on a reconstructed block by considering the correlation between color spaces.
[0246] The prediction mode candidate set can be adaptively determined based on the encoding / decoding settings. The number of candidates can be increased to improve prediction accuracy, or the number of candidates can be reduced to decrease the bit amount according to the prediction mode.
[0247] For example, one of candidate groups such as candidate group A (67, 65 directional modes and 2 non-directional modes), candidate group B (35, 33 directional modes and 2 non-directional modes), and candidate group C (18, 17 directional modes and 1 non-directional mode) can be selected, and the candidate group selection can be adaptively made or determined depending on the size and shape of the block.
[0248] In addition, the composition of the prediction mode candidate group can vary depending on the encoding / decoding settings. For example, as shown in FIG. 9, the prediction mode candidate group may be composed with equal spacing between modes, or the number of modes between modes 18 and 34 in FIG. 9 may be greater than the number of modes between modes 2 and 18. Alternatively, the opposite case may be possible, and the candidate group may be adaptedly composed according to the shape of the block (i.e., square, rectangular (long horizontal shape), rectangular (long vertical shape), etc.). For example, if the width of the current block is greater than the height, the in-frame prediction modes belonging to 2 through 15 are not used and may be replaced by the in-frame prediction modes belonging to 67 through 80. On the other hand, if the width of the current block is smaller than the height, the in-frame prediction modes belonging to 53 through 66 are not used and may be replaced by the in-frame prediction modes belonging to -14 through -1.
[0249] Unless otherwise specified in the present invention, the description assumes that in-frame prediction is performed using a pre-set prediction mode candidate group (Candidate Group A) having equal mode intervals, but the main elements of the present invention may be modified and applied to adaptive in-frame prediction settings as described above.
[0250] FIG. 9 may be a prediction mode supported when the shape of the block is a square or a rectangle. Additionally, the prediction mode supported when the shape of the block is a rectangle may be a prediction mode different from the example above. For instance, the number of prediction mode candidates may differ, or the number of prediction mode candidates may be the same, but the prediction mode may be concentrated on the side with the longer length of the block and dispersed on the opposite side, or vice versa. In the present invention, the prediction mode is described under a prediction mode setting (equal spacing between directional modes) where the prediction mode is supported regardless of the shape of the block as in FIG. 9, but other cases may also be applied.
[0251] The index assigned to the prediction mode can be set using various methods. In the case of the directional mode, the index assigned to each mode can be determined based on priority information that is pre-set according to the angle or tilt information of the prediction mode. For example, a mode corresponding to the x-axis or y-axis (modes 18 and 50 in FIG. 9) may have a higher priority, a diagonal mode (modes 2, 34, and 66) having an angle difference of 45 degrees or -45 degrees relative to the horizontal or vertical mode may have a higher priority, and a diagonal mode having an angle difference of 22.5 degrees or -22.5 degrees relative to the diagonal mode may have a higher priority, and priority information can be set in this manner (next being 11.25 degrees or -11.25 degrees, etc.) or in various other ways.
[0252] Alternatively, indices can be assigned in a specific directional order based on a pre-configured prediction mode. For example, as shown in FIG. 9, indices can be assigned clockwise starting from some diagonal modes (Mode 2). The examples described below are explained under the assumption that indices are assigned clockwise based on a pre-configured prediction mode.
[0253] In addition, index information in the non-directional prediction mode may be assigned prior to the directional mode, assigned between the directional modes, or assigned at the very end, and this can be determined by the encoding / decoding settings. In this example, the non-directional mode is explained by assuming an example where the index is assigned with the highest priority among the prediction modes (assigned to a lower index; Mode 0 is Planar, Mode 1 is DC).
[0254] Although various examples of indices assigned to the prediction mode have been described through the above examples, the examples are not limited to the above, and indices may be assigned under other settings or various variations may be possible.
[0255] In the above example, priority information was described as being used for index allocation in prediction mode; however, priority information may be used not only for index allocation in prediction mode but also for the encoding / decoding process in prediction mode. For example, the above priority information may be used for MPM configuration, and multiple sets of priority information may be supported in the encoding / decoding process of prediction mode.
[0256] We will now examine the method for deriving the in-frame prediction mode of the current block (specifically, the luminance component).
[0257] The current block may use a default mode pre-defined in the video encoding / decoding device. The default mode may be a directional mode or a non-directional mode. For example, the directional mode may include at least one of a vertical mode, a horizontal mode, or a diagonal mode. The non-directional mode may include at least one of a planar mode or a DC mode. If it is determined that the current block uses a default mode, the intra-frame prediction mode of the current block may be set to the default mode.
[0258] Alternatively, the in-screen prediction mode of the current block may be derived based on multiple MPM candidates. First, a predetermined MPM candidate may be selected from the aforementioned group of prediction mode candidates. The number of MPM candidates may be three, four, five, or more. The MPM candidate may be derived based on the in-screen prediction mode of a neighboring block adjacent to the current block. The neighboring block may be a block adjacent to at least one of the left, top, top-left, bottom-left, or top-right sides of the current block.
[0259] Specifically, MPM candidates can be determined by considering whether the intra-pred mode of the left block (candIntraPredModeA) and the intra-pred mode of the top block (candIntraPredModeB) are identical, and whether candIntraPredModeA and candIntraPredModeB are non-directional modes.
[0260] For example, if candIntraPredModeA and candIntraPredModeB are identical and candIntraPredModeA is not a non-directional mode, the MPM candidates for the current block may include at least one of candIntraPredModeA, (candIntraPredModeA-n), (candIntraPredModeA+n), or a non-directional mode. Here, n may be 1, 2, or an integer greater than or equal to 1. The non-directional mode may include at least one of Planar mode or DC mode. For example, the MPM candidates for the current block may be determined as shown in Table 2 below. The indices in Table 2 specify the position or priority of the MPM candidates, but are not limited thereto. For example, index 1 may be assigned to DC mode, or index 4 may be assigned.
[0261] IndexMPM Candidate0candIntraPredModeA12 + ( ( candIntraPredModeA + 61 ) % 64 )22 + ( ( candIntraPredModeA - 1 ) % 64 )3INTRA_DC42 + ( ( candIntraPredModeA + 60 ) % 64 )
[0262] Alternatively, if candIntraPredModeA and candIntraPredModeB are not identical and neither candIntraPredModeA nor candIntraPredModeB is a non-directed mode, the MPM candidates for the current block may include at least one of candIntraPredModeA, candIntraPredModeB, (maxAB-n), (maxAB+n), or a non-directed mode. Here, maxAB represents the maximum value between candIntraPredModeA and candIntraPredModeB, and n may be 1, 2, or an integer greater than or equal to 1. The non-directed mode may include at least one of Planar mode or DC mode. For example, the MPM candidates for the current block may be determined as shown in Table 3 below. The indices in Table 3 specify the position or priority of the MPM candidates, but are not limited thereto. For example, the largest index may be assigned to DC mode. If the difference between candIntraPredModeA and candIntraPredModeB falls within a predetermined threshold range, MPM candidate 1 of Table 3 is applied, and if not, MPM candidate 2 may be applied. Here, the threshold range may mean a range greater than or equal to 2 and less than or equal to 62.
[0263] IndexMPM Candidate 1MPM Candidate 20candIntraPredModeAcandIntraPredModeA1candIntraPredModeBcandIntraPredModeB2INTRA_DCINTRA_DC32 + ( ( maxAB + 61 ) % 64 )2 + ( ( maxAB + 60 ) % 64 )42 + ( ( maxAB - 1 ) % 64 )2 + ( ( maxAB ) % 64 )
[0264] Alternatively, if candIntraPredModeA and candIntraPredModeB are not identical and only one of candIntraPredModeA and candIntraPredModeB is an undirected mode, the MPM candidates for the current block may include at least one of maxAB, (maxAB-n), (maxAB+n), or an undirected mode. Here, maxAB represents the maximum value between candIntraPredModeA and candIntraPredModeB, and n may be 1, 2, or an integer greater than or equal to 1. The undirected mode may include at least one of Planar mode or DC mode. For example, the MPM candidates for the current block may be determined as shown in Table 4 below. The indices in Table 4 specify the position or priority of the MPM candidates, but are not limited thereto. For example, index 0 may be assigned to DC mode, or the largest index may be assigned.
[0265] IndexMPM Candidate0maxAB1INTRA_DC22 + ( ( maxAB + 61 ) % 64 )32 + ( ( maxAB - 1 ) % 64 )42 + ( ( maxAB + 60 ) % 64 )
[0266] Alternatively, if candIntraPredModeA and candIntraPredModeB are not identical and both candIntraPredModeA and candIntraPredModeB are non-directional modes, the MPM candidates for the current block may include at least one of non-directional mode, vertical mode, horizontal mode, (vertical mode-m), (vertical mode+m), (horizontal mode-m), or (horizontal mode+m). Here, m may be an integer of 1, 2, 3, 4, or more. The non-directional mode may include at least one of Planar mode or DC mode. For example, the MPM candidates for the current block may be determined as shown in Table 5 below. The indices in Table 5 specify the position or priority of the MPM candidates, but are not limited thereto. For example, index 1 may be assigned to horizontal mode, or the largest index may be assigned.
[0267] IndexMPM Candidate 0INTRA_DC 1Vertical Mode 2Horizontal Mode 3(Vertical Mode-4) 4(Vertical Mode+4)
[0268] Among the aforementioned multiple MPM candidates, the MPM candidate specified by the MPM index can be set as the in-frame prediction mode of the current block. The MPM index can be encoded and signaled by a video encoding device.
[0269] As described above, an intra-frame prediction mode can be derived by selectively utilizing either the default mode or an MPM candidate. This selection can be made based on a flag signaled by the encoding device. In this case, the flag may indicate whether the intra-frame prediction mode of the current block is set to the default mode. If the flag is the first value, the intra-frame prediction mode of the current block is set to the default mode; otherwise, information such as whether the intra-frame prediction mode of the current block is derived from an MPM candidate and an MPM index may be signaled.
[0270] In the case of the color difference component, it may have the same candidate set as the prediction mode candidate set of the luminance component, or a candidate set composed of some modes from the prediction mode candidate set of the luminance component. In this case, the prediction mode candidate set of the color difference component may have a fixed configuration or a variable (or adaptive) configuration.
[0271] (Fixed candidate pool composition vs. Variable candidate pool composition)
[0272] As an example of a fixed configuration, some modes among the candidate prediction modes for the luminance component (e.g., DC, Planar, Vertical, Horizontal, Diagonal modes) <DL, UL, UR 중 적어도 하나의 모드라 가정. DL은 왼쪽 아래에서 오른쪽 위 방향으로 예측, UL은 왼쪽 위에서 오른쪽 아래 방향으로 예측, UR은 오른쪽 위에서 왼쪽 아래 방향으로 예측. 각각 도 9에서 2번, 34번, 66번 모드라 가정. 그 밖의 다른 대각선 모드도 가능> In-frame prediction can be performed by configuring (assuming) as a candidate group for the prediction mode of the color difference component.
[0273] As an example of a variable configuration, some modes among the candidate prediction modes for the luminance component (e.g., DC, Planar, Vertical, Horizontal, and Diagonal UR modes; assuming that commonly selected modes are configured as the basic prediction mode candidate set) can be configured as the basic prediction mode candidate set for the chrominance component, but there may be cases where the characteristics of the chrominance component are not properly reflected by the modes included in the above candidate set. To improve this, the configuration of the candidate prediction modes for the chrominance component can be changed variably.
[0274] For example, a location identical to or corresponding to the block of the color difference component {for example, if the corresponding location in the luminance component that corresponds to the color difference component <according to the color format> is not composed of a single block but is composed of multiple sub-blocks through block division, etc., it refers to the block at a pre-set location. In this case, the location of the pre-set block is determined from the top-left, top-right, bottom-left, bottom-right, center, top-middle, bottom-middle, left-middle, right-middle, etc. within the luminance component block corresponding to the color difference component block. When distinguished by coordinates within the image, the top-left may be a location including (0,0), the top-right (blk_width-1, 0), the bottom-left (0, blk_height-1), and the bottom-right (blk_width-1, blk_height-1); the center may be one of the coordinates (blk_width / 2-1, blk_height / 2-1), (blk_width / 2, blk_height / 2-1), (blk_width / 2-1, blk_height / 2), or (blk_width / 2, blk_height / 2); the upper-center may be one of the coordinates (blk_width / 2-1, 0) or (blk_width / 2, 0); the lower-center may be one of the coordinates (blk_width / 2-1, blk_height-1) or (blk_width / 2, blk_height-1); and the left-center The location may include one of the coordinates (0, blk_height / 2-1) or (0, blk_height / 2), and the center may include one of the coordinates (blk_width-1, blk_height / 2-1) or (blk_width-1, blk_height / 2). In other words, it refers to the block containing the above coordinate locations. The blk_width and blk_height described above represent the width and height of the luminance block, and the coordinates are not limited to the above cases but may also include other cases.In the following description, the prediction mode of the luminance component <or color mode> added to the prediction mode candidate group of the color difference component may include at least one prediction mode according to the pre-set priority <e.g., assuming the order is top-left-top-right-bottom-right-center>. If two are added, the mode of the top-left block and the mode of the top-right block are added according to the above setting. In this case, if the blocks at the top-left and top-right positions are composed of a single block, the mode of the bottom-left block, which is the next priority, is added. At least one prediction mode of the luminance component block or sub-block located at the following priority may be included in the basic prediction mode candidate group (Example 1 described below) or a new prediction mode candidate group may be formed by replacing some modes (Example 2 described below).
[0275] Alternatively, at least one prediction mode of an adjacent block located to the left, top, top-left, top-right, bottom-left, etc., centered on the current block, or a sub-block of said block (if the adjacent block consists of multiple blocks) (an adjacent block may be designated as a block at a pre-set location, and if multiple modes are included in the prediction mode candidate group of the color difference component, the prediction mode of a block located at a pre-set priority and a sub-block located at a pre-set priority within the sub-block may be included in the candidate group according to priority) may be included in the basic prediction mode candidate group, or a new prediction mode candidate group may be formed by replacing some modes.
[0276] To add further details to the above description, the prediction mode of a block of luminance components or an adjacent block (of the luminance block), as well as at least one mode derived from said prediction mode, may be included as a prediction mode for the color difference component. In the example described below, an example will be given in which the prediction mode of the luminance component is included as a prediction mode for the color difference component. A detailed description of an example in which a prediction mode derived from the prediction mode of the luminance component (e.g., an adjacent mode of said mode. For example, if horizontal mode 18 is the prediction mode of the luminance component, modes 17, 19, 16, etc. correspond to the predicted mode derived from the prediction mode of the color difference component. If multiple prediction modes from the luminance component are configured as a candidate group for the prediction mode of the color difference component, the priority of the candidate group configuration may be set in the order of the prediction mode of the luminance component - the mode derived from the prediction mode of the luminance component) or a prediction mode derived from an adjacent block is included as a candidate group for the prediction mode of the color difference component is omitted, but the same or modified settings described below may be applied.
[0277] For example (1), when the prediction mode of the luminance component matches one of the candidate modes of the color difference component prediction mode, the composition of the candidate group is the same (no change in the number of candidate groups), and when none match, the composition of the candidate group is not the same (the number of candidate groups increases).
[0278] When the composition of the candidate group in the above example is the same, the index of the prediction mode may be kept the same or assigned a different index, which can be determined according to the encoding / decoding settings. For example, when the index of the prediction mode candidate group for the color difference component is Planar (0), DC (1), Vertical (2), Horizontal (3), and Diagonal UR (4), if the prediction mode for the luminance component is Horizontal, the composition of the prediction mode candidate group does not change, and the index of each prediction mode may be kept the same or assigned a different index {in this example, Horizontal (0), Planar (1), DC (2), Vertical (3), and Diagonal UR (4)}. Such index reconstruction may be an example of a process performed for the purpose of generating fewer mode bits (assuming: assigning fewer bits to smaller indices) during the prediction mode encoding / decoding process.
[0279] When the composition of the candidate group in the above examples is not identical, the prediction mode index may be kept the same and added, or a different index may be assigned. For example, when the prediction mode candidate group index setting is the same as in the previous example, if the prediction mode of the luminance component is diagonal DL, the composition of the prediction mode candidate group increases by 1, and the prediction mode index of the existing candidate group is kept the same, and the index of the newly added mode may be placed at the end {diagonal DL (5) in this example} or a different index may be assigned {diagonal DL (0), Planar (1), DC (2), vertical (3), horizontal (4), diagonal UL (5) in this example}.
[0280] For example (2), if the prediction mode of the luminance component matches one of the candidate modes of the color difference component prediction mode, the composition of the candidate group is the same (no change in the mode of the candidate group), and if none match, the composition of the candidate group is not the same (at least one of the modes of the candidate group is replaced).
[0281] When the composition of the candidate group in the above examples is the same, the index of the prediction mode may be kept the same or assigned a different index. For example, when the index of the candidate group of the prediction mode for the color difference component is Planar (0), DC (1), Vertical (2), Horizontal (3), and Diagonal UL (4), when the prediction mode for the luminance component is Vertical, the composition of the candidate group of the prediction mode does not change, and the index of each prediction mode may be kept the same or assigned a different index {in this example, Vertical (0), Horizontal (1), Diagonal UL (2), Planar (3), DC (4). An example where the directional mode is placed first when the mode of the luminance component is directional, and the non-directional mode is placed first when the mode of the luminance component is non-directional. Not limited thereto}.
[0282] When the composition of the candidate group in the above examples is not identical, the index of the prediction mode may be kept the same for the unchanged mode and assigned the index of the replaced mode to the changed mode, or a different index (from the existing one) may be assigned to multiple prediction modes. For example, when the prediction mode candidate group index setting is the same as in the previous example, if the prediction mode of the luminance component is diagonal DL, one of the prediction mode candidate groups (diagonal UL in this example) may be replaced, and the prediction mode index of the existing candidate group may be kept as is and the index of the replaced mode may be assigned to the index of the newly added mode {e.g., diagonal DL (4)} or a different index may be assigned {diagonal DL (0), Planar (1), DC (2), vertical (3), horizontal (4) in this example}.
[0283] In the above description, an example was given of performing index reconstruction for the purpose of allocating fewer mode bits, but this is an example based on the encoding / decoding settings, and other cases are also possible. If there is no change in the index of the prediction mode, binarization can be performed to allocate fewer bits to a small index, or binarization can be performed to allocate bits regardless of the size of the index. For example, when the reconstructed prediction mode candidate group is Planar (0), DC (1), Vertical (2), Horizontal (3), and Diagonal DL (4), a setting can be made to allocate fewer mode bits compared to other prediction modes because Diagonal DL is a mode obtained from the luminance component, even though a large index is allocated.
[0284] The above prediction mode may be a mode that is supported regardless of the image type, or a mode whose support is determined based on some image types (e.g., a mode that is supported for image type I but not for image types P or B).
[0285] The content described in the above example is limited to this example only, and additional or other modified examples may be possible. Furthermore, the encoding / decoding settings described in the above example may be implicitly determined or may explicitly include relevant information in units such as video, video, sequence, picture, slice, tile, etc.
[0286] (Obtaining predicted values in the same color space vs. obtaining predicted values in different color spaces)
[0287] The in-screen prediction mode described through the above example was a description of a prediction mode related to a method of acquiring data for generating prediction blocks from adjacent areas within the same space at the same time (e.g., extrapolation, interpolation, averaging, etc.).
[0288] In addition, a prediction mode related to a method of acquiring data for generating prediction blocks from regions located in different spaces at the same time may be supported.
[0289] For example, a prediction mode regarding a method of acquiring data for generating a prediction block in another color space by utilizing the correlation between color spaces can serve as an example. In this case, the correlation between color spaces may refer to the correlation between Y and Cb, Y and Cr, and Cb and Cr, taking YCbCr as an example. That is, in the case of a color difference component (Cb or Cr), a restored block of the luminance component corresponding to the current block can be generated as the prediction block of the current block (color difference vs. luminance is the default setting of the example described below). Alternatively, a restored block of some color difference components (Cb or Cr) corresponding to the current block of some color difference components (Cr or Cb) can be generated as the prediction block of the said color difference component (Cr or Cb). At this time, a restored block in a different color space can be generated as a prediction block (i.e., no correction is performed), or a block obtained by considering the correlation between color spaces (for example, a correction is performed on the existing restored block. In P = a * R + b, a and b represent values used for correction, and R and P represent values obtained in a different color space and predicted values in the current color space, respectively) can be generated as a prediction block.
[0290] In this example, the explanation assumes that data obtained using the correlation of color spaces is used as the predicted value of the current block; however, it may also be possible to use such data as a correction value that modifies the existing predicted value of the current block (for example, a residual value from a different color space is used as the correction value; that is, if a different predicted value exists, that predicted value is corrected. Although the sum is ultimately a predicted value, this explanation is intended for detailed differentiation). Although the present invention assumes the former case, it is not limited thereto, and the same or modified application as a correction value may be possible.
[0291] The above prediction mode may be a mode that is supported regardless of the image type, or a mode whose support is determined based on some image types (e.g., a mode that is supported for image type I but not for image types P or B).
[0292] (The part that compares to obtain correlation information)
[0293] In the above example, the correlation information between color spaces (a, b, etc.) may explicitly include relevant information or may be implicitly obtained. In this case, the area being compared to obtain the correlation information may be 1) the current block of the color difference component and the corresponding area of the luminance component, or 2) the adjacent area of the current block of the color difference component (e.g., left, top, top-left, top-right, bottom-left blocks, etc.) and the adjacent area of the corresponding block of the luminance component. The former may be an example of an explicit processing, and the latter an implicit processing.
[0294] For example, at least one pixel value in each color space (wherein, the pixel value being compared)<Pixel Value> is one pixel in each color space <pixel>It may be a pixel value obtained from a single pixel, or a pixel value obtained from multiple pixels. In other words, it is a pixel value derived through a filtering process such as a weighted average. That is, the number of pixels referenced or used for a single pixel value used for comparison in each color space may be one pixel vs. one pixel, one pixel vs. multiple pixels, etc. In this case, the former may be the color space generating the prediction value, and the latter may be the referenced color space. The above example may be a case that occurs depending on the color format, or regardless of the color format, it may be possible to compare the pixel value of a single pixel representing the chrominance component with the corresponding pixel value of a single pixel representing the luminance component, and filter the pixel value of a single pixel representing the chrominance component with multiple pixels representing the luminance component.<a-tap separate 1D filter, b x c mask non-separable 2D filter, d-tap directional filter 등> It is possible to compare with the pixel values obtained in this way, and one of the two methods can be used depending on the encoding / decoding settings. Although examples of chrominance and luminance were explained above, chrominance <cb>and color difference <cr>Correlation information can be obtained through comparisons such as (changes such as the example of may be possible).
[0295] In the above example, when correlation information is implicitly obtained, the area being compared may be the nearest pixel line of the current block of the current color component (e.g., pixels included in p[-1,-1] ~ p[blk_width - 1, -1], p[-1,0] ~ p[-1, blk_height - 1]) and the corresponding pixel line of another color space, or multiple pixel lines of the current block of the current color component (e.g., pixels included in multiple pixel lines including p[-2, -2] ~ p[blk_width - 1, -2], p[-2, -1] ~ p[-2, blk_height - 1] in the above case) and the corresponding pixel line of another color space.
[0296] Specifically, assuming the color format is 4:2:0, for comparison with the pixel value of one pixel in the current color space (color difference in this example), the pixel value of one pixel at a pre-set position (selected from top-left, top-right, bottom-left, and bottom-right within 2 x 2 in this example) among the four corresponding pixels in another color space (luminance in this example) (one pixel of the color difference component corresponds to four pixels within 2 x 2 of the luminance component). Alternatively, for comparison with the pixel value of one pixel in the color difference space, a pixel value obtained by performing filtering on multiple pixels in the luminance space (e.g., at least two pixels among the corresponding 2 x 2 pixels) may be used.
[0297] In summary, the parameter information can be derived from the restored pixels of an adjacent region of the current block and the corresponding restored pixels of another color space. That is, based on correlation information, at least one parameter (e.g., a or b, a1, b1 or a2, b2, etc.) can be generated and used as a value to be multiplied or added to the pixels of the restored block in another color space (e.g., a, a1, a2 / b, b1, b2).
[0298] At this time, the comparison process can be performed after verifying the availability of the pixels being compared in the above example. For example, if an adjacent area is available, it can be used as a pixel for comparison, and if it is unavailable, it may be determined according to the encoding / decoding settings. For example, if a pixel in an adjacent area is unavailable, it may be excluded from the process for obtaining color space correlation information, or the unavailable area may be filled before being included in the comparison process, and this may be determined according to the encoding / decoding settings.
[0299] For example, this may be an example where an area containing pixels of at least one color space becomes unusable when excluded during the process of obtaining correlation information between color spaces. Specifically, this may be an example where pixels of one of the two color spaces become unusable or pixels of both color spaces become unusable, and this can be determined according to the encoding / decoding settings.
[0300] Alternatively, when a process for obtaining correlation information between color spaces is performed after filling unusable areas with data for comparison (or an operation similar to a reference pixel padding process), various filling methods may be used. For example, the area may be filled with a pre-set pixel value {e.g., the median of the bit depth, 1 << (bit_depth - 1), a value between the minimum and maximum values of the actual pixels of the image, the average of the actual pixels of the image, the median, etc.}, or the area may be filled with a value obtained by performing filtering on adjacent pixels or adjacent pixels (an operation similar to a reference pixel filtering process), or other methods may be possible.
[0301]
[0302] FIG. 10 shows an example of pixels compared between color spaces to obtain correlation information. For convenience of explanation, the explanation assumes a 4:4:4 color format. At this time, it is assumed that the process described below (i.e., including a conversion process based on the component ratio) is explained by considering the component ratio according to the color format.
[0303] R0 represents an example where both color space regions are available. Since both regions are available, pixels in that region can be used in the comparison process to obtain correlation information.
[0304] R1 represents an example where one of the two color space regions is unusable (in this example, the adjacent region of the current color space is usable, while the corresponding region of the other color space is unusable). The unusable region can be used in the comparison process after being filled using various methods.
[0305] R2 represents an example where one of the two color space regions is unusable (in this example, the adjacent region of the current color space is unusable, and the corresponding region of the other color space is usable). Because an unusable region exists on one side, the corresponding regions of both color spaces cannot be used in the comparison process.
[0306] R3 represents an example where both color space regions are unusable. The unusable regions can be filled using various methods and used in the comparison process.
[0307] R4 represents an example where both color space regions are unusable. Because unusable regions exist on both sides, the corresponding regions of both color spaces cannot be used in the comparison process.
[0308] In addition, unlike Fig. 10, various settings may be possible if the adjacent area of the current block or the corresponding area of another color space are all unusable.
[0309] For example, a and b can be assigned to pre-set values (in this example, a is 1 and b is 0). In this case, it may mean maintaining a mode to fill the current block's prediction block with data from a different color space. In addition, in this case, when performing prediction mode decoding, the setting of the probability of occurrence (or selection) or priority for the mode may be different from the existing case (for example, considering it as having a low probability of selection or placing it in a lower priority, etc. That is, a situation where it can be inferred that the accuracy of the prediction block obtained through this prediction mode is significantly lower because it is correlation information with very low accuracy, and thus it will not be selected as the optimal prediction mode).
[0310] For example, a mode that fills the current block's prediction block with data from a different color space may not be supported because there is no data to compare. That is, the above mode may be supported only if at least one available area exists. Additionally, in this case, when performing prediction mode encoding / decoding, it may be possible to set a configuration that allows or disallows another mode that replaces the corresponding mode. In the former case, the number of prediction mode candidates is maintained, while in the latter case, the number of prediction mode candidates is reduced.
[0311] The above example is not limited to this, and various variations are possible.
[0312] In the above example, cases where it is unusable may occur when the region is not fully encoded or decoded, or when it is located beyond the boundaries of an image (e.g., picture, slice, tile, etc.) (i.e., when the current block and the region are not contained within the same image). Additionally, cases where it is unusable may be added depending on the encoding settings (e.g., constrained_intra_pred_flag, etc. For example, when it is a P or B slice / type and the above flag is 1 and the encoding mode of the region is Inter).
[0313] In the example described below, when generating a prediction block for the current block using restored data from another color space after obtaining correlation information through a comparison of color spaces, the above constraints may occur. That is, if the corresponding region of the other color space corresponding to the current block is determined to be unusable as described above, the use of this mode may be restricted or impossible.
[0314] A predicted value for the current block can be generated using parameters representing correlation information between color spaces obtained through the above process and reconstructed data from another color space corresponding to the current block. In this case, the reconstructed data from another color space used for predicting the current block may be the pixel value of a pixel at a pre-set location or a pixel value obtained through a filtering process.
[0315] For example, in the case of 4:4:4, the pixel value of a corresponding pixel in the luminance space may be used to generate a predicted value for a single pixel in the chrominance space. Alternatively, to generate a predicted value for a single pixel in the chrominance space, a pixel value obtained by performing filtering on multiple pixels in the luminance space may be used (for example, pixels located in directions such as left, right, top, bottom, top-left, top-right, bottom-left, bottom-right, etc., centered on the corresponding pixel. For example, when applying 5-tap or 7-tap filters, this can be understood as cases where there are 2 or 3 pixels each to the left, right, top, and bottom centered on the corresponding pixel).
[0316] For example, in the case of 4:2:0, to generate a predicted value for one pixel in the color difference space, the pixel value of one pixel at a pre-set position (selected from top-left, top-right, bottom-left, and bottom-right) among the four corresponding pixels in the luminance space (one pixel of the color difference component corresponds to a 2 x 2 pixel of the luminance component) may be used. Alternatively, to generate a predicted value for one pixel in the color difference space, a pixel value obtained by performing filtering on multiple pixels in the luminance space (for example, at least two pixels among the corresponding 2 x 2 pixels or pixels located in directions such as left, right, top, bottom, top-left, top-right, bottom-left, and bottom-right centered on the 2 x 2 pixel) may be used.
[0317] In summary, a pixel value obtained by applying (multiplying, adding, etc.) a parameter representing the correlation information obtained through the above process to a pixel value obtained in a different color space can be obtained as a predicted value of a pixel in the current color space.
[0318] In the above examples, cases regarding certain color formats and pixel value acquisition processes were described, but the examples are not limited thereto and identical or modified examples may be possible in other cases.
[0319] The content described in (obtaining predicted values in the same color space vs. obtaining predicted values in a different color space) may also be applied to (fixed candidate set configuration vs. variable candidate set configuration). For example, if a predicted value cannot be obtained in a different color space, an alternative mode for that purpose may be included in the candidate set.
[0320] In the case of the prediction mode described above, as shown in the example above, relevant information (e.g., information on support status, parameter information, etc.) may be included in units such as video, sequence, picture, slice, tile, etc.
[0321] In summary, depending on the encoding / decoding settings, a prediction mode candidate group can be configured with a prediction mode (Mode A) related to a method of acquiring data for generating a prediction block from an adjacent area within the same space at the same time, or additionally, a prediction mode (Mode B) related to a method of acquiring data for generating a prediction block from an area located in a different space at the same time can be included in the prediction mode candidate group.
[0322] In the above example, a prediction mode candidate set can be configured using only Mode A or only Mode B, or a combination of Mode A and Mode B. In this regard, configuration information regarding the configuration of the prediction mode candidate set may be explicitly generated, or information regarding the configuration of the prediction mode candidate set may be implicitly determined in advance.
[0323] For example, it may have the same configuration regardless of some encoding / decoding settings (image type in this example) or individual configurations depending on some encoding / decoding settings (for example, for image type I, a prediction mode candidate set is configured using mode A, mode B_1 <color mode>, and mode B_2 <color copy mode>, for image type P, a prediction mode candidate set is configured using mode A and mode B_1, and for image type B, a prediction mode candidate set is configured using mode A and mode B_2, etc.).
[0324] In the present invention, the prediction mode candidate group for the luminance component is as shown in FIG. 9, and the prediction mode candidate group for the color difference component is described under the assumption that it consists of horizontal, vertical, and diagonal modes (Planar, DC, Color Mode 1, Color Mode 2, Color Mode 3, Color Radiation Mode 1, Color Radiation Mode 2, Neighbor Block Mode 1 (Left Block), Neighbor Block Mode 2 (Upper Block) in FIG. 9, but various other prediction mode candidate groups may be set.
[0325]
[0326] In an image encoding method according to an embodiment of the present invention, the intra-frame prediction may be configured as follows. The intra-frame prediction of the prediction unit may include a reference pixel configuration step, a prediction block generation step, a prediction mode determination step, and a prediction mode encoding step. Additionally, the image encoding device may be configured to include a reference pixel configuration unit, a prediction block generation unit, and a prediction mode encoding unit that implement the reference pixel configuration step, the prediction block generation step, the prediction mode determination step, and the prediction mode encoding step. Some of the aforementioned processes may be omitted, other processes may be added, or the order may be changed from the order described above.
[0327] In addition, in an image decoding method according to an embodiment of the present invention, the intra-frame prediction may be configured as follows. The intra-frame prediction of the prediction unit may include a prediction mode decoding step, a reference pixel configuration step, and a prediction block generation step. Furthermore, the image decoding device may be configured to include a prediction mode decoding unit, a reference pixel configuration unit, and a prediction block generation unit that implement the prediction mode decoding step, the reference pixel configuration step, and the prediction block generation step. Some of the aforementioned processes may be omitted, other processes may be added, or the order may be changed from the order described above.
[0328] In the prediction block generation step described above, intra-frame prediction may be performed on the unit of the current block (e.g., encoding block, prediction block, transformation block, etc.) or on the unit of a predetermined sub-block. To this end, a flag indicating whether the current block is divided into sub-blocks for intra-frame prediction to be performed may be used. The flag may be encoded and signaled by an encoding device. If the flag is a first value, the current block is divided into multiple sub-blocks; otherwise, the current block is not divided into multiple sub-blocks. The division here may be an additional division performed after the aforementioned tree structure-based division. Sub-blocks belonging to the current block share a single intra-frame prediction mode, but different reference pixels may be configured for each sub-block. Alternatively, the sub-blocks may use the same intra-frame prediction mode and reference pixels. Alternatively, the sub-blocks may use the same reference pixels, but different intra-frame prediction modes may be used for each sub-block.
[0329] The above division may be performed in a vertical or horizontal direction. The division direction may be determined based on a flag signaled by the encoding device. For example, if the flag is a first value, it may be divided horizontally, and otherwise, it may be divided vertically. Alternatively, the division direction may be determined based on the size of the current block. For example, if the height of the current block is greater than a predetermined threshold size, it may be divided horizontally, and if the width of the current block is greater than a predetermined threshold size, it may be divided vertically. Here, the threshold size may be a fixed value pre-defined in the encoding / decoding device, or it may be determined based on information regarding block size (e.g., the size of the maximum convert block, the size of the maximum encoding block, etc.). Information regarding block size may be signaled at at least one level among sequence, picture, slice, tile, brick, or CTU row.
[0330] The number of sub-blocks can be variably determined based on the size, shape, depth of division, and in-frame prediction mode of the current block. For example, if the current block is 4x8 or 8x4, the current block may be divided into 2 sub-blocks. Or, if the current block is greater than or equal to 8x8, the current block may be divided into 4 sub-blocks.
[0331] In this invention, the explanation will focus on the encoder, and since the decoder can be derived inversely from the contents of the encoder, a detailed explanation is omitted.
[0332]
[0333] FIG. 11 is an example diagram illustrating the configuration of reference pixels used for intra-frame prediction. The size and shape (M × N) of the current block on which prediction is performed can be obtained from the block partitioning unit, and the description is based on the assumption that the range of 4×4 to 128×128 is supported for intra-frame prediction. Intra-frame prediction may generally be performed in units of prediction blocks, but depending on the settings of the block partitioning unit, it may be performed in units such as encoding blocks or conversion blocks. After verifying the block information, the reference pixel configuration unit can configure the reference pixels used for prediction of the current block. At this time, the reference pixels are in temporary memory (e.g., array <array>It can be managed through 1st, 2nd arrays, etc., and is created and removed during each in-frame prediction process of the block, and the size of the temporary memory can be determined according to the configuration of the reference pixel.
[0334] In this example, the explanation assumes that the left, top, top-left, top-right, and bottom-left blocks are used for the prediction of the current block centered on the current block; however, this is not limited to this, and block candidate sets with other configurations may also be used for the prediction of the current block. For example, the candidate set of neighbor blocks for the reference pixel may be an example of following a raster or Z-scan, and depending on the scan order, some of the candidate sets may be removed or configured to include other block candidate sets (e.g., right, bottom, bottom-right blocks, etc. as additional configurations).
[0335] Alternatively, a block corresponding to the current block in a different color space (e.g., if the current block belongs to Cr, the other color space corresponds to Y or Cb) (e.g., having the same coordinates or corresponding coordinates based on the color component composition ratio in each color space) may be used for the prediction of the current block. Furthermore, for the sake of convenience of explanation, the description assumes an example in which a single block is formed at the aforementioned pre-set locations (left, top, top-left, top-right, bottom-left), but at least one block may exist at those locations. That is, multiple sub-blocks resulting from the block division of the corresponding block may exist at the aforementioned pre-set locations.
[0336] In summary, an adjacent area of the current block may serve as the location of the reference pixel for in-frame prediction of the current block, and depending on the prediction mode, an area corresponding to the current block in a different color space may additionally be considered as the location of the reference pixel. In addition to the above examples, the location of the reference pixel may be determined according to the prediction mode, method, etc. For example, when generating a prediction block through a method such as block matching, the reference pixel location may be considered as the area of the current image prior to the current block where decoding / encoding has been completed, or an area included within the search range (e.g., to the left or top, or top-left or top-right of the current block) of the area where decoding / encoding has been completed.
[0337] As shown in FIG. 11, the reference pixels used for prediction of the current block may be composed of adjacent pixels of the left, top, top-left, top-right, and bottom-left blocks (Ref_L, Ref_T, Ref_TL, Ref_TR, Ref_BL in FIG. 11). In this case, while it is common for the reference pixels to be composed of pixels of the neighboring blocks closest to the current block (Fig. 11 a), it may also be possible to include other pixels (Fig. 11 b and other pixels of the outer lines). That is, at least one of a first pixel line (a) adjacent to the current block, a second pixel line (b) adjacent to the first pixel line, a third pixel line adjacent to the second pixel line, or a fourth pixel line adjacent to the third pixel line may be used. For example, depending on the encoding / decoding settings, the plurality of pixel lines may include all of the first to fourth pixel lines, or may include only the remaining pixel lines excluding the third pixel line. Alternatively, the plurality of pixel lines may include only the first pixel line and the fourth pixel line.
[0338] The current block may perform intra-frame prediction by selectively referencing any one of the plurality of pixel lines. At this time, the selection may be performed based on an index (refIdx) signaled by the encoding device. Alternatively, any one of the plurality of pixel lines may be selectively used based on the size, shape, and subdivision type of the current block, whether the intra-frame prediction mode is a non-directional mode, the angle of the intra-frame prediction mode, etc. For example, if the intra-frame prediction mode is a Planar mode or a DC mode, the use of only the first pixel line may be restricted. Or, if the size (width or height) of the current block is less than or equal to a predetermined threshold value, the use of only the first pixel line may be restricted. Or, if the intra-frame prediction mode is greater than a predetermined threshold angle (or less than a predetermined threshold angle), the use of only the first pixel line may be restricted. The threshold angle may be the angle of the intra-frame prediction mode corresponding to Mode 2 and Mode 66 among the aforementioned prediction mode candidates.
[0339]
[0340] Meanwhile, pixels adjacent to the current block can be classified into at least one reference pixel hierarchy, such that the pixels closest to the current block are ref_0 {pixels with a pixel value difference of 1 from the boundary pixel of the current block, p(-1,-1) ~ p(2m-1,-1), p(-1,0) ~ p(-1,2n-1)}, the next adjacent pixels {pixel value difference of 2 from the boundary pixel of the current block, p(-2,-2) ~ p(2m,-2), p(-2,-1) ~ p(-2,2n)} are ref_1, and the next adjacent pixels {pixel value difference of 3 from the boundary pixel of the current block, p(-3,-3) ~ p(2m+1, -3), p(-3,-2) ~ p(-3,2n+1)} are ref_2, and so on. That is, reference pixels can be classified into multiple reference pixel hierarchys based on the distance between the boundary pixel of the current block and adjacent pixels.
[0341] In addition, the reference pixel hierarchy can be set differently for each adjacent neighbor block. For example, when using the current block and the block adjacent to the top as reference blocks, the reference pixel according to the ref_0 hierarchy can be used, and when using the block adjacent to the upper right as reference blocks, the reference pixel according to the ref_1 hierarchy can be used.
[0342] Here, the reference pixel set generally referenced when performing intra-frame prediction consists of pixels belonging to neighboring blocks adjacent to the current block in the bottom-left, left, top-left, top-right, and tiers ref_0 (pixels closest to the boundary pixel); unless otherwise specified below, it is assumed that these are the pixels. However, only pixels belonging to some of the aforementioned neighboring blocks may be used as the reference pixel set, or pixels belonging to two or more tiers may be used as the reference pixel set. Here, the reference pixel set or tier may be implicitly determined (pre-set by the decoding / encoding device) or explicitly determined (receive determinable information from the encoding device).
[0343] Here, the description is based on the premise that there are up to 3 supported reference pixel layers, but there may also be more values, and the number of reference pixel sets (or reference pixel candidates) based on the number of reference pixel layers and the location of referenceable neighbor blocks <I / P / B. 이때, 영상은 픽쳐, 슬라이스, 타일 등>may be set differently depending on the block size, shape, prediction mode, image type, color component, etc., and the relevant information may be included in units such as sequence, picture, slice, tile, etc.
[0344] The present invention is described under the premise that a lower index (increasing from 0 by 1) is assigned starting from the reference pixel layer closest to the current block, but is not limited thereto. In addition, the reference pixel configuration information described below may be generated under such index settings (such as binarization, which assigns a short bit to a small index when selecting one of a set of multiple reference pixels).
[0345] In addition, if there are two or more supported reference pixel layers, a weighted average can be applied using each reference pixel included in the two or more reference pixel layers.
[0346] For example, a prediction block can be generated using a reference pixel obtained by the weighted sum of pixels located in the ref_0 layer (nearest pixel layer) and the ref_1 layer (next pixel layer) of FIG. 11. In this case, the pixels to which the weighted sum is applied in each reference pixel layer may be integer pixels as well as fractional pixels depending on the prediction mode (e.g., prediction mode directionality). Additionally, a prediction block can be obtained by assigning weights (e.g., 7:1, 3:1, 2:1, 1:1, etc.) to the prediction block obtained using the reference pixel according to the first reference pixel layer and the prediction block obtained using the reference pixel according to the second reference pixel layer, respectively. In this case, the weight may be higher for prediction blocks according to reference pixel layers adjacent to the current block.
[0347] Generally, configuring the nearest pixel of a neighbor block as a reference pixel may be an example of this, but it is not limited to this and various cases may be possible (e.g., when ref_0 and ref_1 are selected as a reference pixel layer and a predicted pixel value is generated through ref_0 and ref_1 using a weighted sum, etc., i.e., an implicit case).
[0348] In addition, information related to the configuration of reference pixels (e.g., selection information for a reference pixel layer or set, etc.) may be configured (e.g., ref_1, ref_2, ref_3, etc.) excluding information that is preset (e.g., when the reference pixel layer is preset as ref_0), but is not limited thereto.
[0349] Through the above example, we have examined some cases regarding reference pixel configuration, which can be combined with various encoding / decoding information to determine the in-frame prediction settings. In this case, the encoding / decoding information may include image type, color components, the size and shape of the current block, and the prediction mode {type of prediction mode (directional, non-directional), direction of prediction mode (vertical, horizontal, diagonal 1, diagonal 2, etc.)}, and the in-frame prediction settings (reference pixel configuration settings in this example) can be determined based on the combination of the encoding / decoding information of neighboring blocks and the encoding / decoding information of the current block and neighboring blocks.
[0350]
[0351] FIG. 12 is an example diagram illustrating the range of reference pixels used for in-frame prediction. Specifically, it shows the range of reference pixels determined by the size and shape of the block, the configuration of the prediction mode (angle information of the prediction mode in this example), etc. In FIG. 12, the position indicated by the arrow represents the pixel used for prediction.
[0352] Referring to Fig. 12, pixels A, A', B, B', and C represent the pixels at the bottom right of the 8×2, 2×8, 8×4, 4×8, and 8×8 blocks, and the reference pixel range of each block can be determined through the pixels (AT, AL, BT, BL, CT, CL) used in the upper and left blocks for the prediction of the corresponding pixels.
[0353] For example, for pixels A and A' (rectangular block), a reference pixel may be located in the range p(0,-1) ~ p(9,-1), p(-1,0) ~ p(-1,9), and p(-1,-1); for pixels B and B' (rectangular block), in the range p(0,-1) ~ p(11,-1), p(-1,0) ~ p(-1,11), and p(-1,-1); and for pixel C (square block), in the range p(0,-1) ~ p(15,-1), p(-1,0) ~ p(-1,15), and p(-1,-1).
[0354] Based on the range information of the reference pixel obtained through the above process {e.g., P(-1,-1), P(M+N-1,-1), P(-1,N+M-1), etc.}, it can be used in an in-frame prediction process (e.g., reference pixel filtering, prediction pixel generation process, etc.). Furthermore, the reference pixel is not supported only in the above cases, and various other cases may be possible.
[0355]
[0356] The reference pixel configuration unit of the in-frame prediction may include a reference pixel generation unit, a reference pixel interpolation unit, a reference pixel filter unit, etc., and may be configured to include all or part of the above configurations.
[0357] In the reference pixel configuration section, available and unavailable reference pixels can be classified by checking the availability of the reference pixels. For example, if a block at a pre-set location (or a reference pixel candidate block) is available, that block can be used as a reference pixel; if it is unavailable, that block cannot be used as a reference pixel.
[0358] The usability of the above reference pixel is determined to be unusable if at least one of the following conditions is satisfied. For example, it may be determined to be unusable if any of the following conditions are satisfied: it is located outside the picture boundary; it does not belong to the same division unit as the current block (e.g., slice, tile, etc.); encoding / decoding is not complete; or its use is restricted according to encoding / decoding settings. In other words, if none of the above conditions are satisfied, it may be determined to be usable.
[0359] Additionally, the use of reference pixels can be restricted by the encoding / decoding settings. For example, the use of reference pixels may be restricted depending on whether constrained intra-pred (e.g., constrained_intra_pred_flag) is performed. Constrained intra-pred can be performed to prohibit the use of blocks restored by referencing from other images as reference pixels when attempting to perform encoding / decoding that is error-robust to external factors such as the communication environment.
[0360] When the above-mentioned restricted intra-pred is disabled (e.g., constrained_intra_pred_flag = 0 in I image type or P or B image type), all reference pixel candidate blocks may be available, and when it is enabled (e.g., constrained_intra_pred_flag = 1 in P or B image type), whether the reference pixel of the block is used may be determined according to the encoding mode (Intra or Inter) of the reference pixel candidate block. That is, if the encoding mode of the block is Intra, it may be available regardless of whether the restricted intra-pred is enabled, and if it is Inter, it may be available (disabled) or unavailable (enabled) depending on whether the restricted intra-pred is enabled.
[0361] Additionally, limited in-frame prediction may be applied based on the encoding mode of the restored block corresponding to the current block in a different color space. For example, if the current block belongs to some chrominance components (Cb, Cr), availability may be determined based on the encoding mode of the block where the encoding / decoding of the luminance component (Y) corresponding to the current block is completed. The above example may apply when a restored block in a different color space is used as a reference pixel. It may also apply when the encoding mode is determined independently of the color space.
[0362] At this time, if the reference pixel candidate block has been encoded / decoded by some prediction method (e.g., predicted by block matching or template matching in the current image), the use of the reference pixel may be determined according to the encoding / decoding settings.
[0363] For example, when encoding / decoding is performed using the above prediction method, if the corresponding encoding mode is set to Intra, the block may be set to available. Alternatively, an exceptional case may be allowed where the block is left in an unavailable state despite being Intra.
[0364] For example, when encoding / decoding is performed using the above prediction method, if the corresponding encoding mode is set to Inter, the block may be designated as unusable. Alternatively, an exceptional case may be allowed where it is left in a usable state despite being Inter.
[0365] In other words, whether or not to allow for exceptions in cases where usage is determined by the encoding mode can be determined by the encoding / decoding settings.
[0366] Limited in-frame prediction may be a setting applied to some image types (e.g., P or B slice / tile types, etc.).
[0367] Reference pixel candidate blocks can be classified into cases where all are available, some are available, or all are unavailable based on reference pixel availability. In all cases except when all are available, reference pixels at the locations of unavailable candidate blocks can be filled or generated.
[0368] If a reference pixel candidate block is available, the pixel at a pre-set location within that block (assuming in this example that it is a pixel adjacent to the current block) can be included in the reference pixel memory of the current block. In this case, the pixel data at the block location can be copied as is or included in the reference pixel memory through a process such as reference pixel filtering.
[0369] If the reference pixel candidate block is unavailable, the pixel obtained through the reference pixel generation process can be included in the reference pixel memory of the current block.
[0370] In summary, if the reference pixel candidate block is available, the reference pixel can be constructed, and if the reference pixel candidate block is unavailable, the reference pixel can be created.
[0371] The following shows examples of filling reference pixels in unusable block locations using various methods.
[0372] For example, a reference pixel may be generated using an arbitrary pixel value, and it may be a single pixel value belonging to a pixel value range (e.g., a value derived from a minimum value, maximum value, median value, etc., in a pixel value adjustment process based on bit depth or a pixel value adjustment process based on image pixel value range information). More specifically, this may be an example applied when all reference pixel candidate blocks are unavailable.
[0373] Alternatively, reference pixels may be generated from an area where the encoding / decoding of the image is completed. Specifically, reference pixels may be generated from at least one available block adjacent to an unusable block. In this case, at least one of methods such as extrapolation, interpolation, or copying may be used, and the direction of reference pixel generation (or copying, extrapolation) may be clockwise or counterclockwise, and may be determined according to the encoding / decoding settings. For example, the direction of reference pixel generation within a block may follow a pre-set direction or a direction adaptively determined based on the location of the unusable block. Alternatively, for an area corresponding to the current block in a different color space, the same method as the above example may be used. The difference is that while the process in the current color space involves filling adjacent reference pixels of the current block, in the case of another color space it involves filling the block (M x N) corresponding to the current block (m x n). Therefore, the corresponding region (which may be referred to as reference pixels in this example) can be generated using various methods, including the method mentioned above (e.g., extrapolation in directions such as vertical, horizontal, and diagonal of surrounding pixels; interpolation such as Planar; averaging, etc. In this case, the filling direction is from the surrounding pixels of the block corresponding to the current block toward the interior of the block corresponding to the current block). This example may be a case where a prediction mode for generating a prediction block from another color space is included in the candidate group rather than being excluded.
[0374] In addition, after completing the reference pixel configuration through a process of verifying the availability of reference pixels, fractional reference pixels can be generated through linear interpolation of the reference pixels. Alternatively, the reference pixel interpolation process can be performed after the reference pixel filtering process. Or, only the filtering process for the configured reference pixels may be performed. In summary, this can be performed prior to the prediction block generation process.
[0375] At this time, interpolation is not performed in horizontal, vertical, and some diagonal modes (e.g., Diagonal down right, Diagonal down left, Diagonal up right), as well as in non-directional mode, color mode, and color radiation mode, while interpolation can be performed in other modes (other diagonal modes).
[0376] The interpolation precision can be determined by the supported prediction mode candidate group (or the total number of prediction modes), prediction mode configuration (e.g., prediction mode direction angle, prediction mode interval), etc.
[0377] For fractional reference pixel interpolation, one pre-set filter (e.g., a 2-tap linear interpolation filter) can be used, and one of a plurality of filter candidates (e.g., a 4-tap cubic filter, a 4-tap Gaussian filter, a 6-tap Wiener filter, an 8-tap Kalman filter, etc.) can be used.
[0378] When using one of multiple filter candidates, filter selection information can be explicitly generated or implicitly determined, and can be determined according to encoding / decoding settings (e.g., interpolation precision, block size, shape, prediction mode, etc.).
[0379] For example, the interpolation filter used can be determined based on the range of block sizes, the interpolation filter used can be determined based on interpolation precision, and the interpolation filter can be determined based on the characteristics of the prediction mode (e.g., directional information, etc.).
[0380] In detail, depending on the block size range, a pre-set interpolation filter (a) may be used in some range (A) and a pre-set interpolation filter (b) may be used in some range (B), one interpolation filter (c) among multiple interpolation filters (C) may be used in some range (C), one interpolation filter (d) among multiple interpolation filters (C) may be used in some range (D), and a pre-set interpolation filter may be used in some range and one of multiple interpolation filters may be used in some range. In this case, the use of one interpolation filter is implicit, and the use of one interpolation filter among multiple interpolation filters is explicit, and the block size dividing the block size range may be M x N (in this example, M and N are 4, 8, 16, 32, 64, 128, etc. That is, M and N can be the minimum or maximum value of each block size range).
[0381] The above interpolation-related information may be included in units such as video, sequence, picture, slice, tile, or block. The interpolation process may be performed in the reference pixel composition unit or may be a process performed in the prediction block generation unit.
[0382] Additionally, filtering may be performed on the reference pixel after configuring it to improve prediction accuracy by reducing remaining degradation through the encoding / decoding process; the filter used in this case may be a low-pass filter. Whether or not filtering is applied may be determined by the encoding / decoding settings, and if filtering is applied, fixed or adaptive filtering may be applied. The encoding / decoding settings may be defined according to the block size, shape, prediction mode, etc.
[0383] Fixed filtering means that a single filter is applied to the reference pixel filter section, and adaptive filtering means that one of a plurality of filters is applied to the reference pixel filter section. In this case, in the case of adaptive filtering, one of the plurality of filters may be implicitly determined or selection information may be explicitly generated depending on the encoding / decoding settings, and the filter candidate group may include filters such as 3-tap (e.g., [1,2,1] / 4) or 5-tap (e.g., [2,3,6,3,2]).
[0384] For example, filtering may not be applied in some settings (block range A).
[0385] For example, filtering is not applied in some settings (block range B, some mode C), and filtering can be applied through a pre-configured filter (3-tap filter) in some settings (block range B, some mode D).
[0386] For example, filtering is not applied in some settings (block range E, some mode F), filtering can be applied through a pre-configured filter (3-tap filter) in some settings (block range E, some mode G), filtering can be applied through a pre-configured filter (5-tap filter) in some settings (block range E, some mode H), and filtering can be applied by selecting one of multiple filters in some settings (block range E, some mode I).
[0387] For example, in some settings (block range J, some mode K), filtering can be applied through a pre-configured filter (5-tap filter), and additionally, filtering can be applied through a pre-configured filter (3-tap filter). That is, multiple filtering processes can be performed. More specifically, additional filtering can be applied based on the results of the preceding filtering.
[0388] In the above example, the block size dividing the block size range may be M x N (in this example, M and N are 4, 8, 16, 32, 64, 128, etc. That is, M and N can be the minimum or maximum values of each block size range). Additionally, the prediction mode can be broadly classified into directional mode, non-directional mode, color mode, color radiation mode, etc., and in detail can be classified into horizontal or vertical mode / diagonal mode (45-degree interval) / mode 1 adjacent to the horizontal or vertical mode, mode 2 adjacent to the horizontal or vertical mode (with a slightly wider interval than the previous one), etc. That is, depending on the mode classified as above, the application of filtering and the type of filtering can be determined.
[0389] Furthermore, although the above example illustrates a case where adaptive filtering is applied based on multiple factors such as block range and prediction mode, such multiple factors are not always required, and an example in which adaptive filtering is performed based on at least one factor is also possible. Additionally, the above example is not limited to various variations, and information related to the reference pixel filter may be included in units such as video, sequence, picture, slice, tile, or block.
[0390] The aforementioned filtering may be performed selectively based on a predetermined flag. Here, the flag may indicate whether filtering is performed on a reference pixel for intra-frame prediction. The flag may be encoded and signaled by an encoding device. Alternatively, the flag may be derived by a decoder based on the encoding parameters of the current block. The encoding parameters may include at least one of the location / region of the reference pixel, block size, component type, whether intra-frame prediction is applied on a sub-block basis, and an intra-frame prediction mode.
[0391] For example, if the reference pixel of the current block is a first pixel line adjacent to the current block, filtering may be performed on the reference pixel, and if not, filtering may not be performed on the reference pixel. Alternatively, if the number of pixels belonging to the current block is greater than a predetermined threshold number, filtering may be performed on the reference pixel, and if not, filtering may not be performed on the reference pixel. The threshold number is a value pre-agreed upon by the encoding / decoding device and may be an integer of 16, 32, 64, or greater. Alternatively, if the current block is greater than a predetermined threshold size, filtering may be performed on the reference pixel, and if not, filtering may not be performed on the reference pixel. The threshold size may be expressed as M x N and is a value pre-agreed upon by the encoding / decoding device, where M and N may be integers of 8, 16, 32, or greater. The threshold number or threshold size may be set as a condition determining whether to perform filtering on the reference pixel by combining one or both elements. Alternatively, if the current block is a luminance component, filtering may be performed on the reference pixel, and if not, filtering may not be performed on the reference pixel. Alternatively, if the current block does not perform the aforementioned sub-block unit intra-frame prediction (i.e., is not divided into multiple sub-blocks of the current block), filtering may be performed on the reference pixel, and if not, filtering may not be performed on the reference pixel. Alternatively, if the intra-frame prediction mode of the current block is a non-directional mode or a predetermined directional mode, filtering may be performed on the reference pixel, and if not, filtering may not be performed on the reference pixel. Here, the non-directional mode may be a Planar mode or a DC mode. However, among the non-directional modes, the DC mode may be restricted so that filtering of the reference pixel is not performed. The above-mentioned directional mode may refer to an intra-frame prediction mode that references an integer pixel.For example, the directional mode may include at least one of the in-frame prediction modes corresponding to modes -14, -12, -10, -6, 2, 18, 34, 50, 66, 72, 78, and 80 shown in FIG. 9. However, the directional mode may be limited so as not to include the horizontal mode and the vertical mode corresponding to modes 18 and 50, respectively.
[0392] When filtering is performed on a reference pixel according to the above flag, filtering may be performed based on a pre-defined filter in the encoding / decoding device. The number of taps in the filter may be 1, 2, 3, 4, 5, or more. The number of filter taps may be determined variably depending on the position of the reference pixel. For example, a 1-tap filter may be applied to a reference pixel corresponding to at least one of the bottom, top, leftmost, or rightmost part of the pixel line, and a 3-tap filter may be applied to the remaining reference pixels. Additionally, the filter strength may be determined variably depending on the position of the reference pixel. For example, a filter strength of s1 may be applied to a reference pixel corresponding to at least one of the bottom, top, leftmost, or rightmost part of the pixel line, and a filter strength of s2 may be applied to the remaining reference pixels (s1 <s2). 상기 필터 강도는 부호화 장치에서 시그날링될 수도 있고, 전술한 부호화 파라미터에 기초하여 결정될 수도 있다. 참조 화소에 n-tap 필터가 적용되는 경우, 현재 참조 화소 및 (n-1)개의 주변 참조 화소에 필터가 적용될 수 있다. 상기 주변 참조 화소는, 현재 참조 화소의 상단, 하단, 좌측 또는 우측 중 적어도 하나의 방향에 위치한 화소를 의미할 수 있다. 상기 주변 참조 화소는, 현재 참조 화소와 동일한 화소 라인에 속할 수도 있고, 주변 참조 화소의 일부는 현재 참조 화소와 상이한 화소 라인에 속할 수도 있다.
[0393] For example, if the current reference pixel is located to the left of the current block, the surrounding reference pixel may be a pixel adjacent in at least one of the directions above or below the current reference pixel. Alternatively, if the current reference pixel is located to the top of the current block, the surrounding reference pixel may be a pixel adjacent in at least one of the directions to the left or right of the current reference pixel. Alternatively, if the current reference pixel is located to the top-left of the current block, the surrounding reference pixel may be a pixel adjacent in at least one of the directions below or to the right of the current reference pixel. The ratio between the coefficients of the filter may be [1:2:1], [1:3:1], or [1:4:1].
[0394] A prediction block generation unit may generate a prediction block according to at least one prediction mode, and may use a reference pixel based on the in-frame prediction mode. At this time, the reference pixel may be used in a method such as extrapolation (directional mode) depending on the prediction mode, or in a method such as interpolation, averaging (DC), or copying (non-directional mode). Meanwhile, as described above, the current block may use a filtered reference pixel or an unfiltered reference pixel.
[0395]
[0396] Figure 13 is a figure showing the current block and adjacent blocks in relation to the generation of predicted blocks.
[0397] For example, in the case of a directional mode, the reference pixels of the bottom-left + left block (Ref_BL, Ref_L in FIG. 13) may be used for the mode between horizontal and some diagonal mode (Diagonal up right. Horizontal is excluded and diagonal is included), the left block for the horizontal mode, the left + top-left + top block (Ref_L, Ref_TL, Ref_T in FIG. 13) for the mode between horizontal and vertical, the top block (Ref_L in FIG. 13) for the vertical mode, and the top + top-right block (Ref_T, Ref_TR in FIG. 13) for the mode between vertical and some diagonal mode (Diagonal down left. Vertical is excluded and diagonal is included). Alternatively, in the case of a non-directional mode, the reference pixels of the left and top blocks (Ref_L, Ref_T in FIG. 13) or the bottom-left, left, top-left, top, and top-right blocks (Ref_BL, Ref_L, Ref_TL, Ref_T, Ref_TR in FIG. 13) may be used. Alternatively, in the case of a mode utilizing the correlation of color spaces (color copy mode), a restored block of another color space (not shown in FIG. 12 but referred to as Ref_Col in the present invention; a Collocated reference meaning a block of another space at the same time) can be used as a reference pixel.
[0398] Reference pixels used for in-frame prediction can be classified into multiple concepts. For example, reference pixels used for in-frame prediction can be classified into a first reference pixel and a second reference pixel, where the first reference pixel may be a pixel directly used to generate the prediction value of the current block, and the second reference pixel may be a pixel indirectly used to generate the prediction value of the current block. Alternatively, the first reference pixel may be a pixel used to generate the prediction value of all pixels of the current block, and the second reference pixel may be a pixel used to generate the prediction value of some pixels of the current block. Alternatively, the first reference pixel may be a pixel used to generate the first prediction value of the current block, and the second reference pixel may be a pixel used to generate the second prediction value of the current block. Alternatively, the first reference pixel may be a pixel located in an area located at the starting point of the prediction direction of the current block (unconditionally), and the second reference pixel may be a pixel not located at the starting point of the prediction direction of the current block (unconditionally).
[0399] As shown in the example above, reference pixels can be distinguished through various definitions; however, depending on the prediction mode, there may be cases where they cannot be distinguished by some definitions. In other words, it should be noted that the definitions used to distinguish reference pixels may differ depending on the prediction mode.
[0400] The reference pixel described in the above example may be the first reference pixel, and additionally, a second reference pixel may be involved in generating the prediction block. For the mode between some diagonal modes (Diagonal up right, excluding horizontal and including diagonal), the reference pixel of the top-left + top-right block (Ref_TL, Ref_T, Ref_TR in FIG. 13) may be used; for the horizontal mode, the reference pixel of the top-left + top-right block (Ref_TL, Ref_T, Ref_TR in FIG. 13); for the vertical mode, the reference pixel of the top-left + left + bottom-left block (Ref_TL, Ref_L, Ref_BL in FIG. 13); and for the mode between the vertical and some diagonal modes (Diagonal down left, excluding vertical and including diagonal), the reference pixel of the top-left + left + bottom-left block (Ref_TL, Ref_L, Ref_BL in FIG. 13) may be used. In detail, the above reference pixel may be used as the second reference pixel. In addition, a prediction block can be generated using a first reference pixel or a first and second reference pixel in non-directional mode and color radiation mode.
[0401] In addition, the second reference pixel may be considered to include not only pixels that have completed encoding / decoding but also pixels within the current block (predicted pixels in this example). That is, the first predicted value may be a pixel used to generate the second predicted value. Although the present invention is described primarily with an example in which pixels that have completed encoding / decoding are considered as the second reference pixel, it is not limited thereto, and variations using pixels that have not completed encoding / decoding (predicted pixels in this example) may also be possible.
[0402] Generating or correcting a prediction block using multiple reference pixels may be performed for the purpose of compensating for the shortcomings of existing prediction modes.
[0403] For example, the directional mode is a mode used for the purpose of predicting the directionality of a block using some reference pixels (first reference pixels), but it may not accurately reflect changes within the block, which may result in a decrease in the accuracy of the prediction. In this case, if additional reference pixels (second reference pixels) are used to generate or correct the prediction block, the accuracy of the prediction can be increased.
[0404] To this end, the following examples will describe cases where prediction blocks are generated using various reference pixels such as those in the above examples, but are not limited to the above examples, and even if terms such as first and second reference pixels are not used, they can be derived and understood from the above definitions.
[0405] The setting for generating prediction blocks using additional reference pixels can be determined explicitly or set implicitly. In the explicit case, units may include video, sequence, picture, slice, tile, etc. The examples described below will explain cases of implicit processing, but are not limited thereto, and other variations (explicit cases or mixed usage cases) may be possible.
[0406] Prediction blocks can be generated in various ways depending on the prediction mode. Specifically, the prediction method can be determined based on the location of the reference pixel used by the prediction mode. Additionally, the prediction method can be determined based on the pixel location within the block.
[0407] Next, we look at the case for horizontal mode.
[0408] For example, when the left block is used as a reference pixel (Ref_L in FIG. 13), the nearest pixel of that block (1300 in FIG. 13) can be used (e.g., extrapolation) to generate a prediction block in the horizontal direction.
[0409] In addition, a prediction block can be generated (or corrected; generation is in the sense that it is involved in the final prediction value, while correction is in the sense that it is not involved in the entire pixel) by utilizing a reference pixel (Ref_T in FIG. 13) adjacent to the current block corresponding to the horizontal direction. Specifically, the prediction value can be corrected using the nearest pixel of the block (1310 in FIG. 13; additionally, 1320 and 1330 may also be considered), and pixel value change or slope information of the said pixels (e.g., reflected in pixel value change or slope information such as R0 - T0, T0 - TL, T2 - TL, T2 - T0, T2 - T1, etc.) can be reflected in the correction process.
[0410] At this time, the pixels to be corrected may be all pixels within the current block or limited to a portion (for example, they may be determined as individual pixels that do not have a specific shape or exist in irregular locations, or as pixels that have a certain shape such as a line, as described in the example below. For convenience of explanation, the example below will assume a pixel unit in line units). If the pixels to be corrected are limited to some pixels, they may be determined as at least one line unit corresponding to the prediction mode direction. For example, pixels corresponding to a to d may be included in the correction target, and additionally, pixels corresponding to e to h may be included in the correction target. Furthermore, the correction information obtained from adjacent pixels of the block may be applied identically regardless of the line position or may be applied differently on a line-by-line basis, and the correction information may be applied less as the distance from the adjacent pixels increases (for example, by setting a larger division value according to distance, such as L1 - TL, L0 - TL, etc.).
[0411] In this case, the pixels included in the correction target may have only one setting in a single image or can be adaptively determined according to various encoding / decoding elements.
[0412] For example, in cases where it is determined adaptively, the pixels subject to correction may be determined according to the block size. In blocks of 8 x 8 or less, no lines are corrected; in blocks of 8 x 8 or more and 32 x 32 or less, correction may be performed on a single pixel line; and in blocks larger than that, correction may be performed on two pixel lines. The definition of the block size range can be derived from the prior description of the present invention.
[0413] Alternatively, the pixels subject to correction may be determined according to the shape of the block (e.g., square, rectangular; more specifically, a horizontally elongated rectangle or a vertically elongated rectangle). For example, in the case of an 8 x 4 block, correction may be performed on two pixel lines (a to h in FIG. 13), and in the case of a 4 x 8 block, correction may be performed on one pixel line (a to d in FIG. 13). This may be because, in the case of an 8 x 4 block, if the horizontal mode is determined from a horizontally elongated block shape, the orientation of the current block may tend to be more dependent on the block above, whereas in the case of a 4 x 8 block, if the horizontal mode is determined from a vertically elongated block shape, the orientation of the current block may tend not to be significantly dependent on the block above. Furthermore, the opposite setting of the above may also be possible.
[0414] Alternatively, the pixels subject to correction may be determined according to the prediction mode. In horizontal or vertical modes, a pixel lines may be subject to correction, and in other modes, b pixel lines may be subject to correction. In some modes (e.g., non-directional mode DC, color radiation mode, etc.), the pixels subject to correction may be specified in pixel units (e.g., a to d, e, i, m) rather than in a rectangular shape as described above. A detailed explanation follows in the examples described below.
[0415] In addition to the above description, adaptive settings may be possible depending on additional encoding / decoding elements. Although the above description focused on the limited cases of horizontal mode, it is not limited to the examples mentioned above, and identical or similar settings may be applied in other modes. Furthermore, cases similar to the above examples may be possible depending on a combination of multiple elements rather than a single encoding / decoding element.
[0416] In the case of the vertical mode, a detailed explanation is omitted because it can be derived by applying only a different direction to the prediction method of the horizontal mode mentioned above. Additionally, the examples described below omit content that can be derived from what overlaps with the explanation in the horizontal mode.
[0417]
[0418] Next, we examine the case for Diagonal mode (Diagonal up right).
[0419] For example, when the left block and the bottom-left block are used as reference pixels (first reference pixel or main reference pixel; Ref_L, Ref_BL in FIG. 13), the nearest pixels of the blocks (1300 and 1340 in FIG. 13) can be used (e.g., extrapolation) to generate prediction blocks in a diagonal direction.
[0420] In addition, a prediction block can be generated (or corrected) by utilizing a reference pixel adjacent to the current block located at a position diametrically opposite in the diagonal direction (a second reference pixel or an auxiliary reference pixel; Ref_T, Ref_TR in FIG. 13). More specifically, the predicted value can be corrected using the nearest pixels of the block (1310 and 1330 in FIG. 13, and additionally 1320 may also be considered), and the weighted average of the auxiliary reference pixel and the primary reference pixel (for example, the weight may be obtained based on at least one of the distance difference between the predicted pixel and the primary reference pixel, or between the predicted pixel and the auxiliary reference pixel, respectively along the x-axis or y-axis). Examples of weights applied to the primary reference pixel and the auxiliary reference pixel may include 15:1 to 8:8. If there are two or more auxiliary reference pixels, the auxiliary reference pixels may have the same weight, such as 14:1:1, 12:2:2, 10:3:3, 8:4:4, etc., or different weights between the auxiliary reference pixels, such as 12:3:1, 10:4:2, 8:6:2, etc. In this case, the differently applied weights depend on how close the slope information, etc., is in the direction of the current prediction mode centered on the corresponding predicted pixel. It can be determined. That is, the slope between the corresponding predicted pixel and each auxiliary reference pixel can be reflected in the above correction process through checking which one is closer to the current prediction mode slope), etc.
[0421] At this time, the filtering may have only one setting in a single image or may be adaptively determined according to various encoding / decoding elements.
[0422] For example, in cases where it is determined adaptively, the pixels to which filtering is applied (e.g., number of pixels, etc.) may be determined based on the location of the pixel to be corrected. When the prediction mode is diagonal mode (mode 2 in this example) and the pixel to be corrected is c, it may be predicted using L3 (the first reference pixel in this example) and corrected using T3 (the second reference pixel in this example). That is, it may be a case where one first reference pixel and one second reference pixel are used for one pixel prediction.
[0423] Alternatively, if the prediction mode is a diagonal mode (mode 3 in this example) and the pixel to be corrected is b, the prediction is made using L1* (or L2*, the first reference pixel in this example) obtained by interpolating fractional pixels between L1 and L2, and can be corrected using T2 (the second reference pixel in this example) or can be corrected using T3. Alternatively, it can be corrected using T2 and T3, or can be corrected using T2* (or T3*) obtained by interpolating fractional pixels between T2 and T3 based on the directionality of the prediction mode. That is, for a single pixel prediction, one first reference pixel (assuming it is L1* in this example; if the pixels used directly are L1 and L2, it can be viewed as two pixels. Or, depending on the filter used to interpolate L1*, it can be viewed as two or more pixels) and two second reference pixels (assuming they are T2 and T3 in this example; if it is L1*, it can be viewed as one pixel) may be used.
[0424] In summary, for a single pixel prediction, at least one first reference pixel and at least one second reference pixel may be used, and this can be determined according to the prediction mode and the location of the predicted pixel.
[0425]
[0426] If the above-mentioned corrected pixels are applied only to some pixels, they may be determined in units of at least one horizontal or vertical line according to the in-frame prediction mode direction. For example, pixels corresponding to a, e, i, m or pixels corresponding to a to d may be included in the correction target, and additionally, pixels corresponding to b, f, j, n or pixels corresponding to e to h may be included in the correction target. Pixels may be corrected in units of horizontal lines for some diagonal modes (Diagonal up right) and in units of vertical lines for some diagonal modes (Diagonal down left), but are not limited thereto.
[0427] In addition to the above description, adaptive settings may be possible depending on additional encoding / decoding elements. Although the above description focused on the limited case of diagonal mode (Diagonal up right), it is not limited to the example above, and identical or similar settings may be applied to other modes. Furthermore, cases similar to the above example may be possible depending on a combination of multiple elements rather than a single encoding / decoding element.
[0428] In the case of the diagonal mode (Diagonal down left), a detailed explanation is omitted because it can be derived by applying only a different direction to the prediction method of the diagonal mode (Diagonal up right).
[0429]
[0430] Next, we examine the case for diagonal mode (Diagonal up left).
[0431] For example, when the left block, the upper left block, and the upper block are used as reference pixels (first reference pixels or main reference pixels; Ref_L, Ref_TL, Ref_T in FIG. 13), the nearest pixels of the blocks (1300, 1310, 1320 in FIG. 13) can be used (e.g., extrapolation) to generate prediction blocks in a diagonal direction.
[0432] Alternatively, a prediction block can be generated (or corrected) by utilizing a reference pixel adjacent to the current block located at a position corresponding to the diagonal direction (a second reference pixel or an auxiliary reference pixel. In FIG. 13, Ref_L, Ref_TL, Ref_T. The positions are the same as the main reference pixel). Specifically, the predicted value can be corrected using pixels other than the nearest pixel of the block (pixels located to the left of 1300 in FIG. 13, pixels located to the left, top, and top left of 1320, pixels located to the top of 1310, etc.), and the weighted average of the auxiliary reference pixel and the main reference pixel (for example, the ratio of weights applied to the main reference pixel and the auxiliary reference pixel may be 7:1 to 4:4, etc. If there are two or more auxiliary reference pixels, the auxiliary reference pixels may have the same weight as examples such as 14:1:1, 12:2:2, 10:3:3, 8:4:4, etc., or different weights between auxiliary reference pixels may be applied as examples such as 12:3:1, 10:4:2, 8:6:2, etc. In this case, the weights applied differently may be determined based on whether they are close to the main reference pixel) or linear extrapolation, etc.
[0433] If the pixels to be corrected are limited to some pixels, they may be determined in units of horizontal or vertical lines adjacent to the reference pixels used in the prediction mode. In this case, horizontal and vertical lines may be considered simultaneously and overlap may be allowed. For example, pixels corresponding to a to d and pixels corresponding to a, e, i, and m (where a overlaps) may be included in the correction target. Additionally, pixels corresponding to e to h and pixels corresponding to b, f, j, and n (where a, b, e, and f overlap) may be included in the correction target.
[0434]
[0435] Next, we examine the case for the non-directional mode (DC).
[0436] For example, when at least one of the left, top, top-left, top-right, and bottom-left blocks is used as a reference pixel, a prediction block can be generated by using the nearest pixel of the block (assumed to be pixels 1300 and 1310 in the example in FIG. 13) (e.g., average, etc.).
[0437] Alternatively, a prediction block can be generated (or corrected) by utilizing adjacent pixels of the above reference pixel (second reference pixel or auxiliary reference pixel. Ref_L, Ref_T in FIG. 13. In this example, the location is the same as the main reference pixel or includes pixels located next adjacent to the main reference pixel. Similar to the case of a diagonal up left). Specifically, the predicted value can be corrected using a pixel at a location identical or similar to the main reference pixel of the block, and can be reflected in the correction process through the weighted average of the auxiliary reference pixel and the main reference pixel (for example, the ratio of weights applied to the main reference pixel and the auxiliary reference pixel may be 15:1 to 8:8, etc. If there are two or more auxiliary reference pixels, the auxiliary reference pixels may have the same weight as examples such as 14:1:1, 12:2:2, 10:3:3, 8:4:4, etc., or different weights between auxiliary reference pixels may be 12:3:1, 10:4:2, 8:6:2, etc.).
[0438] At this time, the filtering may have only one setting in a single image or may be adaptively determined according to various encoding / decoding elements.
[0439] For example, in cases where it is determined adaptively, the filter may be determined based on the size of the block. For pixels located at the top-left, top, and left ends of the current block (assuming that in this example, the pixel at the top-left applies filtering to the pixels to the left and above it, the pixel at the top applies filtering to the pixels above it, and the pixel at the left ends applies filtering to the pixels to the left of it), some filtering settings may be followed in blocks of 16 x 16 or less (filtering applied with weight ratios of 8:4:4 and 12:4 in this example), and some filtering settings may be followed in blocks of 16 x 16 or more (filtering applied with weight ratios of 10:3:3 and 14:2 in this example).
[0440] Alternatively, the filter may be determined based on the shape of the block. For example, in the case of a 16 x 8 block, the pixel located at the top of the current block may follow certain filtering settings (assuming that in this example, filtering is applied to the pixels to the left, top, and right of the corresponding pixel; that is, it can be seen as an example where the pixels to which filtering is applied are changed; filtering is applied with a weighting ratio of 10:2:2:2), and the pixel located at the left end of the current block may follow certain filtering settings (assuming that in this example, filtering is applied to the pixel to the left of the corresponding pixel; filtering is applied with a weighting ratio of 12:4). This may be an example applied under the assumption that it may be dependent on multiple pixels above the block in a horizontally elongated block shape. Additionally, the opposite settings of the above may be possible.
[0441] If the pixels to be corrected are limited to some pixels, they may be determined in units of horizontal or vertical lines adjacent to the reference pixels used in the prediction mode. In this case, horizontal and vertical lines may be considered simultaneously and overlap may be allowed. For example, pixels corresponding to a to d and pixels corresponding to a, e, i, and m (where a overlaps) may be included in the correction target. Additionally, pixels corresponding to e to h and pixels corresponding to b, f, j, and n (where a, b, e, and f overlap) may be included in the correction target.
[0442] In addition to the above description, adaptive settings may be possible depending on additional encoding / decoding elements. Although the above description focused on limited cases regarding non-directional mode, it is not limited to the examples mentioned above, and identical or similar settings may be applied in other modes. Furthermore, cases similar to the above examples may be possible depending on a combination of multiple elements rather than a single encoding / decoding element.
[0443]
[0444] Next, we examine the case for color copy mode.
[0445] In the case of color radiation mode, prediction blocks are generated using a method different from that of the existing prediction mode, but they can be generated (or corrected) using reference pixels in the same or similar way. The details regarding the acquisition of prediction blocks can be derived through the examples described above and below, so they are omitted.
[0446] For example, a prediction block can be generated by using a block corresponding to a different color space from the current block as a reference pixel (first reference pixel or main reference pixel) (e.g., copying, etc.).
[0447] Alternatively, a prediction block can be generated (or corrected) by utilizing the reference pixels of the current block and adjacent blocks (second reference pixels or auxiliary reference pixels. Ref_L, Ref_T, Ref_TL, Ref_TR, Ref_BL in FIG. 13). In detail, the predicted value can be corrected using the nearest pixel of the block (assumed to be 1300 and 1310 in FIG. 13 in this example), and the weighted average of the auxiliary reference pixel and the main reference pixel (for example, the ratio of weights applied to the main reference pixel and the auxiliary reference pixel may be 15:1 to 8:8, etc. If there are two or more auxiliary reference pixels, the auxiliary reference pixels may have the same weight as examples such as 14:1:1, 12:2:2, 10:3:3, 8:4:4, etc., or different weights between auxiliary reference pixels may be 12:3:1, 10:4:2, 8:6:2, etc.) can be reflected in the correction process.
[0448] Alternatively, a prediction block can be generated (or corrected) by utilizing pixels of a block adjacent to a block obtained in a different color space (second reference pixels or auxiliary reference pixels. Assuming the figure in FIG. 13 is a block corresponding to the current block in a different color space, Ref_L, Ref_T, Ref_TL, Ref_TB, Ref_BL, and Ref_R, Ref_BR, Ref_B not shown in FIG. 13). Filtering can be applied to the pixel to be corrected and its surrounding pixels (for example, a first reference pixel in a different color space or a first reference pixel and a second reference pixel. That is, when weighted averaging, etc. are applied within a block, a first reference pixel to be corrected and a first reference pixel to which filtering is applied are required, and at the block boundary, a first reference pixel to be corrected and a first reference pixel to which filtering is applied and a second reference pixel are required) and reflected in the correction process.
[0449] In cases where the two above cases are combined, correction can be performed using pixels in the adjacent block of the current block as well as pixels within the prediction block obtained in a different color space, and filtering can be applied to the pixel to be corrected and its surrounding pixels (e.g., pixels in the adjacent neighbor block of the pixel to be corrected and pixels within the adjacent current block of the pixel to be corrected) (e.g., applying filtering by applying an M x N mask to the location of the pixel to be corrected. In this case, the mask can be applied to all or some pixels including the corresponding pixel and the pixels above, below, left, right, upper left, upper right, lower left, lower right, etc.) and reflected in the correction process.
[0450] This example illustrates a case where the predicted value of the current block is obtained in a different color space and filtering is applied thereafter; however, the predicted value of the current block may also be obtained after filtering has already been applied in that color space prior to obtaining the predicted value. It should be noted that in this case, the target to which filtering is applied remains the same, differing only in the order from the example above.
[0451] At this time, the filtering may have only one setting in a single image or may be adaptively determined according to various encoding / decoding elements.
[0452] As an example of a case where it is determined adaptively, the filtering settings may be determined according to the prediction mode. Specifically, adaptive filtering settings may be provided according to the detailed color copy mode among the color copy modes. For example, some color copy modes (in this example, a set of correlation information in the adjacent area of the current block and the adjacent area of the block corresponding to a different color space)<a와 b> In the case of obtaining ), some filtering settings <1> It can follow, and in some color copy modes (in this example, when multiple sets of correlation information are obtained by comparing with the above modes; i.e., a1 and b1, a2 and b2), some filtering settings <2> Can follow.
[0453] In the above filtering settings, whether or not filtering is applied can be determined. For example, filtering may be applied depending on the filtering settings, or <1> Filtering may not be applied. <2> . Or, you can use an A filter or <1> You can use a B filter. <2> . Or, filtering can be applied to all pixels to the left and above the current block, or <1> Filtering can be applied to some pixels on the left and top sides.
[0454] If the pixels being corrected are limited to some pixels, they may be determined in units of horizontal or vertical lines adjacent to the reference pixels used in the prediction mode (in this example, auxiliary reference pixels; different from the previous example). In this case, horizontal and vertical lines may be considered simultaneously and overlap may be allowed.
[0455] For example, pixels corresponding to a to d and pixels corresponding to a, e, i, and m (where a overlaps) may be included in the correction target. Additionally, pixels corresponding to e to h and pixels corresponding to b, f, j, and n (where a, b, e, and f overlap) may be included in the correction target.
[0456] In summary, the primary reference pixel used for generating the prediction block is obtained from a different color space, and the secondary reference pixel used for correcting the prediction block can be obtained from a block adjacent to the current block in the current color space. Additionally, it can be obtained from a block adjacent to the corresponding block in a different color space. Furthermore, it can be obtained from some pixels of the prediction block of the current block. That is, some pixels within the prediction block can be used for correcting some pixels within the prediction block.
[0457] In addition to the above description, adaptive settings may be possible depending on additional encoding / decoding elements. Although the above description focused on limited cases regarding non-directional mode, it is not limited to the examples mentioned above, and identical or similar settings may be applied in other modes. Furthermore, cases similar to the above examples may be possible depending on a combination of multiple elements rather than a single encoding / decoding element.
[0458] This example is one in which a predicted block of the current block is obtained using the correlation between color spaces, but since the block used to obtain that correlation is obtained from the adjacent area of the current block and the adjacent area of a corresponding block in a different color space, it can be understood as an example of applying filtering to block boundaries.
[0459]
[0460] Generating a prediction block using multiple reference pixels may result in various cases depending on the encoding / decoding settings. Specifically, depending on the encoding / decoding settings, it may be determined whether to support generating or correcting a prediction block using a second reference pixel.
[0461] For example, whether additional pixels are used in the prediction process may be determined implicitly or explicitly. In the case of explicit determination, the information may be included in units such as video, sequence, picture, slice, tile, or block.
[0462] For example, whether additional pixels are used in the prediction process may be applied to all prediction modes or to some prediction modes. In this case, some prediction modes may be at least one of horizontal, vertical, partial diagonal modes, non-directional modes, color radiation modes, etc.
[0463] For example, whether additional pixels are used in the prediction process can be applied to all blocks or to some blocks. In this case, some blocks may be defined according to the size, shape, etc. of the blocks, and the blocks may be M x N (for example, M and N have lengths such as 8, 16, 32, 64, etc. In the case of a square, 8 x 8, 16 x 16, 32 x 32, 64 x 64, etc. In the case of a rectangle, 2:1 rectangle, 4:1 rectangle, etc.).
[0464] Additionally, whether additional pixels are used in the prediction process may be determined based on certain encoding / decoding settings. In this case, the encoding / decoding setting may be constrained_intra_pred_flag, and additional reference pixels may be used restrictively in the prediction process according to the said flag.
[0465] For example, if the area containing the second reference pixel is restricted from use by the flag (i.e., assumed to be an area filled through a process such as reference pixel padding by the flag above), the use of the second reference pixel in the prediction process may be restricted. Alternatively, the second reference pixel may be used in the prediction process regardless of the flag.
[0466] In addition to the cases described in the examples above, various applications and variations, such as the combination of one or more elements, may be possible. Furthermore, although the examples above are described as some cases regarding color radiation modes, they may be examples that can be applied identically or with modifications to prediction modes that generate or correct prediction blocks using multiple reference pixels, in addition to color radiation modes.
[0467] Although the above example described a case where each prediction mode has a single setting for generating or correcting a prediction block using multiple reference pixels, multiple settings may be possible for each prediction mode. That is, a candidate set of filtering settings may consist of multiple options, and selection information regarding them may be generated.
[0468] In summary, information regarding whether filtering is performed may be processed explicitly or implicitly, and if filtering is performed, information regarding the filtering selection may be processed explicitly or implicitly. When such information is processed explicitly, it may be included in units such as video, sequence, picture, slice, tile, or block.
[0469] The prediction block generated above can be corrected, and the correction process of the prediction block will be examined below.
[0470] The above correction process may be performed based on a predetermined reference pixel and weights. In this case, the reference pixel and weights may be determined dependently on the position of the pixel within the current block to be corrected (hereinafter referred to as the current pixel). The reference pixel and weights may also be determined dependently on the in-frame prediction mode of the current block.
[0471] When the in-frame prediction mode of the current block is a non-directional mode, the reference pixels (refL, refT) of the current pixel belong to a first pixel line adjacent to the current block and may be located on the same horizontal / vertical line as the current pixel. The weights may include at least one of a first weight (wL) in the x-axis direction, a second weight (wT) in the y-axis direction, or a third weight (wTL) in the diagonal direction. The first weight may refer to a weight applied to the left reference pixel, the second weight may refer to a weight applied to the top reference pixel, and the third weight may refer to a weight applied to the top-left reference pixel. Here, the first and second weights may be determined based on the position information of the current pixel and a predetermined scaling factor (nScale). The scaling factor may be determined based on the width (W) and height (H) of the current block. For example, the first weight (wL[x]) of the current pixel (predSample[x][y]) can be determined as (32 >> ((x << 1) >> nScale)), and the second weight (wT[x]) can be determined as (32 >> ((y << 1) >> nScale)). The third weight (wTL[x][y]) can be determined as ((wL[x] >> 4) + (wT[y] >> 4)). However, if the intra-frame prediction mode is Planar mode, the third weight can be determined as 0. The scaling factor can be set as ((Log2(nTbW) + Log2(nTbH) - 2) >> 2).
[0472] When the in-frame prediction mode of the current block is vertical / horizontal mode, the reference pixels (refL, refT) of the current pixel belong to the first pixel line adjacent to the current block and may be located on the same horizontal / vertical line as the current pixel. When it is vertical mode, the first weight (wL[x]) of the current pixel (predSample[x][y]) is determined as (32 >> ((x << 1) >> nScale)), the second weight (wT[y]) is determined as 0, and the third weight (wTL[x][y]) may be determined as the same as the first weight. Meanwhile, in horizontal mode, the first weight (wL[x]) of the current pixel (predSample[x][y]) is determined to be 0, the second weight (wT[y]) is determined to be (32 >> ((y << 1) >> nScale)), and the third weight (wTL[x][y]) can be determined to be the same as the second weight.
[0473] When the in-frame prediction mode of the current block is a diagonal mode, the reference pixels (refL, refT) of the current pixel belong to the first pixel line adjacent to the current block and may be located on the same diagonal line as the current pixel. Here, the diagonal line has the same angle as the in-frame prediction mode of the current block. The diagonal line may refer to a diagonal line in the direction from the bottom left to the top right, or a diagonal line in the direction from the top left to the bottom right. At this time, the first weight (wL[x]) of the current pixel (predSample[x][y]) is determined as (32 >> ((x << 1) >> nScale)), the second weight (wT[y]) is determined as (32 >> ((y << 1) >> nScale)), and the third weight (wTL[x][y]) may be determined as 0.
[0474] If the in-frame prediction mode of the current block is equal to or smaller than mode 10, the reference pixels (refL, refT) of the current pixel belong to the first pixel line adjacent to the current block and may be located on the same diagonal line as the current pixel. Here, the diagonal line has the same angle as the in-frame prediction mode of the current block. At this time, the reference pixels may be restricted to using only one of the left or top reference pixels of the current block. The first weight (wL[x]) of the current pixel (predSample[x][y]) may be determined to be 0, the second weight (wT[y]) may be determined as (32 >> ((y << 1) >> nScale)), and the third weight (wTL[x][y]) may be determined to be 0.
[0475] If the in-frame prediction mode of the current block is equal to or greater than mode 58, the reference pixels (refL, refT) of the current pixel belong to the first pixel line adjacent to the current block and may be located on the same diagonal line as the current pixel. Here, the diagonal line has the same angle as the in-frame prediction mode of the current block. At this time, the reference pixels may be restricted to using only one of the left or top reference pixels of the current block. The first weight (wL[x]) of the current pixel (predSample[x][y]) may be determined as (32 >> ((x << 1) >> nScale)), the second weight (wT[y]) may be determined as 0, and the third weight (wTL[x][y]) may be determined as 0.
[0476] Based on the determined reference pixels (refL[x][y], refT[x][y]) and weights (wL[x], wT[y], wTL[x][y]), the current pixel (predSamples[x][y]) can be corrected as shown in the following Equation 1.
[0477] [Mathematical Formula 1]
[0478] predSamples[ x ][ y ] = clip1Cmp( ( refL[ x ][ y ] * wL[ x ] + refT[ x ][ y ] * wT[ y ] - p[ -1 ][ -1 ] * wTL[
[0479] However, the above-described correction process may be performed only when the current block does not perform in-frame prediction at the sub-block level. The above correction process may be performed only when the reference pixel of the current block is the first pixel line. The above correction process may be performed only when the in-frame prediction mode of the current block corresponds to a specific mode. Here, the specific mode may include at least one of a non-directional mode, a vertical mode, a horizontal mode, a mode smaller than a predetermined first threshold mode, and a mode larger than a predetermined second threshold mode. The first threshold mode may be 8, 9, 10, 11, or 12, and the second threshold mode may be 56, 57, 58, 59, or 60.
[0480]
[0481] In the prediction mode determination unit, a process is performed to select the optimal mode among a plurality of prediction mode candidates. Generally, the optimal mode in terms of encoding cost can be determined by using a rate-distortion technique that considers block distortion (e.g., distortion between the current block and the restored block, such as SAD (Sum of Absolute Difference), SSD (Sum of Square Difference), etc.) and the amount of bits generated according to the corresponding mode. The prediction block generated based on the prediction mode determined through the above process can be transmitted to the subtraction unit and the addition unit.
[0482]
[0483] The prediction mode encoding unit can encode a prediction mode selected through the prediction mode determination unit. It can encode index information corresponding to the prediction mode from a group of prediction mode candidates, or it can predict the prediction mode and encode information related thereto. That is, the former refers to a method of encoding the prediction mode as is without prediction, while the latter refers to a method of performing a prediction of the prediction mode to encode mode prediction information and information obtained based thereon. Furthermore, the former is an example where it can be applied to a chrominance component, and the latter to a luminance component, but it is not limited thereto and other cases may also be possible.
[0484] When encoding by predicting the prediction mode, the predicted value (or prediction information) of the prediction mode may be referred to as the Most Probable Mode (MPM). In this case, the MPM may be configured with a pre-set prediction mode (e.g., DC, Planar, Vertical, Horizontal, Diagonal mode, etc.) or a prediction mode of a spatially adjacent block (e.g., Left, Top, Top-left, Top-right, Bottom-left block, etc.). In this example, the diagonal mode refers to Diagonal up right, Diagonal down right, and Diagonal down left, and may be a mode corresponding to modes 2, 34, and 66 of FIG. 9.
[0485] Additionally, modes derived from modes already included in the MPM candidate group can be configured as MPM candidate groups. For example, in the case of directional modes among the modes included in the MPM candidate group, modes having a difference of mode interval a (e.g., a is a non-zero integer such as 1, -1, 2, -2, etc. In FIG. 9, if mode 10 is a mode already included, modes 9, 11, 8, 12, etc. correspond to the modes from which modes are derived) can be newly (or additionally) included in the MPM candidate group.
[0486] The above example may apply to cases where the MPM candidate group is composed of multiple modes, and the MPM candidate group (or number of MPM candidate groups) is determined according to the encoding / decoding settings (e.g., prediction mode candidate group, image type, block size, block shape, etc.) and may be configured to include at least one mode.
[0487] There may be a priority order for prediction modes for configuring the MPM candidate group. The order of prediction modes included in the MPM candidate group may be determined according to the above priority, and the configuration of the MPM candidate group can be completed when the number of prediction modes is filled according to the above priority. At this time, the priority may be determined in the order of prediction modes of spatially adjacent blocks, pre-configured prediction modes, and modes derived from prediction modes included first in the MPM candidate group, but other variations are also possible.
[0488] When performing prediction mode coding of the current block using an MPM, information regarding whether the prediction mode matches the MPM (e.g., most_probable_mode_flag) can be generated.
[0489] If it matches the MPM (e.g., most_probable_mode_flag = 1), additional MPM index information (e.g., mpm_idx) may be generated depending on the configuration of the MPM. For example, if the MPM is configured with a single prediction mode, no additional MPM index information is generated, and if it is configured with multiple prediction modes, index information corresponding to the prediction mode of the current block may be generated from the MPM candidate pool.
[0490] In cases where it does not match MPM (e.g., most_probable_mode_flag = 0), non-MPM index information (e.g., non_mpm_idx) corresponding to the prediction mode of the current block can be generated from the remaining prediction mode candidates (or non-MPM candidates) excluding the MPM candidates from the prediction mode candidates, which may be an example where non-MPM is composed of one group.
[0491] When a non-MPM candidate group is composed of multiple groups, information regarding which group the prediction mode of the current block belongs to can be generated. For example, non-MPM is composed of groups A and B (assuming A has m prediction modes, B has n prediction modes, and non-MPM consists of m+n prediction modes, with n being greater than m; assuming the modes of A are directional modes with equal intervals and the modes of B are directional modes without equal intervals), and if the prediction mode of the current block matches a prediction mode of group A (e.g., non_mpm_A_flag = 1), index information corresponding to the prediction mode of the current block can be generated from the group A candidate group, and if it does not match (e.g., non_mpm_A_flag = 0), index information corresponding to the prediction mode of the current block can be generated from the remaining prediction mode candidate group (or group B candidate group). As in the above example, non-MPM can be composed of at least one prediction mode candidate group (or group), and the non-MPM configuration can be determined according to the prediction mode candidate group. For example, if the number of prediction mode candidates is 35 or fewer, there may be 1, otherwise there may be 2 or more.
[0492] As in the example above, cases where non-MPM is composed of multiple groups may be supported for the purpose of reducing the mode bit amount when the number of predicted modes is large and the predicted modes are not predicted as MPMs.
[0493] When performing prediction mode encoding (or prediction mode decoding) of the current block using MPM, the binarization table applied to each prediction mode candidate group (e.g., MPM candidate group, non-MPM candidate group, etc.) can be generated individually, and the binarization method applied according to each candidate group can also be applied individually.
[0494] Prediction-related information generated through the prediction mode encoding unit can be transmitted to the encoding unit and recorded in the bitstream.
[0495]
[0496] FIG. 14 is an example of a tree-based block partitioning according to one embodiment of the present invention.
[0497] In Fig. 14, i represents quad tree partitioning, ii represents horizontal partitioning during binary tree partitioning, and iii represents vertical partitioning during binary tree partitioning. In the figure, A to C represent the initial block (block before partitioning; for example, assumed to be a Coding Tree Unit), and the number following the character above represents the number of each partition when partitioning is performed. In the case of a quad tree, numbers 0 to 3 are attached to the top-left, top-right, bottom-left, and bottom-right blocks, respectively, and in the case of a binary tree, numbers 0 and 1 are attached to the top-left and bottom-right blocks, respectively.
[0498] Referring to Fig. 14, the division state or information performed to obtain the corresponding block can be checked through the characters and numbers obtained through the division process.
[0499] For example, in i of FIG. 14, block A00 may be the top-left block (0 added to A0) among the four blocks obtained by performing quadtree division on the initial block A and performing quadtree division on A0, which is the top-left block (0 added to A) among the four blocks obtained as a result.
[0500] Alternatively, in ii of FIG. 14, block B10 may be the upper block (0 added to B1) among the two blocks obtained by performing a horizontal split during binary tree splitting on the initial block B, and then performing a horizontal split during binary tree splitting on B1, which is the lower block (1 added to B) among the two blocks obtained as a result.
[0501] Through the above splitting, the splitting status of each block, information (e.g., supported splitting settings <type of tree method, etc.>, block support range such as minimum and maximum sizes <specifically, support range according to the splitting method>, allowable splitting depth <specifically, support range according to the splitting method>, splitting flag <specifically, splitting flag according to the splitting method>, image type , encoding mode <intra inter>It is possible to know the partitioning status of the block, the decoding / encryption settings and information required to verify the information, etc., and to determine which block (parent block) it belonged to in the partitioning stage prior to obtaining the current block (child block). For example, in the case of block A31 in i of FIG. 14, neighboring blocks A30, A12, and A13 may exist, and it can be confirmed (A3) that A30 belonged to the same block as A31 in the partitioning stage immediately prior. In the case of A12 and A13, it can be confirmed that they belonged to a different block (A1) from A31 in the partitioning stage immediately prior (A3), and only in the stage before that did they belong to the same block (A).
[0502] The above example is for a single partitioning operation (four-way partitioning by a quad tree or horizontal / vertical partitioning by a binary tree), so the explanation will continue in the following example to explain cases where multiple tree-based partitioning is performed.
[0503]
[0504] FIG. 15 is an example of a plurality of tree-based block partitioning according to one embodiment of the present invention.
[0505] Referring to FIG. 15, the splitting state or information performed to obtain the corresponding block can be identified through the characters and numbers obtained through the splitting process. In this example, each character does not represent the initial block but indicates information about the split along with a number.
[0506] For example, in the case of block A1A1B0, it means that the upper right block (A1) was obtained when quad tree splitting was performed in the initial block, the upper right block (A1A1) was obtained when quad tree splitting was performed in block A1, and the upper block (A1A1B0) was obtained when horizontal splitting was performed during binary tree splitting in block A1A1.
[0507] Alternatively, in the case of block A3B1C1, it means that the lower right block (A3) was obtained when quad tree splitting was performed in the initial block, the lower block (B1 in A3) was obtained when horizontal splitting was performed during binary tree splitting in block A3, and the right block (C1 in A3B1) was obtained when vertical splitting was performed during binary tree splitting in block A3B1.
[0508] In this example as well, through the above-mentioned division information, information regarding the relationship between the current block and neighboring blocks can be checked, such as which division stage the neighboring blocks belong to, depending on the division stage of each block.
[0509] Figures 14 and 15 are examples of some methods for verifying partition information for each block. Various information and combinations of information (e.g., partition flag, depth information, maximum depth information, range of blocks, etc.) can be used to verify the partition information of each block, and thereby the relationship between blocks can be verified.
[0510]
[0511] Figure 16 is an example diagram showing various cases of block division.
[0512] Generally, because various texture information exists within an image, it is difficult to encode or decode using a single method. For example, some regions may contain strong edge components in a specific direction, while others may contain complex regions without edge components. Block partitioning plays an important role in efficiently encoding this.
[0513] Block partitioning can be performed for the purpose of efficiently dividing regions based on image characteristics. However, if only one type of partitioning method (e.g., quadtree partitioning) is used, it may be difficult to perform partitioning while properly reflecting the image characteristics.
[0514] Referring to FIG. 16, it can be seen that an image containing various textures is divided according to quad tree division and binary tree division. a to e may be cases where only quad tree division is supported, and f to j may be cases where only binary tree division is supported.
[0515] In this example, the partitions resulting from quad tree partitioning are referred to as UL, UR, DL, and DR, and binary tree partitioning is explained based on these.
[0516] Figure 16a may be a texture shape where quad tree partitioning yields optimal performance, and 4 partitions can be performed through a single quad tree partition (Div_HV) to perform decoding and encoding on a block-by-block basis. On the other hand, as shown in Figure 16f, when binary tree partitioning is applied, 3 partitions (2 Div_V and 1 Div_H) may be required compared to quad tree.
[0517] Figure 16b represents a case where the texture is divided into upper and lower regions of a block. In Figure 16b, when quad tree partitioning is applied, one partition (Div_HV) is required, and in Figure 16g, binary tree partitioning can also be performed once (Div_H). Assuming that the quad tree partitioning flag is 1 bit and the binary tree partitioning requires 2 bits or more, the quad tree may be considered efficient in terms of flag bits. However, since decoding information (e.g., information for expressing texture information (information for expressing texture < residual signal, encoding coefficient, information indicating the presence or absence of encoding coefficients, etc.>), prediction information <e.g., information related to intra-frame prediction, information related to inter-frame prediction>, transformation information <e.g., information on the type of transformation, information on the division of transformation, etc.>) occurs in each block unit, and in this case, the texture information is divided into similar regions, so the information may occur redundantly, it may not be as efficient as binary tree partitioning.
[0518] In other cases of Fig. 16, it can be determined which tree is more efficient depending on the type of texture. Through the above example, it can be seen that supporting variable block sizes and various division methods is important for efficient area partitioning based on image characteristics.
[0519]
[0520] Let's perform a detailed analysis of quad tree splitting and binary tree splitting.
[0521] Referring to b in Fig. 16, let's examine the blocks adjacent to the bottom right block (current block) (left and top blocks in this example). It can be seen that the left block is a block with characteristics similar to the current block, while the top block is a block with characteristics different from the current block. It may be a case where the blocks have similar characteristics to some blocks (left block in this example) but have been split due to the nature of quad tree partitioning.
[0522] In this case, the current block may generate similar encoding / decoding information due to characteristics similar to those of the left block. When intra-frame prediction modes (e.g., most probable mode, i.e., information predicting the current block's mode from neighboring blocks for the purpose of reducing the bit size of the current block's prediction mode) or motion information prediction modes (e.g., information intended to reduce mode bits, such as skip mode, merge mode, or contention mode) occur to efficiently reduce this, the information from the left block is efficient to reference. In other words, when referencing the encoding information of the current block (e.g., intra-frame prediction information, inter-frame prediction information, filter information, counting information, etc.) from either the left or the upper block, relatively accurate information (the left block in this example) can be referenced.
[0523] Referring to e in Fig. 16, let us assume a small block at the bottom right (current block, a block divided twice). Since the current block also has similar characteristics to the upper block among the left or upper blocks, coding performance can be improved by referencing coding information from that block (the upper block in this example).
[0524] In contrast, referring to j in Fig. 16, let us examine the left block adjacent to the rightmost block (the current block, a block split twice). Since there is no upper block in this example, if we examine only the left block, we can confirm that the current block has characteristics different from the left block. In the case of quad tree splitting, it may be similar to some of the neighboring blocks, whereas in the case of binary tree splitting, it can be confirmed that there is a high probability of having characteristics different from the neighboring blocks.
[0525] More specifically, in quad tree partitioning, when the pre-partition blocks of a block (x) and its neighbor block (y) are the same, the pre-partition block of a block (x) may have similar characteristics to or different characteristics from the neighbor block (y) that has the same characteristics, whereas in binary tree partitioning, the pre-partition block of a block (x) and its neighbor block (y) are mostly different characteristics. In other words, if they had similar characteristics, the block would not have been partitioned and would have been finalized as the pre-partition block; however, since they have different characteristics, it is highly likely that they were partitioned.
[0526] In addition to the above example, the bottom right block (x) and the bottom left block (y) in i of Fig. 16 may have such a relationship. That is, the upper block is partitioned without being divided because the image characteristics are similar, and the lower block is partitioned into the bottom left and bottom right blocks because the image characteristics are different.
[0527] In summary, in the case of quad tree partitioning, if there is a neighbor block adjacent to the current block whose parent block is identical to the current block, the quad tree is unconditionally partitioned horizontally or vertically in half; therefore, the characteristics may be similar to or different from those of the current block. In the case of binary tree partitioning, if there is a neighbor block adjacent to the current block whose parent block is identical to the current block, the partition can be horizontally or vertically depending on the image characteristics; thus, the fact that it is partitioned implies that it is partitioned because it has different characteristics.
[0528]
[0529] (Quad tree splitting)
[0530] Neighboring blocks that are identical to the current block and the block before splitting (assumed to be left and top in this example, but not limited thereto) may have characteristics similar to or different from the current block.
[0531] In addition, neighboring blocks that are not identical to the current block before splitting may have characteristics similar to or different from the current block.
[0532] (Binary Tree Splitting)
[0533] Neighboring blocks that are identical to the current block and the block prior to splitting (in this example, left or top; since it is binary, there is a maximum of one candidate) may have different characteristics.
[0534] In addition, neighboring blocks that are not identical to the current block before splitting may have characteristics similar to or different from the current block.
[0535] It is assumed that the above assumption is the main assumption of the present invention and will be described below. According to the above, the characteristics of the neighboring block can be classified into cases where the characteristics may be similar to or different from the current block (1) and cases where the characteristics may be different (2).
[0536]
[0537] Refer again to Fig. 15.
[0538] For example (where the current block is A1A2), among the adjacent blocks, block A0 (left block) can be classified as a general case (i.e., a case where it is unknown whether the characteristics are similar to or different from the current block) because the block before division of block A1A2 (A1) and the block before division of block A0 (initial block) are different.
[0539] The adjacent A1A0 block (upper block) checks the partitioning method because the pre-partition block (A1) of A1A2 block and the pre-partition block (A1) of A1A0 block are identical. In this case, since it is a quad tree partition (A), it is classified as a general case.
[0540] For example (where the current block is A2B0B1), the splitting method is checked because the A2B0B0 block (upper block) among the adjacent blocks is identical to the pre-split block (A2B0) of the A2B0B1 block. In this case, since it is a binary tree split (B), it is classified as an exceptional case.
[0541] For example (current block is A3B1C0), among adjacent blocks, block A3B0 (upper right block in this example) can be classified as a general case because the block before division of block A3B1C0 (A3B1) and the block before division of block A3B0 (A3) are different.
[0542] As shown in the example above, neighboring blocks can be classified into general and exceptional cases. In the general case, it is uncertain whether it is advisable to use the encoding information of the neighboring block as the encoding information for the current block, whereas in the exceptional case, it is strongly determined that using the encoding information of the neighboring block as the encoding information for the current block would be undesirable.
[0543] Based on the above classification, this can be applied to a method of obtaining prediction information for the current block from neighboring blocks.
[0544]
[0545] To summarize the above process,
[0546] Check if the pre-division blocks of the current block and the adjacent block are the same. (A)
[0547] If A yields the same result, check the partitioning method of the current block (since in the same case, not only the current block but also adjacent blocks are partitioned using the same partitioning method, check only the current block). (B)
[0548] If A yields a non-identical result, terminate. (End)
[0549] If B is a quad tree split, mark adjacent blocks as normal and terminate. (End)
[0550] If B results in a binary tree split, mark the adjacent block as an exception and terminate. (End)
[0551]
[0552] The above example can be applied to and configured for the in-screen prediction mode prediction candidate group setting (related to most probable mode, etc.) of the present invention. General prediction candidate group settings (e.g., candidate group configuration priority, etc.) and exceptional prediction candidate group settings may be supported, and depending on the state of the adjacent block above, the candidate group derived from that block may be pushed back in priority or excluded.
[0553] In this case, the adjacent blocks to which the above setting is applied may be limited to spatial cases (same space), or may also be applicable when derived from a different color space of the same image, such as in color copy mode. That is, such a setting can be established by considering the segmentation state of the block derived through color copy mode.
[0554] In addition to the above example, the following example may be an example in which encoding / decoding settings (e.g., setting a prediction candidate group, setting a reference pixel, etc.) are adaptively determined according to the relationship between blocks (the relative relationship between the current block and other blocks is identified in the above example as block partitioning information, etc.).
[0555]
[0556] This example (luminance component) is described under the assumption that there are 35 pre-defined intra-frame prediction modes in the encoding / decoding device, and that a total of 3 candidates from adjacent blocks (left and top in this example) constitute the MPM candidate group.
[0557] One candidate can be added to each of the left block (L0) and the top block (T0) to form a total of two candidates. If a candidate group cannot be formed in each block, it can be filled by replacing it with candidates such as DC, Planar, Vertical, Horizontal, and Diagonal modes. If two candidates are filled through the above process, the remaining one candidate can be filled by considering various cases.
[0558] For example, if the filled candidates in each block are identical, they can be replaced with adjacent modes of the above mode so as not to overlap with modes included in the candidate set (e.g., if the identical mode is k_mode, then k_mode-2, k_mode-2, k_mode+1, k_mode+2, etc.). Alternatively, if the filled candidates in each block are identical or not identical, the candidate set can be constructed by adding Planar, DC, Vertical, Horizontal, Diagonal modes, etc.
[0559] Through the above process, a set of intra-frame prediction mode candidates (general case) can be constructed. The candidates created in this way represent examples corresponding to the general case, and an adaptive intra-frame prediction mode candidate set (exceptional case) can be constructed based on adjacent blocks.
[0560] For example, if an adjacent block is marked as an exception, the candidate group derived from that block may be excluded. If the left block is marked as an exception, the candidate group may be composed of the upper block and a pre-configured prediction mode (e.g., DC, Planar, Vertical, Horizontal, Diagonal, mode derived from the upper block, etc.). The above example may be applied in the same or similar way even when the MPM candidate group consists of three or more MPM candidates.
[0561]
[0562] This example (chrominance component) is explained by assuming that there are 5 intra-frame prediction modes {DC, Planar, Vertical, Horizontal, and Color modes in this example}, and that the prediction modes are adaptedly prioritized and sorted to form a candidate set for encoding / decoding (i.e., encoding / decoding directly without using MPM).
[0563] First, the highest priority is assigned to the color mode (index 0 in this example, 1 bit '0' is assigned), and the other modes (Planar, Vertical, Horizontal, DC in this example) are assigned lower priority (indexes 1 to 4 in this example, 3 bits each '100', '101', '110', '111' are assigned).
[0564] If the color mode matches one of the other prediction modes (DC, Planar, Vertical, Horizontal) in the candidate group, a pre-set prediction mode (e.g., Diagonal mode, etc.) can be assigned to the priority given to the matching prediction mode (indexes 1 through 4 in this example), and if it does not match, the candidate group configuration is completed immediately.
[0565] Through the above process, a candidate set for the in-screen prediction mode can be constructed. The example created in this way corresponds to a general case, and an adaptive candidate set can be constructed depending on the block that acquired the color mode.
[0566] In this example, it is assumed that among other color spaces that have acquired a color mode, if the block corresponding to the current block consists of a single block (i.e., an undivided state), it is considered a general case, and if it consists of multiple blocks (i.e., divided into two or more), it is considered an exceptional case. Although this differs from the example classified according to whether the parent block of the current block and the adjacent block is the same, the division method, etc. as described above, if the corresponding block in another color space consists of multiple blocks, it may be an example set under the assumption that it is more likely to have characteristics different from those of the current block. In other words, it can be understood as an example where the encoding / decoding settings are adaptively determined according to the relationship between blocks.
[0567] Based on the above assumption, for example, if a corresponding block is marked as an exception state, the prediction mode derived from that block may be relegated to a lower priority. In this case, one of the other prediction modes (Planar, DC, Vertical, Horizontal) may be assigned to a high priority, and other prediction modes and color modes not included in the above priority may be assigned a low priority.
[0568] The above examples are described under certain assumptions for the convenience of explanation, but are not limited thereto, and the same or similar applications may be possible to the various embodiments described above of the present invention.
[0569] In summary, candidate group A can be used if there is a block marked as an exception among adjacent blocks, and candidate group B can be used if there is a block marked as an exception. Although the above distinction is divided into two cases, it can be understood that the configuration of in-screen prediction candidate groups at the block level can be adaptive depending on the existence of an exception state and block location.
[0570]
[0571] Furthermore, the aforementioned example was described assuming a tree-based partitioning method among the partitioning methods, but is not limited thereto. Specifically, the aforementioned example may be described under the assumption that a block obtained using at least one tree-based partitioning method is set as an encoding block, and that prediction, transformation, etc. are performed directly without the block being divided into prediction blocks, transformation blocks, etc.
[0572] As another example of a partitioning setting, the encoding block is obtained through tree-based partitioning, and it may be an applicable example where at least one prediction block is obtained based on the obtained encoding block.
[0573] For example, it is assumed that the encoding block (2N x 2N) can be obtained through tree-based partitioning (quad tree in this example) and the prediction block is obtained through type-based partitioning (assuming that the supported candidate forms in this example are 2N x 2N, 2N x N, N x 2N, and N x N). In this case, when a single encoding block (assuming this block is the parent block in this example) is divided into multiple prediction blocks (assuming these blocks are child blocks in this example), the aforementioned exception state, etc., may be set between the prediction blocks as well.
[0574] In g of Fig. 16, if there are two prediction blocks (separated by thin lines) within an encoding block (thick solid line), the lower block is considered to have different characteristics from the upper block, so the encoding information of the upper block is not referenced or its priority is delayed.
[0575]
[0576] FIG. 17 illustrates an example of block partitioning according to an embodiment of the present invention. Specifically, it illustrates an example in which a basic encoding block (maximum encoding block, 8N x 8N) is obtained through quad tree-based partitioning to obtain an encoding block (hatched block, 2N x 2N), and the obtained encoding block is divided into at least one prediction block (2N x 2N, 2N x N, N x 2N, N x N) through type-based partitioning.
[0577] At this time, the setting of in-screen prediction mode candidates for cases where a rectangular block is obtained (2N x N, N x 2N) will be described later.
[0578]
[0579] Figure 18 shows various examples of setting up in-screen prediction mode candidates for a block where prediction information occurs (prediction block in this example, 2N x N).
[0580] This example (luminance component) is explained under the assumption that when there are 67 in-frame prediction modes, a total of 6 candidates from adjacent blocks (left, top, top-left, top-right, and bottom-left in this example) constitute the MPM candidate group.
[0581] Referring to FIG. 13, four candidates can be configured in the order of L3 - T3 - B0 - R0 - TL, and two candidates can be configured with a pre-set mode (e.g., Planar, DC). If the maximum number (6 in this example) is not met even with the above configuration, the configuration can include prediction modes derived from prediction modes already included in the candidate group (e.g., k_mode-2, k_mode-2, k_mode+1, k_mode+2, etc. in the case of k_mode), pre-set modes (e.g., vertical, horizontal, diagonal, etc.).
[0582] In this example, the candidate priority is assumed to be the order of the current block's left - top - bottom-left - top-right - top-left blocks for spatially adjacent blocks (specifically, the bottom of the left block and the right sub-block of the top block), and the order of the pre-configured modes is assumed to be Planar - DC - Vertical - Horizontal - Diagonal modes.
[0583] Referring to Fig. 18a, the current block (2N x N. PU0) can be configured with candidate sets in the order of l1 - t3 - l2 - tr - tl, just like the candidate set in the above example (other details are omitted as they are redundant). In this example, the adjacent block of the current block may be a block that has already been encoded / decoded (encoded block, i.e., a prediction block within another encoded block).
[0584] Unlike the above, if the current block's location corresponds to PU1, it may be possible to set at least one intra-frame prediction mode candidate group. Generally, since blocks adjacent to the current block are more likely to have characteristics similar to the current block, it may be advantageous to form the candidate group from that block (1). Meanwhile, it may be necessary to form the candidate group for purposes such as parallel processing of encoding / decoding (2).
[0585] In the case where the current block is PU1, the candidate group configuration can be configured in the order of l3 - c7 - bl - k - l1 as shown in Fig. 18 b, and in the candidate group configuration configuration as shown in Fig. 18 c, the candidate group can be configured in the order of l3 - bl - k - l1 as shown in Fig. 18 c, where k can be derived from tr, etc. The difference between the two examples above is whether the intra-frame prediction mode of the upper block is included in the candidate group. That is, in the former case, the intra-frame prediction mode of the upper block is included in the candidate group for the purpose of increasing the efficiency of intra-frame prediction mode encoding / decoding, and in the latter case, the intra-frame prediction mode of the upper block, which cannot be referenced because it is a state where encoding / decoding may not yet be completed, is excluded from the candidate group for the purpose of parallel processing, etc.
[0586]
[0587] Figure 19 shows various examples of setting up in-screen prediction mode candidates for a block where prediction information occurs (prediction block in this example, N x 2N).
[0588] This example (luminance component) is explained under the assumption that when there are 67 in-frame prediction modes, a total of 6 candidates from adjacent blocks (left, top, top-left, top-right, and bottom-left in this example) constitute the MPM candidate group.
[0589] Referring to FIG. 13, one candidate can be formed in the order of L3 - L2 - L1 - L0 (top block), one in the order of T3 - T2 - T1 - T0 (top block), and two candidates in the order of B0 - R0 - TL (top-left, top-right, bottom-left blocks), and two candidates can be formed with pre-set modes (e.g., Planar, DC). If the maximum number is not met even with the above configuration, the candidate can be configured by including a prediction mode derived from a prediction mode already included in the candidate group, a pre-set mode, etc. In this case, the priority of the candidate group configuration may be in the order of Left - Top - Planar - DC - Bottom-left - Top-right - Top-left blocks.
[0590] Referring to FIG. 19a, the current block (N x 2N. PU0) can be configured with one candidate in the order of l3 - l2 - l1 - l0, one candidate in the order of t1 - t0, and two candidates in the order of bl - t2 - tl, in the same way as the candidate group setting of the above example. In this example, the adjacent block of the current block may be a block that has already been encoded / decoded (encoded block, i.e., a prediction block within another encoded block).
[0591] Unlike the above, if the current block's position corresponds to PU1, at least one in-screen prediction mode candidate set may be possible. Configurations (1) and (2) in the example above may be possible.
[0592] In the case where the current block is PU1, in the candidate group configuration setting as in (1), one candidate can be configured in the order of c13 - c9 - c5 - c1 as in Fig. 19 b, one candidate in the order of t3 - t2, and two candidate groups in the order of k - tr - t1 (k can be derived from bl or c13, etc.), and in the candidate group configuration setting as in (2), one candidate can be configured in the order of t3 - t2, and two candidate groups in the order of k - tr - t1 (k can be derived from bl, etc.) as in Fig. 19 c. The difference between the two examples above is whether the candidate group of the in-screen prediction mode of the upper block is included. In other words, in the former case, the intra-frame prediction mode of the left block was included in the candidate pool to increase the efficiency of intra-frame prediction mode encoding / decoding, whereas in the latter case, the intra-frame prediction mode of the left block was excluded from the candidate pool for purposes such as parallel processing, as it could not be referenced because encoding / decoding might not yet be complete.
[0593]
[0594] In this way, the candidate group configuration can be determined according to the settings for the candidate group configuration. In this example, the candidate group configuration settings (in this example, the candidate group configuration settings for parallel processing) may be implicitly determined or the relevant information may be explicitly recorded in units such as video, sequence, picture, slice, tile, etc.
[0595] The above can be summarized as follows. It is assumed that relevant information is implicitly determined or explicitly generated.
[0596] Verify the configuration settings for the intra-frame prediction mode candidate pool during the initial encoding / decoding stage. (A)
[0597] If the verification result for A is a setting that allows referencing a previous prediction block within the same encoding block, configure by including the in-frame prediction mode of that block in the candidate pool. (End)
[0598] If the verification result for A is a setting that prohibits referencing the previous prediction block within the same encoding block, configure by excluding the in-frame prediction mode of that block from the candidate pool. (End)
[0599] The above examples are described under certain assumptions for the convenience of explanation, but are not limited thereto, and the same or similar applications may be possible to the various embodiments described above of the present invention.
[0600]
[0601] FIG. 20 illustrates an example of block partitioning according to an embodiment of the present invention. Specifically, it illustrates an example in which a basic encoding block (maximum encoding block) obtains an encoding block (hatched block, A x B) through binary tree-based partitioning (or multiple tree-based partitioning), and the obtained encoding block is set as a prediction block.
[0602] At this time, the setting of candidate sets for motion information prediction for the case where a rectangular block is obtained (A x B, A ≠ B) will be described later.
[0603]
[0604] Figure 21 shows various examples of setting up in-frame prediction mode candidates for a block where prediction information is generated (in this example, the encoding block, 2N x N).
[0605] This example (luminance component) is explained under the assumption that when there are 67 prediction modes within the screen, a total of 6 candidates from adjacent blocks (left, top, top-left, top-right, and bottom-left in this example) constitute the MPM candidate group, and when the non-MPM candidate group is composed of multiple groups (A and B in this example, assuming that A belongs to a mode with a higher probability of predicting the current block's prediction mode within the non-MPM candidate group), a total of 16 candidates are composed of Group A and a total of 45 candidates are composed of Group B.
[0606] At this time, Group A may include candidates classified according to certain rules (e.g., directional modes with equal spacing) among modes not included in the MPM candidate group, or candidates not included in the final MPM candidate group according to the priority of the MPM candidate group. Group B may consist of candidates from the MPM candidate group and non-MPM candidate group that were not included in Group A.
[0607] Referring to FIG. 21a, the current block (CU0 or PU0, 2N x N in size, assumed to be 2:1 aspect ratio) is l1 - t3 - planar - DC - l2 - tr - tl - l1 * - t3 * - l2 * - tr * - tl * A candidate group of 6 candidates can be formed in the order of - vertical - horizontal - diagonal modes. In the above example, * represents a mode derived from the prediction mode of each block (e.g., an additive mode such as +1, -1, etc.).
[0608] Meanwhile, referring to Fig. 21b, the current block (CU1 or PU1, size 2N x N) is l3 - c7 - planar - DC - bl - k - l1 - l3 * - c7 * - bl * - k * - l1 * - Six candidates can be formed into a candidate group in the order of vertical, horizontal, and diagonal modes.
[0609] In this example, when the same setting as (2) in FIG. 18 is applied, b in FIG. 21 is l3 - planar - DC - bl - k - l1 - c7 - l3 * - bl * - k * - l1 * - Vertical - Horizontal - Diagonal - c7 * As such, the prediction mode of the upper block may be excluded from the candidate pool or pushed back in priority and included in group A.
[0610] However, the difference from the case of Fig. 18 is that even though it is a rectangular block, it is divided into encoding units and is set directly into prediction units without additional division. Therefore, the candidate group configuration setting as in Fig. 18 may not be applied in this example {setting as in (2)}.
[0611] However, in the case of k in Fig. 21, since it is a location where decoding / encoding is not completed, it may be derived from an adjacent block where decoding / encoding is completed.
[0612]
[0613] Figure 22 shows various examples of setting up in-frame prediction mode candidates for a block where prediction information is generated (encoding block in this example, N x 2N).
[0614] This example (color difference component) is explained by assuming a case where there are 5 intra-frame prediction modes (DC, Planar, Vertical, Horizontal, and Color Radiation modes in this example), and the priority is adaptively determined to form a sorted prediction mode as a candidate set for encoding / decoding.
[0615] In this case, the explanation is based on the assumption that priority is determined from adjacent blocks (left, top, top-left, top-right, bottom-left in this example).
[0616] Referring to FIG. 13, a first-rank candidate can be determined from the left (L3), top (T3), top-left (TL), top-right (R0), and bottom-left (B0) blocks. At this time, the mode having the most frequent value among the prediction modes of the blocks can be determined as the first-rank candidate. If multiple modes have the most frequent value, a pre-set priority (e.g., Color diffraction mode - Planar - Vertical - Horizontal - DC) is assigned.
[0617] This example may be viewed as similar to an MPM candidate setting (in this example, the first bit is determined to be 0 or 1. If it is 1, an additional 2 bits are required. Depending on the decoding / encoding setting, the first bit may be bypassed or regular coding, and the remaining bits may be bypassed) in that the prediction of the prediction mode is obtained from neighboring blocks and the priority (i.e., the amount of bits allocated is determined accordingly. For example, '0' for the first priority, and '100', '101', '110', '111' for the second to fourth priorities) is determined.
[0618] Referring to Fig. 22a, the current block (CU0 or PU0, size N x 2N, and assumed to be width / height 1:2) can determine the first-ranked candidate in the l3, t1, tl, t2, bl blocks.
[0619] Meanwhile, referring to Fig. 22b, the current block (CU1 or PU1, size N x 2N) can determine the first-ranked candidate in blocks c13, t3, t1, tr, and k.
[0620] In this example, when a setting such as (2) in FIG. 19 is applied, the current block (CU1) can determine the first priority candidate from blocks t3, t1, tr, and k. Additionally, if the block corresponding to the current block among other color spaces that have acquired the color copy mode is not composed of a single block, the priority of the color copy mode can be shifted backward from the previously set priority and changed to Planar - Vertical - Horizontal - DC - Color Copy Mode, etc. Although this differs from the example of classifying based on whether the parent block of the current block and the adjacent block is the same, the partitioning method, etc., as described above, if the corresponding block in another color space is composed of multiple blocks, it may be an example set under the assumption that it is more likely to have characteristics different from those of the current block. In other words, it can be understood as an example where the encoding / decoding setting is adaptively determined according to the relationship between blocks.
[0621] However, the difference from the case of Fig. 19 is that even though it is a rectangular block, it is partitioned into encoding units and is set directly into prediction units without additional division. Therefore, the candidate group configuration settings as in Fig. 19 may not be applicable in this example.
[0622] However, in the case of k in Fig. 22, since it is a location where decoding / encoding is not completed, it can be derived from an adjacent block where decoding / encoding is completed.
[0623] The above examples are described under certain assumptions for the convenience of explanation, but are not limited thereto, and the same or similar applications may be possible to the various embodiments described above of the present invention.
[0624]
[0625] In the embodiments of FIGS. 21 and 22, it was assumed that M x N (M ≠N) blocks that can be partitioned by binary tree partitioning occur consecutively.
[0626] Among the cases where the above may occur, a case such as the one described in FIGS. 14 to 16 (an example of setting a candidate group by checking the relationship between the current block and neighbor blocks, etc.) may occur. That is, in FIGS. 21 and 22, the neighbor blocks of CU0 and CU1 before splitting are identical to each other, and CU0 and CU1 can be obtained through horizontal or vertical splitting of the binary tree.
[0627] In addition, in the embodiments of FIGS. 18 and 19, it was assumed that a plurality of rectangular prediction blocks occur within the encoding block after the encoding block partition.
[0628] In the above case, a case similar to the one described in FIGS. 14 to 16 may occur. Then, a case that conflicts with the example in FIGS. 18 and 19 may occur. For example, in FIG. 18, in the case of PU1, it is determined whether or not to use the information of PU0, but in FIGS. 14 to 16, it is a case where the information of PU0 is not used because the characteristics of PU1 and the characteristics of PU0 are different.
[0629] Regarding the above matters, a candidate pool can be configured without conflicts depending on the settings of the initial encoding / decoding stage. Additionally, it may be possible to configure a candidate pool for motion information prediction based on various other encoding / decoding settings.
[0630]
[0631] Through the aforementioned example, we examined the case regarding the setting of prediction mode candidates. Additionally, it may be possible to set a restriction on the use of reference pixels used for the prediction of the current block from neighboring blocks marked as exception states.
[0632] For example, when generating a prediction block by distinguishing between a first reference pixel and a second reference pixel according to one embodiment of the present invention, if the second reference pixel is included in a block marked as an exception state, generating the prediction block using the second reference pixel may be restricted. That is, the prediction block may be generated using only the first reference pixel.
[0633] In summary, the above example can be considered as one of the elements of the encoding / decoding settings regarding the use of the second reference pixel.
[0634]
[0635] We examine various cases regarding the setting of the in-screen prediction mode candidate group of the present invention.
[0636] In this example, it is assumed that there are 67 intra-frame prediction modes, consisting of 65 directional modes and 2 non-directional modes (Planar and DC), but this is not limited to this, and other intra-frame prediction mode settings may be possible. In this example, it is assumed that 6 candidates are included in the MPM candidate group. However, this is not limited to this, and 4, 5, or 7 candidates may constitute the MPM candidate group. Furthermore, the candidate group has a setting in which there are no duplicate modes. Additionally, priority refers to the order in which inclusion in the MPM candidate group is determined, but it can be considered as a factor in determining binarization, entropy encoding / decoding settings, etc., for each candidate belonging to the MPM candidate group. The example described below focuses on the luminance component, but the same, similar, or modified application may also be possible for the chrominance component.
[0637] The modes included in the in-frame prediction mode candidate group of the present invention (e.g., MPM candidate group, etc.) may be composed of prediction modes of spatially adjacent blocks, pre-configured prediction modes, and prediction modes derived from prediction modes already included in the candidate group. In this case, the rules for configuring the candidate group (e.g., priority, etc.) may be determined according to the encoding / decoding settings.
[0638]
[0639] The example described below explains the case of a fixed candidate group composition.
[0640] For example (1), a fixed priority may be supported for configuring a prediction mode candidate group within the screen (MPM candidate group in this example). For example, when adding prediction modes of spatially adjacent blocks to the candidate group, a sequence of blocks such as Left (L3 in FIG. 13) - Top (T3 in FIG. 13) - Bottom Left (B0 in FIG. 13) - Top Right (R0 in FIG. 13) - Top Left (TL in FIG. 13) may be supported, and when adding a pre-set prediction mode to the candidate group, a sequence of pre-set priority such as Planar - DC - Vertical - Horizontal - Diagonal modes may be supported. Additionally, a mixed configuration of the above example may be possible, such as the sequence of blocks such as Left - Top - Planar - DC - Bottom Left - Top Right - Top Left. If the number of candidates in the candidate pool is not fully filled even after proceeding to the candidates in the preceding priority, the derived modes of the included prediction modes (e.g., +1, -1 of the left block mode, +1, -1 of the top block mode, etc.) and vertical, horizontal, and diagonal modes may have the next priority.
[0641] In the above example, when adding the prediction modes of spatially adjacent blocks to the candidate group, the priority order is left - top - bottom-left - top-right - top-left blocks, and the prediction mode of the left block (L3 in FIG. 12) is included in the candidate group in order; however, if the corresponding mode does not exist, the prediction mode of the top block (T3 in FIG. 13), which has the next highest priority, is included in the candidate group. In this way, the prediction modes of the corresponding blocks are included in the candidate group in order, but if the prediction mode of the corresponding block is unavailable or overlaps with a mode already included, the order is passed to the next block.
[0642] As another example, when adding prediction modes of spatially adjacent blocks to the candidate pool, the candidate pool can be configured in the order of Left - Top - Bottom Left - Top Right - Top Left blocks. In this case, for the prediction mode of the left block, the prediction mode of the block located at L3 is considered first; however, if there are unavailable or overlapping modes, the prediction mode candidates for the left block can be filled in the order of the next sub-blocks of the left block, L2, L1, and L0. In this way, the same or similar settings are applied to the Top (T3 - T2 - T1 - T0), Bottom Left (B0 - B1 - B2 - B3), Top Right (R0 - R1 - R2 - R3), and Top Left (TL) blocks. For example, even if the prediction mode of the left block proceeds in the order of L3 - L2 - L1 - L0, if it is not possible to add it to the candidate pool, it can proceed to the next ranked block.
[0643]
[0644] The following example describes the case of adaptive candidate pool configuration. Adaptive candidate pool configuration can be determined based on the state of the current block (e.g., block size and shape, etc.), the state of neighboring blocks (e.g., block size and shape, prediction mode, etc.), or the relationship between the current block and neighboring blocks.
[0645] For example (3), an adaptive priority can be supported for configuring the in-screen prediction mode candidate group (MPM candidate group in this example). The priority can be set according to frequency. That is, prediction modes that occur frequently may have a high priority, and prediction modes that occur infrequently may have a low priority.
[0646] For example, priority can be set based on the frequency of prediction modes of spatially adjacent blocks.
[0647] Pre-configured priorities may be supported for cases where the frequencies are the same. For example, assuming that among the prediction modes of the Left, Top, Bottom Left, Top Right, and Top Left blocks (assuming one prediction mode is obtained per block), modes 6 and 31 occur twice and mode 14 occurs once, in this example (assuming that in the order of Left - Top - Bottom Left - Top Right - Top Left blocks, the mode that occurred in the block with the higher priority is mode 6), modes 6 and 31 are included as the first and second candidates.
[0648] Planar and DC are included as the third and fourth candidates, and mode 14, a prediction mode with a frequency of 1, is included as the fifth candidate. Then, modes 5 and 7, and modes 30 and 32, derived from the first and second candidates, can be placed as the next priority. Additionally, modes 13 and 15, derived from the fifth candidate, are placed as the next priority, followed by vertical, horizontal, diagonal modes, etc.
[0649] That is, prediction modes with a frequency of 2 or more are assigned priority before Planar and DC modes, prediction modes with a frequency of 1 are assigned priority after Planar and DC modes, and induction modes such as the example above, pre-set modes, etc. may follow.
[0650] In summary, the pre-set priority (e.g., the order of Left - Top - Planar - DC - Bottom Left - Top Right - Top Left blocks) may be a priority set by considering the statistical characteristics of the general image, and the adaptive priority (in this example, the case of including in the candidate group based on frequency) may be an example of applying partial modification to the pre-set priority by considering the partial characteristics of the image (in this example, Planar and DC are fixed, and modes with two or more frequencies are placed in front, and modes with one frequency are placed behind).
[0651] In addition, not only is the priority for configuring the candidate group determined based on frequency, but binarization and entropy decoding settings for each candidate belonging to the MPM candidate group can also be determined based on frequency. For example, binarization for an MPM candidate with m frequencies can be determined based on frequency. Alternatively, context information for the candidate can be adaptively determined based on frequency. That is, in this example, context information with a high selection probability for the above mode can be used. That is, context information can be set differently when m is 1 to 4 (in this example, since 2 of the total 6 candidates include non-directional modes, the maximum frequency is 4).
[0652] If, among the 0s and 1s of a certain bin index (Bin index: the order of each bit when one or more bits are configured according to binary conversion. For example, if a certain MPM candidate is composed of '010', the first to third bins could be 0, 1, 0), 0 represents a bin signifying that the candidate is selected as an MPM (a situation where whether it is the final MPM is determined by a single bin, or a situation where additional bins must be checked to determine the final MPM. That is, in the '010' above, if the first bin is 0, the second and third must be checked further to confirm whether that mode is the final MPM. If it is a single bin, whether it is the final MPM can be immediately confirmed based on the 0 and 1 of that bin), and 1 represents a bin signifying that the candidate is not selected as an MPM, then context information with a high probability of 0 occurrence (i.e., when performing binary arithmetic, the probability of 0 occurrence can be set to 90% and the probability of 1 occurrence to 10%. The base case is the probability of 0 and 1 occurrence Context-adaptive binary arithmetic coding (CABAC) can be applied by applying (assuming 60% and 40%).
[0653]
[0654] For example (4), an adaptive priority for configuring the in-screen prediction mode candidate group (MPM candidate group in this example) may be supported. For example, the priority can be set according to the direction of the prediction mode of spatially adjacent blocks.
[0655] At this time, the categories for directionality can be classified into a mode group facing upward to the right (modes 2 to 17 in FIG. 9), a horizontal mode group (mode 18 in FIG. 9), a mode group facing downward to the right (modes 19 to 49 in FIG. 9), a vertical mode group (mode 50 in FIG. 9), a mode group facing downward to the left (modes 51 to 66 in FIG. 9), and a non-directional mode group (Planar, DC mode). Alternatively, they can be divided into a horizontal directional mode group (modes 2 to 34 in FIG. 9), a vertical directional mode group (modes 35 to 66 in FIG. 9), and a non-directional mode group (Planar, DC mode), and various examples of configurations are possible.
[0656] For example, if the basic candidate group priority is in the order of Left - Top - Planar - DC - Bottom Left - Top Right - Top Left blocks, the candidate group can be configured in the order above. However, after the candidate group is configured, the priority for binarization and entropy decoding settings for each candidate can be determined based on the above categories. For example, among the MPM candidate groups, binarization with fewer bits allocated can be performed in categories that contain many candidates. Alternatively, context information according to the category can be adaptively determined separately. That is, context information can be determined according to the number of modes included in each category (e.g., when there are m in Category 1 and n in Category 2, context information is determined according to the combination of m and n).
[0657]
[0658] For example (5), adaptive priority can be supported based on the size and shape of the current block. For example, priority can be determined based on the size of the block, and priority can be determined based on the shape of the block.
[0659] If the block size is 32 x 32 or larger, it may be included in the candidate group in the order of Left - Top - Planar - DC - Bottom Left - Top Right - Top Left blocks, and if it is less than 32 x 32, it may be included in the candidate group in the order of Left - Top - Bottom Left - Top Right - Top Left - Planar - DC modes.
[0660] Alternatively, if the block shape is square, it may be included in the candidate group in the order of Left - Top - Planar - DC - Bottom Left - Top Right - Top Left blocks; if the block shape is rectangular (long horizontally), it may be included in the candidate group in the order of Top - Top Right - Top Left - Planar - DC - Left - Bottom Left blocks; and if the block shape is rectangular (long vertically), it may be included in the candidate group in the order of Left - Bottom Left - Top Left - Planar - DC - Top - Top Right blocks. This example can be understood as a case where a higher rank is given to a block adjacent to the longer side of the block.
[0661]
[0662] For example (6), adaptive priority can be supported based on the relationship between the current block and neighboring blocks.
[0663] Referring to FIG. 23a, an example of generating a prediction block according to the prediction mode of a neighbor block (a sub-block of the left block in this example) is shown.
[0664] In the case where the prediction mode of the left block (the upper sub-block of the left block in this example) is a directional mode (a mode tilted to the right from a vertical mode) existing at numbers 51 to 66 in Fig. 9 (used for convenience of explanation; only the content regarding directionality needs to be referred to), the pixels referenced for generating the prediction block of the left block correspond to the shaded area in the figure (part of TL, T). In the case of the right region of the left block, it may actually be a region with high correlation to the left region of the current block, but since the encoding / decoding of the current block has not yet proceeded, the prediction block can be generated through the reference pixels of the lower region of T, where encoding / decoding has already been completed.
[0665] Even though the prediction was performed from the lower region of a part of the T block as described above, the fact that the final prediction mode was determined to be a mode existing in numbers 51 to 66 may mean that a part of the current block (2400) also has a directionality (or edge, etc.) that is the same or similar to the prediction mode above.
[0666] That is, it may mean that the probability of the prediction mode of the left block being selected as the prediction mode of the current block is slightly higher. Or, it may mean that the probability of modes 51 through 66 being selected is slightly higher.
[0667] Fig. 23b, by deriving the relevant explanation from the above example, implies that if the prediction mode of the upper block exists in modes 2 through 17 of Fig. 9 (modes tilted downward from horizontal mode), some areas of the current block also have the same directionality as the prediction mode.
[0668] That is, it may mean that the probability of the prediction mode of the upper block being selected is slightly higher than the prediction mode of the current block. Or, it may mean that the probability of mode 2 through 17 being selected is slightly higher.
[0669] As described above, priority can be determined adaptively by identifying the probability of the current block's prediction mode occurring through the prediction mode of neighboring blocks.
[0670] For example, if the priority order of the basic candidate group composition is in the order of Left - Top - Planar - DC - Bottom Left - Top Right - Top Left blocks, there is no change in priority in the case of Fig. 23a, and in the case of Fig. 23b, a change may occur in the order of Top - Left - Planar - DC - Bottom Left - Top Right - Top Left blocks.
[0671] Alternatively, in the case of Fig. 23a, the sequence may be Left - (Left + 1) - (Left - 1) - Top - Planar - DC - Bottom Left - Top Right - Bottom Left block, and in the case of Fig. 23b, the sequence may be Top - (Top - 1) - (Top + 1) - Left - Planar - DC - Bottom Left - Top Right - Bottom Left block.
[0672]
[0673] Referring to FIG. 24a, an example of generating a prediction block according to the prediction mode of a neighboring block (left block in this example) is shown.
[0674] In the case where the prediction mode of the left block is a directional mode existing at numbers 51 to 66 in Fig. 9, the pixels referenced for generating the prediction block of the left block correspond to the shaded area (TL, T) in the figure. In the case of the right region of the left block, it may actually be a region with high correlation to the left region of the current block, but since the encoding / decoding of the current block has not yet proceeded, the prediction block can be generated through the reference pixels of the lower region of T, where encoding / decoding has already been completed.
[0675] Even though predictions were made from the lower region of block T as described above, the fact that the final prediction mode was determined to be a mode existing in numbers 51 to 66 may mean that some regions of the current block also have the same or similar directionality (or edges, etc.) as the prediction mode above.
[0676] That is, it may mean that there is a high probability that the prediction mode of the left block will be selected as the prediction mode of the current block. Or, it may mean that there is a high probability that modes 51 through 66 will be selected.
[0677] Fig. 24b, by deriving the relevant explanation from the above example, may mean that if the prediction mode of the upper block exists in modes 2 through 17 in Fig. 9, some areas of the current block also have the same directionality as the prediction mode.
[0678] That is, it may mean that there is a high probability that the prediction mode of the upper block will be selected as the prediction mode of the current block. Or, it may mean that there is a high probability that modes 2 through 17 will be selected.
[0679] As described above, priority can be determined adaptively by identifying the probability of the current block's prediction mode occurring through the prediction mode of neighboring blocks.
[0680] For example, if the priority order of the basic candidate group configuration is Left - Top - Planar - DC - Bottom Left - Top Right - Top Left blocks, there is no change in priority in the case of FIG. 24a, and a change may occur in the order of Top - Left - Planar - DC - Bottom Left - Top Right - Top Left blocks in the case of FIG. 24b. In this case, while the priority remains the same as before in the case of FIG. 24a, the binarization and entropy encoding / decoding settings for the prediction mode candidate of the left block among the MPM candidate group can be adaptively determined. For example, binarization (assigning shorter bits) for the prediction mode candidate of the left block can be determined. Alternatively, context information can be adaptively determined. That is, in this example, context information that sets a high probability of the above candidate being selected can be used.
[0681] If 0 of a bin index represents a bin where the candidate is selected as an MPM (a situation where it is determined whether it is the final MPM by a single bin, or a situation where additional bins must be checked to determine the final MPM), and 1 represents a bin where the candidate is not selected as an MPM, then contextual information with a high probability of 0 occurrence can be applied and CABAC can be applied.
[0682] Alternatively, in the case of Fig. 24a, the sequence may be Left - (Left + 1) - (Left - 1) - (Left + 2) - (Left - 2) - Top - Planar - DC - Bottom Left - Top Right - Bottom Left block, and in the case of Fig. 24b, the sequence may be Top - (Top - 1) - (Top + 1) - (Top - 2) - (Top + 2) - Left - Planar - DC - Bottom Left - Top Right - Bottom Left block.
[0683]
[0684] Referring to FIG. 25a, an example of generating a prediction block ac...
Claims
1. A step of deriving an on-screen prediction mode of the current block; A step of determining a pixel line for prediction within the screen of the current block among a plurality of pixel lines; and An intra-screen prediction method comprising a step of performing intra-screen prediction of the current block based on the intra-screen prediction mode and the determined pixel line.
2. In paragraph 1, An intra-screen prediction method further comprising the step of filtering a first reference pixel of the determined pixel line.
3. An intra-screen prediction method in which the above filtering step is selectively performed based on a first flag indicating whether filtering is performed on a first reference pixel for intra-screen prediction.
4. In paragraph 3, The above first flag is derived in the decoding device based on the encoding parameter of the current block, An intra-picture prediction method wherein the encoding parameter includes at least one of a block size, a component type, an intra-picture prediction mode, or whether intra-picture prediction in sub-block units is applied.
5. In paragraph 1, An intra-screen prediction method further comprising a step of correcting a predicted pixel of the current block according to the intra-screen prediction.
6. In the fifth paragraph, the step of correcting, An intra-screen prediction method further comprising a step of determining at least one of a second reference pixel or a weight for the correction based on the position of the predicted pixel of the current block.
7. In paragraph 6, An intra-screen prediction method in which the above-mentioned correcting step is selectively performed by considering at least one of the position of the pixel line of the current block, the intra-screen prediction mode of the current block, or whether intra-screen prediction is performed in units of sub-blocks of the current block.
8. In paragraph 1, The above on-screen prediction is performed in units of sub-blocks of the current block, A method for predicting within a screen, wherein the sub-block is determined based on at least one of a second flag regarding whether to split, split direction information, or split number information.
9. In paragraph 1, The above current block's on-screen prediction mode is a method for on-screen prediction derived based on a predetermined default mode or multiple MPM candidates.