Video encoding / decoding method and device
The method addresses the inefficiencies in video encoding/decoding by optimizing intra prediction using reference pixels and performing affine model-based inter prediction on sub-blocks, resulting in improved coding performance and prediction accuracy.
Patent Information
- Application Number
- JP2023103684
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-07-24
- Filing Date
- 2023-06-23
- Publication Date
- 2025-06-05
- Estimated Expiration
- 2039-03-25
AI Technical Summary
Existing video encoding/decoding technologies face challenges in efficiently compressing and transmitting high-resolution images, particularly in effectively utilizing reference pixels for intra prediction and performing motion compensation.
The proposed method involves identifying a reference pixel area, determining a processing setting based on the usability of the reference pixels, and performing intra prediction accordingly. Additionally, it generates a candidate list for motion information prediction, derives a control point vector, and uses it to obtain a motion vector for inter prediction, all while considering affine candidates and sub-block motion compensation.
This approach improves coding performance by optimizing intra prediction based on available reference pixels, enhances video encoding/decoding efficiency through affine model-based inter prediction, and increases prediction accuracy by performing inter prediction on a sub-block basis.
Smart Images

Figure 0007689158000017 
Figure 0007689158000018 
Figure 0007689158000019
Abstract
Description
[Technical field]
[0001] The present invention relates to a video encoding / decoding method and apparatus. [Background technology]
[0002] 2. Description of the Related Art Recently, the demand for high-resolution, high-quality images such as high definition (HD) images and ultra high definition (UHD) images has been increasing in various application fields, and therefore highly efficient image compression techniques have been discussed.
[0003] There are various video compression technologies, such as inter-frame prediction technology that predicts pixel values contained in the current picture from pictures before or after the current picture, intra-frame prediction technology that predicts pixel values contained in the current picture using pixel information in the current picture, and entropy coding technology that assigns short codes to values that occur frequently and long codes to values that occur less frequently. Using these video compression technologies, video data can be effectively compressed and transmitted or stored. Summary of the Invention [Problem to be solved by the invention]
[0004] SUMMARY OF THE PRESENTLY PREFERRED EMBODIMENTS In order to solve the above problems, an object of the present invention is to provide an image encoding / decoding method and apparatus for performing intra prediction based on the availability of reference pixels.
[0005] An object of the present invention is to provide an inter-prediction method and device.
[0006] SUMMARY OF THE PRESENT EMBODIMENT An object of the present invention is to provide a method and apparatus for sub-block based motion compensation.
[0007] The present invention aims to provide a method and apparatus for determining affine candidates. [Means for solving the problem]
[0008] To achieve the above object, a method for performing intra-screen prediction according to one embodiment of the present invention may include a step of identifying a reference pixel area designated to obtain correlation information, a step of determining a reference pixel processing setting based on a judgment of the usability of the reference pixel area, and a step of performing intra-screen prediction based on the determined reference pixel processing setting.
[0009] The video encoding / decoding method and apparatus according to the present invention may generate a candidate list for motion information prediction of a current block, derive a control point vector of the current block based on the candidate list and a candidate index, derive a motion vector of the current block based on the control point vector of the current block, and perform inter prediction on the current block using the motion vector.
[0010] In the video encoding / decoding apparatus according to the present invention, the candidate list can include a plurality of affine candidates.
[0011] In the video encoding / decoding apparatus according to the present invention, the affine candidates may include at least one of spatial candidates, temporal candidates, or configuration candidates.
[0012] In the video encoding / decoding apparatus according to the present invention, the motion vector of the current block may be derived in units of sub-blocks of the current block.
[0013] In the video encoding / decoding apparatus according to the present invention, the spatial candidates may be determined in consideration of whether a boundary of the current block borders a coding tree block boundary (CTU boundary).
[0014] In the video encoding / decoding apparatus according to the present invention, the configuration candidates may be determined based on at least two combinations of control point vectors corresponding to each corner of the current block. Effect of the Invention
[0015] When the method of performing intra prediction based on the availability of reference pixels according to the present invention is used, the coding performance can be improved.
[0016] According to the present invention, it is possible to improve video encoding / decoding performance by using inter prediction based on an affine model.
[0017] According to the present invention, it is possible to improve prediction accuracy by performing inter prediction on a sub-block basis.
[0018] According to the present invention, the encoding / decoding efficiency of inter prediction can be improved by efficient affine candidate determination. [Brief description of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram of a video encoding and decoding system according to an embodiment of the present invention; [Diagram 2] 1 is a block diagram showing the configuration of a video encoding device according to an embodiment of the present invention; [Diagram 3] 1 is a block diagram of a video decoding device according to an embodiment of the present invention; [Figure 4] 1 is an example diagram illustrating an intra-frame prediction mode according to an embodiment of the present invention; [Diagram 5] FIG. 2 is a conceptual diagram illustrating intra-screen prediction for directional and non-directional modes according to an embodiment of the present invention. [Figure 6] FIG. 1 is a conceptual diagram showing intra-screen prediction for a color copy mode according to an embodiment of the present invention. [Figure 7] 4 is an exemplary diagram showing corresponding blocks and adjacent areas of each color space in relation to a color copy mode according to an embodiment of the present invention; FIG. [Figure 8] 5 is an exemplary diagram illustrating an example of setting an area for obtaining correlation information in a color copy mode according to an embodiment of the present invention. [Figure 9] 2 is an exemplary diagram illustrating a configuration of reference pixels used in intra-frame prediction according to an embodiment of the present invention; [Figure 10] 1 is a conceptual diagram illustrating blocks adjacent to a current block of intra-screen prediction according to an embodiment of the present invention; [Figure 11] 11 is an exemplary diagram illustrating the availability of reference pixels in a color copy mode according to an embodiment of the present invention; FIG. [Figure 12] 11 is an exemplary diagram illustrating the availability of reference pixels in a color copy mode according to an embodiment of the present invention; FIG. [Figure 13] 11 is an exemplary diagram illustrating the availability of reference pixels in a color copy mode according to an embodiment of the present invention; FIG. [Figure 14] 4 is a flowchart illustrating an intra-screen prediction method in a color copy mode according to an embodiment of the present invention. [Figure 15] 11 is an exemplary diagram illustrating prediction in a color copy mode according to an embodiment of the present invention; FIG. [Figure 16] 11 is an exemplary diagram illustrating prediction in a color copy mode according to an embodiment of the present invention; FIG. [Figure 17] 11 is an exemplary diagram illustrating prediction in a color copy mode according to an embodiment of the present invention; FIG. [Figure 18] 4 is a flowchart of a process for performing correction in a color copy mode according to an embodiment of the present invention. [Figure 19] 5 is an example diagram illustrating types of filters applied to a correction target pixel according to an embodiment of the present invention; FIG. [Figure 20] 5 is an example diagram illustrating types of filters applied to a correction target pixel according to an embodiment of the present invention; FIG. [Figure 21] FIG. 2 is a diagram illustrating an inter-frame prediction method according to an embodiment of the present invention. [Figure 22] FIG. 2 is a diagram illustrating a method for deriving affine candidates from spatial / temporal neighboring blocks according to an embodiment of the present invention. [Diagram 23] 13 is a diagram illustrating a method for deriving a candidate constructed by combining motion vectors of spatial / temporal neighboring blocks according to an embodiment of the present invention. FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0020] The video encoding / decoding method and apparatus according to the present invention may include a step of identifying a reference pixel area designated for obtaining correlation information, a step of determining a reference pixel processing setting based on a determination of the usability of the reference pixel area, and a step of performing intra-screen prediction based on the determined reference pixel processing setting.
[0021] The video encoding / decoding method and apparatus according to the present invention may generate a candidate list for motion information prediction of a current block, derive a control point vector of the current block based on the candidate list and a candidate index, derive a motion vector of the current block based on the control point vector of the current block, and perform inter prediction on the current block using the motion vector.
[0022] In the video encoding / decoding apparatus according to the present invention, the candidate list can include a plurality of affine candidates.
[0023] In the video encoding / decoding apparatus according to the present invention, the affine candidates may include at least one of spatial candidates, temporal candidates, or configuration candidates.
[0024] In the video encoding / decoding apparatus according to the present invention, the motion vector of the current block may be derived in units of sub-blocks of the current block.
[0025] In the video encoding / decoding apparatus according to the present invention, the spatial candidates may be determined in consideration of whether a boundary of the current block borders a coding tree block boundary (CTU boundary).
[0026] In the video encoding / decoding apparatus according to the present invention, the configuration candidates may be determined based on at least two combinations of control point vectors corresponding to each corner of the current block.
[0027] Although the present invention can be modified in various ways and has various embodiments, a specific embodiment will be illustrated in the drawings and described in detail. However, this is not intended to limit the present invention to a specific embodiment, and it should be understood that the present invention includes all modifications, equivalents, and alternatives within the spirit and technical scope of the present invention.
[0028] Terms such as first, second, A, B, etc. may be used to describe various elements, but the elements should not be limited to these terms. The terms are used only to distinguish one element from another element. For example, a first element may be termed a second element, and similarly, a second element may be termed a first element, without departing from the scope of the present invention. The term "and / or" includes a combination of multiple associated listed items or any item of multiple associated listed items.
[0029] When a component is referred to as being "coupled" or "connected" to another component, it should be understood that the component may be directly coupled or connected to the other component, but there may also be other components in between. On the other hand, when a component is referred to as being "directly coupled" or "directly connected" to another component, it should be understood that there are no other components in between.
[0030] The terms used in the present invention are used only to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless otherwise clearly indicated in the context. In the present invention, terms such as "include" or "have" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0031] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which the present invention belongs. Terms as defined in commonly used dictionaries should be interpreted as meanings that they have in the context of the relevant art, and unless expressly defined in the present invention, they should not be interpreted as ideal or overly formal.
[0032] Typically, it may be composed of one or more color spaces depending on the color format of an image. It may be composed of one or more pictures having a certain size or one or more pictures having different sizes depending on the color format. For example, in a YCbCr color configuration, color formats such as 4:4:4, 4:2:2, 4:2:0, monochrome (composed of only Y) and the like may be supported. For example, in the case of YCbCr4:2:0, it may be composed of one luminance component (Y in this example) and two chrominance components (Cb / Cr in this example). Here, the composition ratio of the chrominance component to the luminance component may be 1:2 horizontally and vertically. For example, in the case of 4:4:4, the horizontal and vertical lengths may have the same composition ratio. When composed of one or more color spaces as in the above example, the picture may be divided into each color space.
[0033] Images can be classified into I, P, B, etc., depending on the image type (e.g., picture type, slice type, tile type, etc.). The I image type may mean an image that is encoded / decoded by itself without using a reference picture, the P image type may mean an image that is encoded / decoded using a reference picture but allows only forward prediction, and the B image type may mean an image that is encoded / decoded using a reference picture and allows forward / backward prediction, but some of the types may be combined (combining P and B) or other image types may be supported depending on the encoding / decoding setting.
[0034] FIG. 1 is a conceptual diagram of a video encoding and decoding system according to an embodiment of the present invention.
[0035] Referring to FIG. 1, the video encoding device 105 and the decoding device 100 may be a user terminal such as a personal computer (PC), a notebook computer, a personal digital assistant (PDA), a portable multimedia player (PMP), a PlayStation Portable (PSP), a wireless communication terminal, a smart phone, or a TV, or a server terminal such as an application server or a service server, and may include various devices including a communication device such as a communication modem for communicating with various devices or wired and wireless communication networks, memories 120, 125 for storing various programs and data for inter or intra prediction to encode or decode a video, or processors 110, 115 for executing programs for calculation and control, etc.
[0036] In addition, the image encoded into a bitstream by the video encoding device 105 may be transmitted to the video decoding device 100 in real time or non-real time via a wired or wireless communication network (Network) such as the Internet, a short-range wireless communication network, a wireless LAN network, a WiBro network, or a mobile communication network, or via various communication interfaces such as a cable or a Universal Serial Bus (USB), and may be decoded by the video decoding device 100 to restore the image and be played back. In addition, the image encoded into a bitstream by the video encoding device 105 may be transmitted from the video encoding device 105 to the video decoding device 100 via a computer-readable recording medium.
[0037] The video encoding device and the video decoding device may be separate devices, but may be implemented as a single video encoding / decoding device. In this case, some components of the video encoding device may be substantially the same technical elements as some components of the video decoding device, and may be implemented to include at least the same structure or perform at least the same function.
[0038] Therefore, in the following detailed description of technical elements and their operating principles, the overlapping description of corresponding technical elements will be omitted. Also, since a video decoding device corresponds to a computer device that applies a video encoding method performed in a video encoding device to decoding, the following description will focus on the video encoding device.
[0039] The computer device may include a memory for storing a program or software module for implementing the video encoding method and / or the video decoding method, and a processor coupled to the memory for executing the program. Here, the video encoding device may be referred to as an encoder, and the video decoding device may be referred to as a decoder.
[0040] FIG. 2 is a block diagram showing the configuration of a video encoding device according to an embodiment of the present invention.
[0041] Referring to FIG. 2, the video encoding device 20 may include a prediction unit 200, a subtraction unit 205, a transformation unit 210, a quantization unit 215, an inverse quantization unit 220, an inverse transformation unit 225, an addition unit 230, a filter unit 235, an encoding picture buffer 240, and an entropy encoding unit 245.
[0042] The prediction unit 200 may be implemented using a prediction module, which is a software module, and may generate a prediction block for a block to be encoded using an intra prediction method or an inter prediction method. The prediction unit 200 may generate a prediction block by predicting a current block to be encoded in an image. In other words, the prediction unit 200 may generate a prediction block having a predicted pixel value of each pixel generated by predicting a pixel value of each pixel of a current block to be encoded in an image using intra prediction or inter prediction. In addition, the prediction unit 200 may transmit information required for generating a prediction block, such as information about a prediction mode, such as an intra prediction mode or an inter prediction mode, to an encoding unit, and may cause the encoding unit to encode information about the prediction mode. Here, a processing unit in which prediction is performed and a prediction method and a processing unit in which specific contents are determined may be determined according to encoding / decoding settings. For example, a prediction method, a prediction mode, etc. may be determined in a prediction unit, and prediction may be performed in a conversion unit.
[0043] The inter-frame prediction unit may be classified into a translation motion model and an affine motion model according to a motion prediction method. In the translation motion model, prediction is performed considering only translation, and in the affine motion model, prediction is performed considering not only translation but also rotation, perspective, zoom in / out, etc. When assuming unidirectional prediction, one motion vector may be required in the translation motion model, but one or more motion vectors may be required in the affine motion model. In the affine motion model, each motion vector may be information applied to a preset position of the current block, such as the upper left vertex or upper right vertex of the current block, and the position of the area to be predicted of the current block may be obtained in pixel units or sub-block units according to the corresponding motion vector. In the inter-frame prediction unit, some processes described below may be commonly applied depending on the motion model, and some processes may be individually applied.
[0044] The inter prediction unit may include a reference picture construction unit, a motion estimation unit, a motion compensation unit, a motion information determination unit, and a motion information coding unit. The reference picture construction unit may include pictures coded before or after the current picture in a reference picture list (L0, L1). A prediction block may be obtained from a reference picture included in the reference picture list, and the current picture may also be constructed from reference pictures according to a coding setting and may be included in at least one of the reference picture lists.
[0045] In the inter prediction unit, the reference picture constructing unit may include a reference picture interpolating unit, and may perform an interpolation process for a small number of unit pixels according to the interpolation accuracy. For example, an interpolation filter based on 8-tap DCT may be applied to a luminance component, and an interpolation filter based on 4-tap DCT may be applied to a chrominance component.
[0046] In the inter-frame prediction section, the motion estimation section is a process of searching for blocks that have a high correlation with the current block using a reference picture, and can use various methods such as FBMA (Full search-based block matching algorithm) and TSS (Three step search), and the motion compensation section refers to the process of obtaining a prediction block through the motion estimation process.
[0047] In the inter prediction unit, the motion information determination unit may perform a process for selecting optimal motion information of the current block, and the motion information may be coded according to a motion information coding mode such as skip mode, merge mode, and competition mode. The modes may be configured by combining modes supported by a motion model, and examples thereof may include skip mode (moving), skip mode (non-moving), merge mode (moving), merge mode (non-moving), competition mode (moving), and competition mode (non-moving). Depending on the coding setting, some of the modes may be included in the candidate group.
[0048] The motion information coding mode can obtain a predicted value of motion information (motion vector, reference picture, prediction direction, etc.) of the current block from at least one candidate block, and when two or more candidate blocks are supported, optimal candidate selection information can be generated. The skip mode (without residual signal) and the merge mode (with residual signal) can use the predicted value as the motion information of the current block as it is, and the competitive mode can generate difference value information between the motion information of the current block and the predicted value.
[0049] The candidate group for the motion information prediction value of the current block can have various adaptive configurations according to the motion information coding mode. The candidate group can include motion information of blocks spatially adjacent to the current block (e.g., left, upper, upper left, upper right, lower left blocks, etc.), motion information of blocks temporally adjacent to the current block, and combination motion information of spatial candidates and temporal candidates can be included in the candidate group.
[0050] The temporally adjacent blocks include blocks in other images that correspond to (or are equivalent to) the current block, and may refer to blocks located on the left, right, upper, lower, upper left, upper right, lower left, lower right, etc., around the current block. The combined motion information may refer to information obtained as an average, median, etc. from the motion information of spatially adjacent blocks and the motion information of temporally adjacent blocks.
[0051] There may be a priority order for constructing a motion information predictor candidate group. The order of steps included in constructing a predictor candidate group may be determined according to the priority order, and the candidate group construction may be completed when the number of candidates (determined according to a motion information coding mode) is filled according to the priority order. Here, the priority order may be determined in the order of motion information of spatially adjacent blocks, motion information of temporally adjacent blocks, and combined motion information of spatial candidates and temporal candidates, but other variations are also possible.
[0052] For example, among spatially adjacent blocks, the candidate group may include blocks in the order of left-top-top-top-bottom-left-top block, etc., and among temporally adjacent blocks, the candidate group may include blocks in the order of bottom-right-middle-right-bottom block, etc.
[0053] The subtraction unit 205 may subtract the predicted block from the current block to generate a residual block. In other words, the subtraction unit 205 may calculate the difference between the pixel value of each pixel of the current block to be encoded and the predicted pixel value of each pixel of the predicted block generated by the prediction unit to generate a residual block that is a residual signal in a block form. In addition, the subtraction unit 205 may generate a residual block in a unit other than the block unit obtained by the block division unit described later.
[0054] The transform unit 210 can transform a signal belonging to the spatial domain into a signal belonging to the frequency domain, and a signal obtained by the transform process is called a transformed coefficient. For example, a residual block having a residual signal transmitted from the subtraction unit can be transformed to obtain a transformed block having a transformed coefficient. The input signal is determined by a coding setting, and is not limited to a residual signal.
[0055] The transform unit may transform the residual block using a transform technique such as a Hadamard transform, a discrete sine transform (DST Based-Transform), or a discrete cosine transform (DCT Based-Transform), but is not limited thereto, and may use various transform techniques that are improvements or modifications of these.
[0056] At least one of the transformation techniques may be supported, and at least one detail transformation technique may be supported for each of the transformation techniques, where the detail transformation techniques may be configured such that a portion of the basis vector is different in each transformation technique.
[0057] For example, in the case of DCT, one or more detail conversion techniques of DCT-I to DCT-VIII can be supported, and in the case of DST, one or more detail conversion techniques of DST-I to DST-VIII can be supported. A part of the detail conversion techniques can be configured as a conversion technique candidate group. For example, DCT-II, DCT-VIII, and DST-VII can be configured as a conversion technique candidate group to perform conversion.
[0058] The transformation can be performed in the horizontal / vertical direction. For example, a one-dimensional transformation can be performed in the horizontal direction using the DCT-II transformation technique, and a one-dimensional transformation can be performed in the vertical direction using the DST-VIII transformation technique, resulting in a total two-dimensional transformation, which can transform pixel values in the spatial domain into the frequency domain.
[0059] The conversion can be performed using one fixed conversion technique, or the conversion technique can be adaptively selected according to the encoding / decoding setting. Here, in the adaptive case, the conversion technique can be selected using an explicit or implicit method. In the explicit case, each conversion technique selection information or conversion technique set selection information applied in the horizontal and vertical directions can be generated in units such as blocks. In the implicit case, the encoding setting can be defined according to the image type (I / P / B), color components, block size, shape, intra-frame prediction mode, etc., and a predefined conversion technique can be selected accordingly.
[0060] Also, some of the conversions may be omitted depending on the encoding settings, which means that one or more horizontal / vertical units may be omitted explicitly or implicitly.
[0061] In addition, the transform unit can transmit the information necessary to generate a transform block to the encoding unit so as to encode it, and the resulting information can be recorded in a bitstream and transmitted to a decoder, and the decoding unit of the decoder can parse the information and use it in the inverse transform process.
[0062] The quantizer 215 may quantize an input signal. Here, a signal obtained through the quantization process is called a quantized coefficient. For example, a residual block having a residual transform coefficient transmitted from a transformer may be quantized to obtain a quantized block having a quantized coefficient. The input signal is determined according to a coding setting, and is not limited to the residual transform coefficient.
[0063] The quantization unit may quantize the transformed residual block using a quantization technique such as Dead Zone Uniform Threshold Quantization, Quantization Weighted Matrix, etc., but is not limited thereto, and may use various improved and modified quantization techniques.
[0064] The quantization process may be omitted depending on the encoding settings. For example, the quantization process may be omitted (including the inverse process) depending on the encoding settings (e.g., the quantization parameter is 0, i.e., lossless compression environment). As another example, the quantization process may be omitted when the compression performance by quantization is not exhibited due to the characteristics of the image. Here, the area in the quantization block (M×N) where the quantization process is omitted may be the entire area or a partial area (M / 2×N / 2, M×N / 2, M / 2×N, etc.), and the quantization omission selection information may be determined implicitly or explicitly.
[0065] The quantization unit can transmit the information necessary to generate a quantization block to the encoding unit so that it can be encoded, and the resulting information can be recorded in a bitstream and transmitted to the decoder, and the decoding unit of the decoder can parse the information and use it in the inverse quantization process.
[0066] In the above example, the explanation was given under the assumption that the residual block is transformed and quantized by the transform unit and the quantization unit, but the residual signal may be transformed to generate a residual block having transform coefficients without performing the quantization process, the residual signal of the residual block may be not transformed into transform coefficients and only the quantization process may be performed, or neither the transform nor the quantization process may be performed, which can be determined according to the encoder settings.
[0067] The inverse quantization unit 220 inverse quantizes the residual block quantized by the quantization unit 215. That is, the inverse quantization unit 220 inverse quantizes the quantized frequency coefficient sequence to generate a residual block having frequency coefficients.
[0068] The inverse transform unit 225 inversely transforms the residual block dequantized by the inverse quantization unit 220. That is, the inverse transform unit 225 inversely transforms the frequency coefficients of the dequantized residual block to generate a residual block having pixel values, i.e., a restored residual block. Here, the inverse transform unit 225 can perform inverse transform by using the transform method used by the transform unit 210 inversely.
[0069] The adder 230 reconstructs a current block by adding the predicted block predicted by the prediction unit 200 and the residual block reconstructed by the inverse transform unit 225. The reconstructed current block is stored as a reference picture (or reference block) in the coding picture buffer 240 and can be used as a reference picture when coding the next block of the current block or other future blocks or pictures.
[0070] The filter unit 235 may include one or more post-processing filter processes, such as a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF). The deblocking filter can remove block noise generated at the boundary between blocks from the restored picture. The ALF can perform filtering based on a value obtained by comparing an image restored after a block is filtered through a deblocking filter with an original image. The SAO can restore an offset difference between the original image and a residual block to which the deblocking filter is applied in pixel units. Such post-processing filters can be applied to a restored picture or block.
[0071] The coding picture buffer 240 may store a block or a picture reconstructed through the filter unit 235. The reconstructed block or picture stored in the coding picture buffer 240 may be provided to a prediction unit 200 that performs intra prediction or inter prediction.
[0072] The entropy coding unit 245 scans the generated quantized frequency coefficient sequence in various scanning methods to generate a quantized coefficient sequence, and outputs the quantized coefficient sequence by encoding it using an entropy coding technique, etc. The scan pattern can be set to one of various patterns such as zigzag, diagonal, raster, etc. Also, the entropy coding unit 245 can generate encoded data including encoding information transmitted from each component and output it as a bit stream.
[0073] FIG. 3 is a block diagram of a video decoding device according to an embodiment of the present invention.
[0074] Referring to FIG. 3, the video decoding device 30 may include an entropy decoding unit 305, a prediction unit 310, an inverse quantization unit 315, an inverse transform unit 320, an adder / subtractor 325, a filter 330, and a decoded picture buffer 335.
[0075] Furthermore, the prediction unit 310 may further include an intra-frame prediction module and an inter-frame prediction module.
[0076] First, when a video bitstream is received from the video encoding device 20 , it can be transmitted to the entropy decoding unit 305 .
[0077] The entropy decoder 305 may decode the bitstream to generate decoded data including quantized coefficients and decoding information transmitted to each component.
[0078] The prediction unit 310 may generate a prediction block based on data transferred from the entropy decoding unit 305. Here, the prediction unit 310 may also configure a reference picture list using a default configuration technique based on a reference image stored in the decoded picture buffer 335.
[0079] The inter prediction unit may include a reference picture construction unit, a motion compensation unit, and a motion information decoding unit, some of which may perform the same process as the encoder, and some of which may perform a reverse guidance process.
[0080] The inverse quantization unit 315 may inverse quantize the quantized transform coefficients provided as a bitstream and decoded by the entropy decoding unit 305 .
[0081] The inverse transform unit 320 may apply an inverse transform technique such as an inverse DCT, an inverse integer transform, or a similar concept to the transform coefficients to generate residual blocks.
[0082] Here, the inverse quantization unit 315 and the inverse transform unit 320 may be implemented in various ways by inversely performing the processes performed by the transform unit 210 and the quantization unit 215 of the video encoding device 20. For example, the inverse quantization unit 315 and the inverse transform unit 320 may use the same processes and inverse transforms shared by the transform unit 210 and the quantization unit 215, or may inversely perform the transform and quantization processes using information about the transform and quantization processes from the video encoding device 20 (e.g., transform size, transform shape, quantization type, etc.).
[0083] The residual block that has undergone the inverse quantization and inverse transformation processes can be added to the prediction block derived by the prediction unit 310 to generate a reconstructed image block.
[0084] The filter 330 may apply a deblocking filter to the reconstructed image block to remove blocking phenomena if necessary, and other loop filters may be additionally used before or after the decoding process to improve video quality.
[0085] The restored and filtered image block may be stored in a decoded picture buffer 335 .
[0086] Although not shown in the drawing, the video encoding / decoding apparatus may further include a picture dividing unit and a block dividing unit.
[0087] The picture divider can divide (or partition) the picture into at least one processing unit such as a color space (e.g., YCbCr, RGB or XYZ), a tile, a slice, a basic coding unit (or a maximum coding unit), and the block divider can divide the basic coding unit into at least one processing unit (e.g., encoding, prediction, transform, quantization, entropy and in-loop filter units, etc.).
[0088] The basic coding unit may be obtained by dividing a picture at regular intervals in the horizontal and vertical directions. Based on this, division into tiles, slices, etc. may be performed, but is not limited to this. The division units such as tiles and slices may be configured as integer multiples of the basic coding block, but exceptional cases may occur for division units located at image boundaries. For this reason, adjustment of the basic coding block size may occur.
[0089] For example, a picture may be partitioned into basic coding units and then partitioned into the units, or a picture may be partitioned into the units and then partitioned into the basic coding units. In the present invention, the partition and partition order of each unit is assumed to be the former case, but the present invention is not limited thereto, and the latter case is also possible depending on the encoding / decoding setting. In the latter case, the size of the basic coding unit may be modified to be adaptive depending on the division unit (tile, etc.). That is, it means that basic coding blocks having different sizes for each division unit can be supported.
[0090] In the present invention, the following examples will be described with a default setting of dividing a picture into basic coding units. The default setting may mean a case where a picture is not divided into tiles or slices, or a picture is one tile or one slice. However, it should be understood that the following various embodiments can be applied in the same way or with modifications even when, as described above, each division unit (tile, slice, etc.) is first divided and divided into basic coding units according to the obtained unit (i.e., when each division unit is not an integer multiple of the basic coding unit).
[0091] Among the division units, a slice may be composed of at least one block that is continuous according to a scan pattern, a tile may be composed of a rectangular block of spatially adjacent blocks, and other additional division units may be supported and configured according to the definitions therefor. Slices and tiles may be division units supported for purposes such as parallel processing, and for this reason, references between division units may be restricted (i.e., cannot be referenced).
[0092] For slices, division information for each unit can be generated with information about the starting positions of consecutive blocks, and for tiles, information about horizontal and vertical division lines or tile position information (e.g., upper left, upper right, lower left, lower right positions) can be generated.
[0093] Here, slices and tiles can be divided into multiple units according to encoding / decoding settings.
[0094] For example, some units can be a unit that contains configuration information that affects the encoding / decoding process (i.e., a tile header or slice header), and some units can be a unit that does not contain configuration information, or some unit is a unit that cannot refer to other units during the encoding / decoding process, and some units can be a unit that can be referenced. Also, some units is another unit Can be a hierarchical relationship involving some units is another unit and can have an equal relationship.
[0095] Here, A and B can be a slice and a tile (or a tile and a slice), or A and B can consist of one slice and one tile. For example, A can be a slice / tile <type 1> and B can be a slice / tile <type 2>.
[0096] Here, type 1 and type 2 can each be one slice or tile, or type 1 can be multiple slices or tiles (including type 2) (a slice or tile collection) and type 2 can be one slice or tile.
[0097] As already described above, the present invention will be described assuming that a picture is composed of one slice or tile, but if two or more division units are generated, the above description can be applied to the following embodiments. Also, A and B are examples of characteristics that a division unit may have, and examples in which A and B are combined are also possible.
[0098] Meanwhile, the image may be divided into blocks of various sizes by the block division unit. Here, the block may be composed of one or more blocks (e.g., one luminance block and two chrominance blocks) according to the color format, and the size of the block may be determined according to the color format. For convenience of explanation, the following description will be given based on a block of one color component (luminance component).
[0099] It should be understood that the following description is for one color component, but that it can be changed and applied to other color components in proportion to the ratio according to the color format (for example, in the case of YCbCr4:2:0, the ratio of the width and height of the luminance component and the chrominance component is 2:1). It should also be understood that although block division that depends on other color components (for example, Cb / Cr depending on the result of block division of Y) is possible, independent block division is possible for each color component. It should also be understood that while one common block division setting (considering that it is proportional to the length ratio) can be used, it is also possible to use individual block division settings for each color component.
[0100] A block may have a variable size such as M×N (M and N are integers such as 4, 8, 16, 32, 64, 128, etc.) and may be a unit for performing coding (coding block). In detail, it may be a basic unit for prediction, transformation, quantization, entropy coding, etc., and is generally referred to as a block in the present invention. Here, a block does not only mean a rectangular block, but should be understood as a broad concept including regions of various shapes such as a triangle and a circle, but the present invention will mainly be described in the case of a rectangular block.
[0101] The block division unit may be set in relation to each component of the image encoding device and the decoding device, and may determine the size and shape of the block through this process. Here, the set block may be defined differently depending on the component, and may correspond to a prediction block in the case of a predictor, a transform block in the case of a transformer, and a quantization block in the case of a quantizer. However, without being limited thereto, a block unit may be additionally defined according to other components. In the present invention, it is assumed that the input and output of each component are blocks (i.e., rectangles), but some components may have inputs / outputs of other shapes (e.g., squares, triangles, etc.).
[0102] The size and shape of the initial (or starting) block of the block division unit can be determined from a higher unit. For example, in the case of a coding block, a basic coding block can be the initial block, and in the case of a prediction block, a coding block can be the initial block. Also, in the case of a transformation block, a coding block or a prediction block can be the initial block, which can be determined by the coding / decoding setting.
[0103] For example, if the coding mode is intra, the prediction block may be a superordinate unit of the transform block, and if the coding mode is inter, the prediction block may be a unit independent of the transform block. An initial block, which is a starting unit of division, may be divided into blocks of small size, and when an optimal size and shape according to the division of the block is determined, the block may be determined as an initial block of the lower unit. An initial block, which is a starting unit of division, may be regarded as an initial block of the upper unit. Here, the upper unit may be a coding block, and the lower unit may be a prediction block or a transform block, but is not limited thereto. Once the initial block of the lower unit is determined as in the above example, a division process may be performed to find a block of optimal size and shape like the upper unit.
[0104] In summary, the block division unit can divide the basic coding unit (or the largest coding unit) into at least one coding unit (or sub-coding unit). The coding unit can also be divided into at least one prediction unit and can be divided into at least one transform unit. The coding unit can be divided into at least one coding block, the coding block can be divided into at least one prediction block and can be divided into at least one transform block. The prediction unit can be divided into at least one prediction block, and the transform unit can be divided into at least one transform block.
[0105] Here, in the case of some blocks, one division process may be performed by combining with other blocks. For example, when a coding block and a transform block are combined into a single unit, a division process is performed to obtain an optimal block size and shape, which may be the optimal size and shape of the coding block as well as the optimal size and shape of the transform block. Alternatively, a coding block and a transform block may be combined into a single unit, a prediction block and a transform block may be combined into a single unit, a coding block, a prediction block and a transform block may be combined into a single unit, and other block combinations are possible. However, the combination is not applied collectively within an image (picture, slice, tile, etc.), but may be adaptively determined according to detailed conditions of a block unit (e.g., image type, coding mode, block size / shape, prediction mode information, etc.).
[0106] As described above, when a block having an optimal size and shape is found, mode information (e.g., partition information, etc.) for the block can be generated. The mode information can be recorded in a bitstream together with information generated in a component to which the block belongs (e.g., prediction-related information and transformation-related information, etc.) and transmitted to a decoder, where it can be parsed in the same level unit and used in the video decoding process.
[0107] The following describes a division method, and for convenience of explanation, it is assumed that the initial block is square, but this is not limited to this, as the same or similar application can be applied when the initial block is rectangular.
[0108] Although various methods for block division can be supported, the present invention will be described with emphasis on tree-based division, and at least one tree division can be supported. Here, the tree method can be a quad tree (QT), a binary tree (BT), a ternary tree (TT), etc. When one tree method is supported, it is called a single tree division, and when two or more tree methods are supported, it is called a multi-tree method.
[0109] Quad tree partitioning refers to a method of dividing a block into two parts horizontally and vertically, binary tree partitioning refers to a method of dividing a block into two parts in one direction, horizontally or vertically, and ternary tree partitioning refers to a method of dividing a block into three parts horizontally or vertically.
[0110] In the present invention, it is assumed that when a block before division is M×N, the quad tree division is divided into four M / 2×N / 2 blocks, the binary tree division is divided into two M / 2×N blocks or M×N / 2 blocks, and the ternary tree division is divided into M / 4×N / M / 2×N / M / 4×N blocks or M×N / 4 / M×N / 2 / M×N / 4 blocks. However, the division results are not limited to the above, and various modified examples are possible.
[0111] Depending on the encoding / decoding configuration, one or more types of tree partitioning may be supported, for example, quad tree partitioning may be supported, or quad tree partitioning and binary tree partitioning may be supported, or quad tree partitioning and ternary tree partitioning may be supported, or quad tree partitioning, binary tree partitioning and ternary tree partitioning may be supported.
[0112] The above example is an example of a case where the basic partitioning method is a quad tree, and binary tree partitioning and ternary tree partitioning are included as additional partitioning methods depending on whether other trees are supported, but various modifications are possible. Here, information on whether other trees are supported (bt_enabled_flag, tt_enabled_flag, bt_tt_enabled_flag, etc., which can have a value of 0 or 1, where 0 means not supported and 1 means supported) can be implicitly determined by encoding / decoding settings or explicitly determined in units of sequences, pictures, slices, tiles, etc.
[0113] The division information may include information on whether division is possible (tree_part_flag, or qt_part_flag, bt_part_flag, tt_part_flag, bt_tt_part_flag. It may have a value of 0 or 1, where 0 means no division and 1 means division). In addition, depending on the division method (binary tree, ternary tree), information on the division direction (dir_part_flag, or bt_dir_part_flag, tt_dir_part_flag, bt_tt_dir_part_flag. It may have a value of 0 or 1, where 0 means <horizontal> and 1 means <vertical>) may be added, which may be information that can be generated when division is performed.
[0114] When multiple tree divisions are supported, various division information configurations are possible. Next, an example of how division information is configured at a one-depth level (i.e., the supported division depth is set to one or more, and recursive division is possible, but for the sake of convenience of explanation, the following will be described).
[0115] In the example (1), information on whether or not division is possible is confirmed. If division is not to be carried out, division is terminated.
[0116] If splitting is to be performed, the selection information for the split type (for example, tree_idx. If it is 0, it is QT, if it is 1, it is BT, if it is 2, it is TT) is checked. Here, the split direction information is further checked according to the selected split type, and the next step is started (if additional splitting is possible due to reasons such as the split depth not reaching the maximum, it starts again from the beginning, if splitting is not possible, it ends the splitting).
[0117] In example (2), information regarding whether or not a partial tree method (QT) can be split is checked, and then the next step is started. If splitting is not to be performed at this stage, information regarding whether or not a partial tree method (BT) can be split is checked. If splitting is not to be performed at this stage, information regarding whether or not a partial tree method (TT) can be split is checked. If splitting is not to be performed at this stage, the splitting is terminated.
[0118] If a partial tree type (QT) split is to be performed, the process proceeds to the next stage. If a partial tree type (BT) split is to be performed, the process checks the split direction information and proceeds to the next stage. If a partial tree split type (TT) split is to be performed, the process checks the split direction information and proceeds to the next stage.
[0119] In example (3), information on whether or not splitting is possible for some tree methods (QT) is checked. If splitting is not to be performed, information on whether or not splitting is possible for some tree methods (BT and TT) is checked. If splitting is not to be performed, splitting is terminated.
[0120] If a partial tree type (QT) split is to be performed, the process proceeds to the next step. If a partial tree type (BT and TT) split is to be performed, the process checks the split direction information and proceeds to the next step.
[0121] The above examples are cases where tree division priority exists (examples 2 and 3) or does not exist (example 1), but various variations are possible. Also, the above examples are examples that explain the case where the division of the current stage is independent of the division result of the previous stage, but it is also possible to set the division of the current stage to be dependent on the division result of the previous stage.
[0122] For example, in the case of examples 1 to 3, if some tree-based splitting (QT) was performed in the previous stage and then moved to the current stage, the same tree-based splitting (QT) can be supported in the current stage.
[0123] On the other hand, if a partial tree-based division (QT) was not performed in the previous stage and another tree-based division (BT or TT) was performed before moving to the current stage, it is also possible to set the partial tree-based division (BT and TT) to be supported in subsequent stages including the current stage, except for the partial tree-based division (QT).
[0124] In this case, it means that the tree structure supporting block division is adaptive, which means that the above-mentioned division information structure can also be configured differently (assuming the following example is the third example). That is, in the above example, if the division of a part of the tree type (QT) was not performed in the previous step, the division process can be performed in the current step without considering the part of the tree type (QT). Also, the division information for the related tree type (e.g., information on whether or not to divide, division direction information, etc. In this example, <qt>In this case, the information about whether or not the data can be divided can be removed.
[0125] The above example is for adaptive partition information configuration when block partitioning is allowed (e.g., the block size is within the range between the maximum and minimum values, and the partitioning depth of each tree method does not reach the maximum depth <allowed depth>), but adaptive partition information configuration is also possible when block partitioning is restricted (e.g., the block size is not within the range between the maximum and minimum values, and the partitioning depth of each tree method reaches the maximum depth).
[0126] As described above, in the present invention, the tree-based partitioning can be performed in a recursive manner. For example, if the partition flag of a coding block with partition depth k is 0, the coding block is encoded in a coding block with partition depth k, and if the partition flag of a coding block with partition depth k is 1, the coding block is encoded in N sub-coding blocks (where N is an integer equal to or greater than 2, such as 2, 3, or 4) with a partition depth k+1 according to a partitioning scheme.
[0127] The sub-coding block can be further set as a coding block (k+1) and divided into a sub-coding block (k+2) through the above process, and such a hierarchical division method can be determined according to division settings such as the division range and the allowable division depth.
[0128] Here, the bitstream structure for expressing the partition information can be selected from one or more scanning methods. For example, the bitstream of the partition information can be configured based on the partition depth order or based on whether partitioning is possible.
[0129] For example, in the case of a partition depth order, partition information at a current level of depth is obtained based on a first block, and then partition information at a next level of depth is obtained, and in the case of a partition possibility, additional partition information of a partitioned block is preferentially obtained based on a first block, and other additional scanning methods can be considered. In the present invention, a case is assumed in which a bit stream of partition information is constructed based on whether partition is possible.
[0130] As described above, various cases for block division have been described. Fixed or adaptive settings for block division can be supported.
[0131] Here, the setting for block partitioning may explicitly include related information in units of a sequence, a picture, a slice, a tile, etc. Or, the block partitioning setting may be implicitly determined by the encoding / decoding setting. Here, the encoding / decoding setting may be configured by one or a combination of two or more of various encoding / decoding elements such as an image type (I / P / B), a color component, a partition type, a partition depth, etc.
[0132] In the video encoding method according to an embodiment of the present invention, the intra prediction may be configured as follows. The intra prediction of the predictor may include a reference pixel configuration step, a prediction block generation step, a prediction mode determination step, and a prediction mode encoding step. Also, the video encoding device may be configured to include a reference pixel configuration unit, a prediction block generation unit, and a prediction mode encoding unit implementing the reference pixel configuration step, the prediction block generation step, the prediction mode determination step, and the prediction mode encoding step. Some of the above-described steps may be omitted or other steps may be added, and the steps may be changed to an order other than the above-described order.
[0133] FIG. 4 is a diagram illustrating an example of an intra-frame prediction mode according to an embodiment of the present invention.
[0134] Referring to FIG 4, 67 prediction modes are configured as a prediction mode candidate group for intra prediction. It is assumed that 65 of the prediction modes are directional modes and 2 are non-directional modes (DC, Planar), but the present invention is not limited thereto and various configurations are possible. Here, the directional modes may be classified according to gradient (e.g., dy / dx) or angle information (Degree). In addition, all or a part of the prediction modes may be included in a prediction mode candidate group for a luminance component or a chrominance component, and additional modes may be included in the prediction mode candidate group.
[0135] In the present invention, the direction of the directional mode may mean a straight line, and a curved directional mode may also be configured as a prediction mode. In addition, the non-directional mode may include a DC mode that obtains a prediction block from the average (or weighted average, etc.) of pixels of neighboring blocks (e.g., left, upper, upper left, upper right, lower left blocks, etc.) adjacent to the current block, and a planar mode that obtains a prediction block by linearly interpolating pixels of the neighboring blocks.
[0136] Here, in the case of DC mode, the reference pixels used to generate the prediction block can be obtained from blocks consisting of various combinations such as the left side, upper side, left side + upper side, left side + lower left, upper side + upper right, and left side + upper side + lower left + upper right, and the reference pixel acquisition block position can be determined according to the encoding / decoding setting defined by the image type, color component, block size / shape / position, etc.
[0137] Here, in the case of the planar mode, the pixels used for generating the predicted block can be obtained from a region composed of reference pixels (e.g., left side, upper side, upper left side, upper right side, lower left side, etc.) and a region not composed of reference pixels (e.g., right side, lower side, lower right side, etc.), and in the case of a region not composed of reference pixels (i.e., not encoded), it can be obtained implicitly by using one or more pixels from the region composed of reference pixels (e.g., directly copying, weighted average, etc.), or information about at least one pixel of the region not composed of reference pixels can be explicitly generated. Thus, a predicted block can be generated using the region composed of reference pixels and the region not composed of reference pixels in this way.
[0138] FIG. 5 is a conceptual diagram illustrating intra-frame prediction for directional and non-directional modes according to an embodiment of the present invention.
[0139] Referring to (a) of Fig. 5, intra-frame prediction in vertical (5a), horizontal (5b), and diagonal (5c to 5e) modes is shown. Referring to (b) of Fig. 5, intra-frame prediction in DC mode is shown. Referring to (c) of Fig. 5, intra-frame prediction in Planar mode is shown.
[0140] The present invention may include additional non-directional modes other than those described above. Although the present invention will be described with a focus on the linear directional mode and the DC and planar non-directional modes, modifications and application to other cases are possible.
[0141] 4 may be prediction modes that are fixedly supported regardless of the size of the block, or the prediction modes supported depending on the size of the block may be different from those in FIG.
[0142] For example, the number of prediction mode candidates may be adaptive (e.g., the angles between prediction modes are equally spaced but the angles are set differently, such as 9, 17, 33, 65, 129, etc. based on directional mode) or the number of prediction mode candidates may be fixed but have other configurations (e.g., directional mode angles, non-directional types, etc.).
[0143] 4 may be fixed prediction modes supported regardless of the block type, or the prediction modes supported depending on the block type may be different from those shown in FIG.
[0144] For example, the number of prediction mode candidates may be adaptive (e.g., the number of prediction modes derived in the horizontal or vertical direction may be set smaller or larger depending on the width-to-height ratio of the block), or the number of prediction mode candidates may be fixed but have another configuration (e.g., the number of prediction modes derived in the horizontal or vertical direction may be set more finely depending on the width-to-height ratio of the block).
[0145] Alternatively, a larger number of prediction modes may be supported on the side of the block with a longer length, and a smaller number of prediction modes may be supported on the side of the block with a shorter length, and when the block length is long, the prediction mode interval may be supported as modes located to the right of mode 66 in Fig. 4 (e.g., modes having angles of +45 degrees or more based on mode 50, i.e., modes having numbers 67 to 80, etc.) or modes located to the left of mode 2 (e.g., modes having angles of -45 degrees or more based on mode 18, i.e., modes having numbers -1 to -14, etc.). This can be determined according to the ratio of the width and height of the block, and the opposite situation is also possible.
[0146] In the present invention, the description will be focused on the case where the prediction mode is fixedly supported (regardless of any encoding / decoding elements) as shown in FIG. 4, but it is also possible to set a prediction mode that is adaptively supported depending on the encoding setting.
[0147] In addition, when classifying prediction modes, horizontal and vertical modes (modes 18 and 50), some diagonal modes (Diagonal up right <mode 2>, Diagonal down right <mode 34>, Diagonal down left <mode 66>, etc.) can be used as criteria, and this may be a classification method based on some directionality (or angles of 45 degrees, 90 degrees, etc.).
[0148] In addition, some of the directional modes located at both ends (modes 2 and 66) may be reference modes for prediction mode classification, which is a possible example when the intra-frame prediction mode configuration is as shown in Fig. 4. That is, when the prediction mode configuration is adaptive, the reference mode may be changed. For example, mode 2 may be replaced with a mode having a number smaller or larger than 2 (-2, -1, 3, 4, etc.), and mode 66 may be replaced with a mode having a number smaller or larger than 66 (64, 66, 67, 68, etc.).
[0149] In addition, additional prediction modes related to color components may be included in the prediction mode candidates. Next, a color copy mode and a color mode will be described as examples of the prediction modes.
[0150] (Color copy mode)
[0151] Prediction modes related to the manner in which data for generating a prediction block is obtained from a region located in another color space may be supported.
[0152] For example, a prediction mode for acquiring data for generating a prediction block in another color space using correlation between color spaces can be an example of this.
[0153] 6 is a conceptual diagram showing intra-frame prediction for a color copy mode according to an embodiment of the present invention. Referring to FIG. 6, a current block C in a current color space M can be predicted using data of a corresponding area D in another color space N at the same time t.
[0154] Here, the correlation between color spaces may refer to the correlation between Y and Cb, Y and Cr, or Cb and Cr in the case of YCbCr as an example. That is, in the case of a chrominance component (Cb or Cr), a restored block of a luminance component (Y) corresponding to a current block may be used as a prediction block of the current block (chrominance vs. luminance is the default setting in the example described below). Alternatively, a restored block of some chrominance components (Cb or Cr) corresponding to a current block of some chrominance components (Cr or Cb) may be used as a prediction block of the current block.
[0155] Here, the area corresponding to the current block may have the same absolute position in each image in some color formats (e.g., YCbCr4:4:4, etc.), or may have the same relative position in each image in some color formats (e.g., YCbCr4:2:0, etc.). This can be determined by the ratio of width and height according to the color format, and the pixel of the current image and the corresponding pixel of another color space can be obtained by multiplying or dividing each component of the coordinates of the current pixel by the ratio of width and height according to the color format.
[0156] For ease of explanation, the following description focuses on the case of a certain color format, 4:4:4, but it should be understood that the position of the corresponding area in other color spaces can be determined by the horizontal to vertical ratio of the color format.
[0157] In color copy mode, a restored block of another color space can be used as a predicted block as it is, or a block obtained by considering the correlation between color spaces can be used as a predicted block. A block obtained by considering the correlation between color spaces means a block obtained by performing correction on an existing block. In detail, in the formula {P=a*R+b}, a and b mean values used for correction, and R and P can respectively mean a value obtained in another color space and a predicted value of the current color space. Here, P means a block obtained by considering the correlation between color spaces.
[0158] In this embodiment, it is assumed that the data obtained using the correlation between color spaces is used as a predicted value of the current block, but the data can also be used as a correction value applied to a predicted value of an existing current block, i.e., the predicted value of the current block can be corrected with a residual value of another color space.
[0159] In the present invention, the former case will be assumed for explanation, but the present invention is not limited to this, and the same or modified application is possible to the case where the value is used as a correction value.
[0160] The color copy mode may be explicitly or implicitly determined whether it is supported depending on the encoding / decoding settings. Here, the encoding / decoding settings may be defined by one or more combinations of image type, color components, block position / size / shape, block width / length ratio, etc. In addition, when it is explicit, it may include related information in units of sequence, picture, slice, tile, etc. In addition, depending on the encoding / decoding settings, whether the color copy mode is supported may be implicitly determined in some cases, and related information may be explicitly generated in other cases.
[0161] In color copy mode, the correlation information between color spaces (such as a and b) can be generated explicitly or obtained implicitly by the encoding / decoding settings.
[0162] Here, the areas compared (or referenced) to obtain the correlation information can be the current block (C in FIG. 6) and the corresponding area in another color space (D in FIG. 6), or the adjacent area of the current block (left side, upper side, upper left side, upper right side, lower left block, etc., centered on C in FIG. 6) and the adjacent area of the corresponding area in another color space (left side, upper side, upper left side, upper right side, lower left block, etc., centered on D in FIG. 6).
[0163] In the above description, the former case corresponds to an example of explicitly processing related information since correlation information must be directly obtained using data of a corresponding block of the current block. That is, it may be a case where correlation information must be generated since data of the current block has not yet been encoded. The latter case corresponds to an example of implicitly processing related information since correlation information can be indirectly obtained using data of an adjacent area of the corresponding block to an adjacent area of the current block.
[0164] In summary, in the former case, the current block is compared with the corresponding block to obtain correlation information, and in the latter case, the current block is compared with the corresponding block to obtain correlation information, and the correlation information is applied to the corresponding block to obtain data that can be used as a prediction pixel for the current block.
[0165] In the former case, the correlation information can be directly encoded, or the correlation information obtained by comparing adjacent regions can be used as a predicted value, and information on the difference value can be encoded. The correlation information can be information that can be generated when a color copy mode is selected as a prediction mode.
[0166] Here, the latter case can be understood as an example of an implicit case where there is no additional information generated except that the color copy mode is selected as the optimal mode from among the prediction mode candidates, that is, this can be an example that is possible under a setting that supports one correlation information.
[0167] If two or more pieces of correlation information are supported, the color copy mode is selected as the optimal mode, and selection information on the correlation information may be required. As in the above example, it is also possible to combine the explicit and implicit cases according to the encoding / decoding settings.
[0168] In the present invention, the case where correlation information is indirectly acquired will be mainly described. The number of pieces of correlation information acquired here may be N or more (N is an integer equal to or greater than 1, such as 1, 2, 3, etc.). The setting information regarding the number of pieces of correlation information may be included in units of sequences, pictures, slices, tiles, etc. It should be understood that in some of the examples described below, supporting k pieces of correlation information may mean the same as supporting k color copy modes.
[0169] 7 is an example diagram showing corresponding blocks and adjacent areas of each color space in relation to a color copy mode according to an embodiment of the present invention. Referring to FIG. 7, corresponding examples (p and q) between pixels in a current color space M and another color space N are shown, and can be understood as examples that can occur in the case of some color formats (4:2:0). Corresponding relationship 7a can be confirmed for obtaining correlation information, and corresponding relationship 7b can be confirmed for applying a predicted value.
[0170] Next, the acquisition of correlation information in color copy mode will be described. To acquire correlation information, pixel values of pixels in a predefined area (all or part of each adjacent area of a block corresponding to a current block) of each color space can be compared (or used) (i.e., a 1:1 pixel value comparison process is performed). Here, the pixel values to be compared can be acquired based on corresponding pixel positions in each color space. The pixel values can be values derived from at least one pixel in each color space.
[0171] For example, in the case of some color formats such as 4:4:4, the pixel value of one pixel in the chrominance space and the pixel value of one pixel in the luminance space can be used as the pixel value corresponding to the correlation information acquisition process, or in the case of some color formats such as 4:2:0, the pixel value of one pixel in the chrominance space and a pixel value derived from one or more pixels in the luminance space (i.e., obtained by a downsampling process) can be used as the pixel value corresponding to the correlation information acquisition process.
[0172] In detail, in the former case, p[x,y] in the chrominance space can be compared with q[x,y] in the luminance space, where the pixel value can be the brightness value of one pixel as it is, and in the latter case, p[x,y] in the chrominance space can be compared with q[2x,2y], q[2x,2y+1], q[2x+1,2y], q[2x+1,2y+1], etc. in the luminance space.
[0173] Here, since a 1:1 pixel value comparison must be performed, in the case of a luminance space, one of the plurality of pixels can be used as a value to be compared with the pixel value of a chrominance pixel. That is, the brightness value of one of the plurality of pixels is used as is. Or, one pixel value can be derived from two or more pixels (k pixels, k being an integer equal to or greater than 2, such as 2, 4, 6, etc.) among the plurality of pixels and used as a value to be compared. That is, a weighted average (weights can be assigned to each pixel evenly or non-evenly) can be applied to two or more pixels.
[0174] In the case where there are multiple corresponding pixels as in the above example, the pixel value of one previously set pixel or a pixel value derived from two or more pixels can be used as a comparison value. Here, one of the above two methods for deriving pixel values to be compared in each color space according to encoding / decoding settings can be used alone or in combination.
[0175] Next, it is assumed that the pixel value of one pixel in the current color space is used for comparison, and one or more pixels in another color space can be used to derive a pixel value. For example, assume that the color format is YCbCr4:2:0, the current color space is a chrominance space, and the other color space is a luminance space. The method for deriving a pixel value will be described focusing on the other color space.
[0176] For example, it may be determined based on the shape of the block (ratio of width to height). For a detailed example, p[x, y] in the chrominance space adjacent to the longer side of the current block (or the block to be predicted) may be compared with q[2x, 2y] in the luminance space, and p[x, y] in the chrominance space adjacent to the shorter side may be compared with an average of q[2x, 2y] and q[2x+1, 2y] in the luminance space.
[0177] Here, adaptive settings are possible, such as applying the above content to some block shapes (rectangles) regardless of the width-to-height ratio, or applying it only when the width-to-height ratio is equal to or exceeds a certain ratio (k:1 or 1:k, for example, k is 2 or more, such as 2:1, 4:1, etc.).
[0178] For example, it can be determined based on the size of the block. As a detailed example, the size of the current block is a certain size (M×N. For example, 2 m ×2 n If the size is greater than or equal to a certain size (where m and n are integers greater than or equal to 1, such as 2 through 6), p[x, y] in the chrominance space can be compared with q[2x+1, 2y] in the luminance space, and if the size is less than or equal to a certain size, p[x, y] in the chrominance space can be compared with the average of q[2x, 2y], q[2x, 2y+1] in the luminance space.
[0179] Here, the boundary value for size comparison may be one as in the above example, or may be two or more (M 1 ×N 1 , M 2 ×N 2 Adaptive settings such as changing the default setting are possible.
[0180] The above examples are only some of the cases that can be considered in terms of the amount of calculation, and various modified examples including the opposite cases to the above examples are possible.
[0181] For example, it may be determined according to the position of the block. As a detailed example, if the current block is located inside a preset region (assumed to be the maximum coding block in this example), p[x, y] in the chrominance space may be compared with an average of q[2x, 2y], q[2x+1, 2y], q[2x, 2y+1], and q[2x+1, 2y+1] in the luminance space, and if the current block is located on the boundary of the preset region (assumed to be the upper left boundary in this example), p[x, y] in the chrominance space may be compared with q[2x+1, 2y+1] in the luminance space. The preset region may refer to a region set based on a slice, a tile, a block, etc. In particular, it may be acquired based on an integer multiple of a slice, a tile, or a maximum coding / prediction / transformation block.
[0182] As another example, if the current block is located on a part of the boundary of the region (assuming it is the upper boundary in this example), p[x, y] in the chrominance space adjacent to the part of the boundary (top) can be compared with q[2x+1, 2y+1] in the luminance space, and p[x, y] in the chrominance space adjacent to the interior (left) of the region can be compared with the average of q[2x, 2y], q[2x+1, 2y], q[2x, 2y+1], q[2x+1, 2y+1] in the luminance space.
[0183] The above examples are only some of the cases that can be considered in terms of memory, and various modifications including the opposite cases to the above examples are possible.
[0184] Based on the above examples, various cases regarding the derivation of pixel values to be compared in each color space have been described. As in the above examples, the setting of pixel value derivation for obtaining the correlation information can be determined by considering various encoding / decoding factors as well as the size / shape / position of the block.
[0185] Based on the above example, the area to be compared to obtain correlation information is described as using one or two reference pixel lines of the current block and the corresponding block as shown in Fig. 7. That is, in the case of YCbCr4:4:4, one reference pixel line is used for each, and in the case of other formats, some color spaces use one reference pixel line as shown in Fig. 7.<color N> However, the present invention is not limited to this and various modifications are possible.
[0186] The following description focuses on the reference pixel lines of the current color space, and it should be understood that for other color spaces, the reference pixel lines can be determined according to the color format, for example, the same number of reference pixel lines can be used or twice as many reference pixel lines can be used.
[0187] In the color copy mode of the present invention, k reference pixel lines (where k is an integer equal to or greater than 1, such as 1 or 2) can be used (or compared) to obtain correlation information. In addition, the k reference pixel lines can be used either fixedly or adaptively. Various examples of setting the number of reference pixel lines will now be described.
[0188] For example, it can be determined based on the shape of the block (ratio of width to height). For a detailed example, two adjacent reference pixel lines on the longer side of the current block can be used, and one adjacent reference pixel line on the shorter side can be used.
[0189] Here, the above content can be applied to some block shapes (rectangles) regardless of the ratio of width to height, or can be applied only when the ratio of width to height is equal to or exceeds a certain ratio (k:1 or 1:k, where k is an integer equal to or greater than 2, for example, 2:1, 4:1, etc.). In addition, there are two or more boundary values for the ratio of width to height, and it is possible to extend the range of values, such as using two reference pixel lines adjacent to the longer side (or shorter side) in the case of 2:1 or 1:2, and using three reference pixel lines adjacent to the longer side (or shorter side) in the case of 4:1 or 1:4.
[0190] The above example may be an example where the longer side (or shorter side) uses s reference pixel lines and the shorter side (or longer side) uses t reference pixel lines depending on the ratio of width to height, where s is greater than or equal to t (i.e., s and t are integers greater than or equal to 1).
[0191] For example, it can be determined based on the size of the block. As a detailed example, the size of the current block is a certain size (M×N. For example, 2 m ×2 n If m and n are greater than or equal to a certain size (m and n are integers greater than or equal to 1, such as 2 to 6), two reference pixel lines can be used, and if m and n are less than or equal to a certain size, one reference pixel line can be used.
[0192] Here, the boundary value for size comparison may be one as in the above example, or may be two or more (M 1 ×N 1 , M 2 ×N 2 Adaptive settings such as changing the default setting are possible.
[0193] For example, it can be determined according to the position of the block. For a detailed example, if the current block is located inside a preset region (which can be guided by the previous description related to obtaining correlation information, assumed to be the largest coding block in this example), two reference pixel lines can be used, and if the current block is located on the boundary of the preset region (assumed to be the upper left boundary in this example), one reference pixel line can be used.
[0194] As another example, if the current block is located at a partial boundary of the region (assuming the upper boundary in this example), one reference pixel line adjacent to the partial boundary (top) can be used, and two reference pixel lines adjacent inside the region (left side) can be used.
[0195] The above examples are only some of the cases that can be considered in terms of accuracy of correlation information and memory, and various modified examples including the opposite cases to the above examples are possible.
[0196] Based on the above examples, various cases regarding the setting of the reference pixel line used to obtain correlation information in each color space have been described. As in the above examples, the setting of the reference pixel line for obtaining correlation information can be determined taking into consideration various encoding / decoding factors as well as the size / shape / position of the block.
[0197] Next, other cases of the area to be compared (or referenced) for obtaining correlation information will be described. The area to be compared may be adjacent pixels at positions such as the left side, upper side, upper left side, upper right side, lower left side, etc. adjacent to the current block in the current color space.
[0198] Here, the area to be compared may be set to include all blocks at the left, upper, upper left, upper right, and lower left positions. Alternatively, the reference pixel area may be configured as a combination of blocks at some positions. For example, the area to be compared may be configured as a combination of adjacent blocks such as left / upper / left+upper / left+upper+upper left / left+lower left / upper+upper right / left+upper left+lower left / upper+upper left+upper right / left+upper+upper right / left+upper+lower left.
[0199] FIG. 8 is an exemplary diagram of a region setting for obtaining correlation information in a color copy mode according to an embodiment of the present invention. In the cases of a to e in FIG. 8, it may correspond to the above-mentioned examples (left side + upper side, upper side + upper right side, left side + lower left side, left side + upper side + upper right side, left side + upper + lower left side), and it is possible to use an example (f, g in FIG. 8) in which a block at a certain position is divided into one or more sub-blocks, some of which are set as a region for obtaining correlation information. That is, it is possible to set a region for obtaining correlation information by one or more sub-blocks located in a certain direction. Alternatively, it is possible to set a region for obtaining correlation information by one or more blocks located in a certain direction (a) and one or more sub-blocks located in a certain direction (b) (here, a and b mean other directions). In addition, it is also possible to use an example (h, i in FIG. 8) in which a region for obtaining correlation information is set in non-consecutive blocks.
[0200] In summary, the regions to be compared for obtaining correlation information may be configured as predefined regions, or may be configured as various combinations of some regions. That is, the regions to be compared may be configured either fixedly or adaptively depending on the encoding / decoding settings.
[0201] Next, various examples of which direction adjacent regions are configured as reference regions with respect to a current block in a current color space will be described. Here, it is assumed that which direction adjacent regions are configured as reference regions in a corresponding block in another color space is determined according to the reference region configuration of the current color block. It is also assumed that the basic reference region is configured from the left and upper blocks.
[0202] For example, it may be determined based on the shape of the block (ratio of width to height). For a detailed example, if the current block is long in the horizontal direction, the left, upper, and upper right blocks may be set as the reference region, and if the current block is long in the vertical direction, the left, upper, and lower left blocks may be set as the reference region.
[0203] Here, the above content can be applied to some block shapes (rectangles) regardless of the ratio of width to height, or can be applied only when the ratio of width to height is equal to or exceeds a certain ratio (k:1 or 1:k, where k is an integer equal to or greater than 2, for example, 2:1, 4:1, etc.). In addition, there are two or more boundary values for the ratio of width to height, and in the case of 2:1 (or 1:2), the left, upper, and upper right blocks (or the left, upper, and lower left) are set as the reference region, and in the case of 4:1 (or 1:4), the upper and upper right blocks (or the left and lower left) are set as the reference region.
[0204] For example, it can be determined based on the size of the block. As a detailed example, the size of the current block is a certain size (M×N. For example, 2 m ×2 n If m and n are integers of 1 or more, such as 2 to 6, and are greater than or equal to a certain size, the left and upper blocks are set as the reference region, and if m and n are less than or equal to a certain size, the left, upper, and upper left blocks can be set as the reference region.
[0205] Here, the boundary value for size comparison may be one as in the above example, or may be two or more (M 1 ×N 1 , M 2 ×N 2 Adaptive settings such as changing the default setting are possible.
[0206] For example, it can be determined according to the position of the block. As a detailed example, if the current block is located inside a preset region (which can be derived from the previous description related to obtaining correlation information, and is assumed to be the largest coding block in this example), the left, upper, upper left, upper right, and lower left blocks can be set as reference regions, and if the current block is located on the boundary of the preset region (which is assumed to be the upper left boundary in this example), the left and upper blocks can be set as reference regions.
[0207] As another example, if the current block is located at a certain boundary of the region (assumed to be the upper boundary in this example), the adjacent left and lower left blocks inside the region, except for the block adjacent to the certain boundary (upper boundary), can be set as the reference region. That is, the left and lower left blocks can be set as the reference region.
[0208] The above examples are only some of the cases that can be considered in terms of the amount of calculation, memory, etc., and various modified examples including the opposite cases to the above examples are possible.
[0209] Based on the above examples, various cases have been described for setting a reference area used to obtain correlation information in each color space. As in the above examples, the setting of a reference area for obtaining correlation information can be determined taking into consideration various encoding / decoding factors as well as the size / shape / position of a block.
[0210] Also, the compared region may be pixels adjacent to the current block in the current color space, where all or some of the reference pixels may be used to obtain correlation information.
[0211] For example, (with reference to color M in FIG. 7) if the current block is a block (i.e., 8×8) having a pixel range of (a, b) to (a+7, b+7), the area to be compared (corresponding blocks are omitted as the explanation can be guided depending on the color format) is assumed to be one reference pixel line of the blocks to the left and above the current block.
[0212] Here, the area to be compared may include all pixels within the ranges of (a, b-1) to (a+7, b-1) and (a-1, b) to (a-1, b+7). Or, it may include partial pixels within the ranges, that is, (a, b-1), (a+2, b-1), (a+4, b-1), (a+6, b-1), (a-1, b), (a-1, b+2), (a-1, b+4), (a-1, b+6). Or, it may include partial pixels within the range, that is, (a, b-1), (a+4, b-1), (a-1, b), (a-1, b+4).
[0213] The above example can be applied for the purpose of reducing the amount of calculation required for correlation acquisition. As with the many examples described above, the setting of reference pixel sampling of the compared area for correlation information acquisition can take into account various encoding / decoding factors such as block size / shape / position, and related application examples can be derived from the previous examples, so detailed description will be omitted.
[0214] Based on the above-mentioned various examples, various factors that affect correlation information acquisition (corresponding pixel value induction, number of reference pixel lines, reference area direction setting, reference pixel sampling, etc.) have been described. Various cases are possible in which the above examples affect correlation information acquisition, either alone or in combination.
[0215] The above description can be understood as a pre-setting process for acquiring one correlation information. Also, as mentioned above, one or more correlation information can be supported by the encoding / decoding setting. Here, two or more correlation information can be supported by setting two or more of the pre-settings (i.e., combinations of factors that affect the acquisition of correlation information).
[0216] In summary, parameter information based on correlation information can be derived from adjacent regions of the current block and the corresponding block. That is, at least one parameter (e.g.,<a1、b1> ,<a2、b2> ,<a3、b3> , etc.), which can be used as values to multiply or add to the pixels of the reconstructed block in other color spaces.
[0217] Next, a linear model applied in the color copy mode will be described. By applying the parameters obtained by the above process, prediction based on the following linear model can be performed.
[0218] pred_sample_C(i, j)=a×rec_sample_D(i, j)+b
[0219] In the above formula, pred_sample_C means a predicted pixel value of the current block in the current color space, and rec_sample_D means a restored pixel value of the corresponding block in another color space. a and b can be obtained by minimizing the regression error between the adjacent areas of the current block and the adjacent areas of the corresponding block, and can be calculated by the following formula.
number
[0220] In the above formula, D(n) means the adjacent area of the corresponding block, C(n) means the adjacent area of the current block, and N means a value set based on the width or height of the current block (in this example, assumed to be twice the minimum value of the width and height).
[0221] Also, various methods such as a straight-line equation that obtains correlation information based on the minimum and maximum values of adjacent areas of each color space can be used. Here, the model for obtaining correlation information can be one pre-set model or one of multiple models can be selected. Here, selecting one of multiple models means that model information can be considered as an encoding / decoding element for parameter information based on correlation information. In other words, when multiple parameters are supported, it can mean that even if the remaining correlation information related settings are the same, different parameter information can be distinguished by using different models for obtaining correlation.
[0222] For some color formats (not 4:4:4), one pixel in the current block can correspond to one or more (two, four, etc.) pixels in the corresponding block. For example, for 4:2:0, p[x, y] in chrominance space can correspond to q[2x, 2y], q[2x, 2y+1], q[2x+1, 2y], q[2x+1, 2y+1], etc. in luma space.
[0223] For one predicted pixel value, a pixel value (or predicted value) of one pixel already set from the corresponding pixels, or one pixel value from two or more pixels, can be derived (7b). That is, for obtaining one predicted pixel value, a restored value before applying correlation information can be obtained from one or more corresponding pixels in another color space. Various cases are possible depending on the encoding / decoding setting, and the relevant explanation can be derived in the part related to the corresponding pixel value derivation (7a) for obtaining correlation information, so detailed explanation will be omitted. However, the same or different settings can be applied to 7a and 7b.
[0224] (Color mode)
[0225] It is possible to support prediction modes related to a method of obtaining a prediction mode for generating a prediction block from a region located in another color space.
[0226] For example, a prediction mode for a method of obtaining a prediction mode for generating a prediction block in another color space using correlation between color spaces can be an example of this. That is, a color mode does not have a specific prediction direction or prediction method, but can be a mode that uses an existing prediction direction and method and is adaptively determined according to the prediction mode of a corresponding block in another color space.
[0227] Here, various color modes can be obtained by setting block division.
[0228] For example, in a setting where block division for some color components (chrominance) is implicitly determined by the block division result for some color components (luminance) (i.e., block division for luminance component is explicitly determined), one block of some color components (chrominance) may correspond to one block of some color space (luminance). Therefore, (assuming the case of 4:4:4. For other formats, the explanation of this example can be guided by the ratio of width to height), if a current block (chrominance) has a pixel range of (a, b) to (a+m, b+n), any pixel position within the pixel range of (a, b) to (a+m, b+n) of the corresponding block (luminance) indicates one block, so one prediction mode can be obtained from the block including the corresponding pixel.
[0229] Alternatively, when individual block division is supported for each color component (i.e., when block division for each color space is explicitly determined), one block of some color components (chrominance) may correspond to one or more blocks of some color space (luminance). Therefore, even if the current block (chrominance) has the same pixel range as the above example, the corresponding block (luminance) may be composed of one or more blocks depending on the block division result. Therefore, it is also possible to obtain other prediction modes (i.e., one or more modes) from the corresponding block indicated by the corresponding pixel depending on the pixel position within the pixel range of the current block.
[0230] If one color mode is supported in the intra prediction mode candidate group for the chrominance component, it is possible to set from which position of the corresponding block the prediction mode is to be taken.
[0231] For example, prediction modes can be obtained from positions such as the center, upper left, upper right, lower left, and lower right of the corresponding block. That is, prediction modes are obtained in the above order, but if the corresponding block is not used (for example, the coding mode is Inter), a prediction mode from a position corresponding to the next order can be obtained. Alternatively, a prediction mode with a high frequency (2 or more times) can be obtained from the block in the above position.
[0232] Alternatively, when multiple color modes are supported, it is possible to set from which prediction mode to bring according to the priority order. Alternatively, a combination is possible in which some prediction modes are brought according to the priority order and some prediction modes with high frequency are brought from the block at the position. Here, the above-mentioned case of priority order is one example, and various modified examples are possible.
[0233] The color mode and the color copy mode may be prediction modes that can be supported for the chrominance component. For example, a prediction mode candidate group for the chrominance component may be configured including horizontal, vertical, DC, planar, diagonal modes, etc. Alternatively, a prediction mode candidate group for the intra-screen prediction mode may be configured including the color mode and the color copy mode.
[0234] That is, it may be composed of directional + non-directional + color mode, or directional + non-directional + color copy mode, or it may be composed of directional + non-directional + color mode + color copy mode, or it may include modes for additional color difference components.
[0235] Depending on the encoding / decoding settings, it can be determined whether the color mode and color copy mode are supported, and this can be processed implicitly or explicitly. Or, a combination of explicit and implicit configuration can be processed. In addition, detailed settings related to the color mode and color copy mode (e.g., the number of modes supported) can be included and processed implicitly or explicitly.
[0236] For example, the related information may explicitly include related information in units of a sequence, a picture, a slice, a tile, a block, etc., or may be implicitly determined by various encoding / decoding factors (e.g., image type, block position, block size, block shape, block width-to-height ratio, etc.), or may be implicitly determined under some conditions or explicitly generated under some conditions depending on the encoding / decoding factors.
[0237] 9 is a diagram illustrating an example of a configuration of reference pixels used in intra prediction according to an embodiment of the present invention, in which the size and shape (M×N) of a prediction block can be obtained by a block division unit.
[0238] Block range information defined as the sizes of the minimum block and the maximum block for intra prediction may include information related to units such as a sequence, a picture, a slice, a tile, etc. Generally, size information may be set by specifying the width and the height (e.g., 32x32, 64x64, etc.), but size information may also be set in the form of the product of the width and the height. For example, when the product of the width and the height is 64, the size of the minimum block may correspond to 4x16, 8x8, 16x4, etc.
[0239] Also, the size information can be set by specifying the width and height, or the size information can be set in the form of a product. For example, if the product of the width and height is 4096 and the maximum value of one of the two lengths is 64, the maximum block size can be 64 x 64.
[0240] As in the above example, the size and shape of the predicted block can be finally determined by combining the size information of the minimum block and the maximum block as well as the block division information. In the present invention, the product of the width and the height of the predicted block must be greater than or equal to s (e.g., s is a multiple of 2 such as 16, 32, etc.), and one of the width and the height must be greater than or equal to k (e.g., k is a multiple of 2 such as 4, 8, etc.). In addition, the width and the height of the block can be defined under the setting that they are smaller than or equal to v and w (e.g., v and w are a multiple of 2 such as 16, 32, 64, etc.), but is not limited thereto, and various block ranges can be set.
[0241] Intra-frame prediction is generally performed in units of prediction blocks, but can be performed in units of coding blocks, transform blocks, etc., depending on the setting of the block division unit. After checking the block information, the reference pixel construction unit can construct reference pixels to be used in the prediction of the current block. Here, the reference pixels are stored in a temporary memory (for example, an array <array>The temporary memory can be managed by a first array, a second array, etc., and can be generated and removed for each intra-prediction process of a block, and the size of the temporary memory can be determined according to the configuration of the reference pixels.
[0242] In this example, it is assumed that the left, upper, upper left, upper right, and lower left blocks are used to predict the current block, but the present invention is not limited to this, and other configurations of block candidates can be used to predict the current block. For example, the neighboring block candidates for the reference pixels may be an example of a raster or Z scan, and some of the candidates may be removed depending on the scan order, or other block candidates (for example, right, lower, lower right blocks, etc. may be added) may be included.
[0243] In addition, when a certain prediction mode (color copy mode) is supported, a certain area of another color space can be used for predicting the current block, and therefore this can also be considered as a reference pixel. The existing reference pixels (spatially adjacent areas of the current block) and the additional reference pixels can be managed as one or managed separately (e.g., reference pixel A and reference pixel B. In other words, the reference pixel memory can be named separately just as the temporary memory is used separately).
[0244] For example, the temporary memory of the basic reference pixels may have a size of <2×blk_width+2×blk_height+1> (based on one reference pixel line), and the temporary memory of the additional reference pixels may have a size of <2×blk_width+2×blk_height+1> (in the case of 4:4:4).<blk_width×blk_height> (In the case of 4:2:0, blk_width / 2×blk_height / 2 is required.) The above size of the temporary memory is merely an example and is not limited thereto.
[0245] In addition, to obtain correlation information, adjacent areas of the current block and the corresponding block to be compared (or referenced) can be managed as reference pixels, which means that additional reference pixels can be managed according to the color copy mode.
[0246] In summary, adjacent areas of the current block can be included as reference pixels for intra-screen prediction of the current block, and depending on the prediction mode, even a corresponding block in another color space and its adjacent areas can be included as reference pixels.
[0247] 10 is a conceptual diagram showing blocks adjacent to a target block of intra prediction according to an embodiment of the present invention. In detail, the left side of FIG. 10 shows blocks adjacent to a current block in a current color space, and the right side shows a corresponding block in another color space. For convenience of explanation, the following description will be given on the assumption that the blocks adjacent to the current block in the current color space are basic reference pixel configurations.
[0248] As shown in Figure 9, the reference pixels used for predicting the current block can be configured from adjacent pixels in the left, upper, upper left, upper right, and lower left blocks (Ref_L, Ref_T, Ref_TL, Ref_TR, and Ref_BL in Figure 9). Here, the reference pixels are generally configured from pixels in the neighboring block closest to the current block (a in Figure 9, which is referred to as the reference pixel line), but other pixels (b in Figure 9 and pixels on other outer lines) can also be configured as reference pixels.
[0249] Pixels adjacent to the current block can be classified into at least one reference pixel line, and the pixel closest to the current block can be classified as ref_0 {e.g., pixels whose distance between the boundary pixel of the current block and the pixel is 1. p(-1,-1) to p(2m-1,-1), p(-1,0) to p(-1,2n-1)}, the next adjacent pixel {e.g., pixels whose distance between the boundary pixel of the current block and the pixel is 2. p(-2,-2) to p(2m,-2), p(-2,-1) to p(-2,2n)} can be classified as ref_1, and the next adjacent pixel {e.g., pixels whose distance between the boundary pixel of the current block and the pixel is 3. p(-3,-3) to p(2m+1,-3), p(-3,-2) to p(-3,2n+1)} can be classified as ref_2, etc. That is, the reference pixel lines can be classified according to the distance between the boundary pixel of the current block and the adjacent pixel.
[0250] Here, the number of supported reference pixel lines may be N or more, and N may be an integer of 1 or more, such as 1 to 5. Here, the reference pixel line candidates are generally included in the reference pixel line candidate group in order starting from the reference pixel line closest to the current block, but are not limited thereto. For example, when N is 3,<ref_0、ref_1、ref_2> or<ref_0、ref_1、ref_3> ,<ref_0、ref_2、ref_3> ,<ref_1、ref_2、ref_3> The candidate set may be formed in a non-sequential manner or by excluding the most adjacent reference pixel line.
[0251] Prediction can be performed using all reference pixel lines in the candidate set, or prediction can be performed using a portion (one or more) of the reference pixel lines.
[0252] For example, one of a plurality of reference pixel lines may be selected according to encoding / decoding settings, and intra-frame prediction may be performed using the corresponding reference pixel line, or two or more of the plurality of reference pixel lines may be selected, and intra-frame prediction may be performed using the corresponding reference pixel line (e.g., by applying a weighted average to data of each reference pixel line).
[0253] Here, the selection of the reference pixel line may be determined implicitly or explicitly. For example, the implicit case means that it is determined according to an encoding / decoding setting defined by one or a combination of two or more factors such as an image type, a color component, a size / shape / position of a block, etc., and the explicit case means that reference pixel line selection information may be generated in units such as a block.
[0254] In the present invention, the case where intra-screen prediction is performed using the most adjacent reference pixel line will be mainly described, but it should be understood that the various embodiments described below can be applied in the same or similar manner when multiple reference pixel lines are used.
[0255] The reference pixel construction unit for intra-screen prediction of the present invention may include a reference pixel generation unit, a reference pixel interpolation unit, a reference pixel filter unit, etc., and may include all or a part of the above components.
[0256] The reference pixel configuration unit can check the availability of the reference pixels and classify the reference pixels into available and unavailable ones. Here, the availability of the reference pixels is determined to be unavailable when at least one of the following conditions is satisfied. Of course, the availability of the reference pixels can be determined based on additional conditions not mentioned in the example described below, but the present invention will be described assuming that the availability of the reference pixels is limited to the exemplary conditions described below.
[0257] For example, if it is located outside the picture boundary, if it does not belong to the same division unit as the current block (e.g., units that cannot refer to each other, such as slices and tiles. However, exceptions are made even if the units, such as slices or tiles, have the characteristic that they can refer to each other, even if they are not the same division unit), if encoding / decoding is not completed, it can be determined that it is unusable. In other words, if any of the above conditions are not met, it can be determined that it is usable.
[0258] In addition, the use of reference pixels may be restricted depending on the encoding / decoding settings. For example, even if it is determined that the use of reference pixels is possible according to the above conditions, the use of reference pixels may be restricted depending on whether or not restricted intra prediction (e.g., constrained_intra_pred_flag) is performed. The restricted intra prediction may be performed when it is desired to prohibit the use of a block reconstructed by referring to another image as a reference pixel when performing error-resilient encoding / decoding in an external context such as a communication environment.
[0259] When constrained intra prediction is deactivated (e.g., I picture type, or constrained_intra_pred_flag is set to 0 for P or B picture types), reference pixel candidate blocks (but only if they meet the aforementioned conditions, such as being located inside the picture boundary) may be usable.
[0260] Alternatively, when constrained intra prediction is activated (e.g., constrained_intra_pred_flag is set to 1 for P or B video type), the reference pixel candidate block can be determined whether it can be used depending on the coding mode (Intra or Inter). Generally, it can be used in Intra mode and cannot be used in Inter mode. In the above example, it is assumed that the use is determined depending on the coding mode, but the use can also be determined depending on various other coding / decoding factors.
[0261] Since reference pixels are composed of one or more blocks, after checking the reference pixel possibility, they can be classified into three cases: <all usable>, <some usable>, and <all unusable>. In the remaining cases except for the case where all can be used, reference pixels at unusable candidate block positions can be filled or generated (A). Alternatively, reference pixels at unusable candidate block positions cannot be used in the prediction process, and prediction mode encoding / decoding can be performed excluding prediction modes that perform prediction from reference pixels at the corresponding positions (B).
[0262] If the reference pixel candidate block is available, the pixel at the corresponding position can be included in the reference pixel memory of the current block by directly copying the pixel data or by using a process such as reference pixel filtering or reference pixel interpolation.
[0263] When the reference pixel candidate block is unusable, processing can be performed under reference pixel processing setting A or B. Next, a processing example when the reference pixel candidate block is unusable by each setting will be described.
[0264] (A) When the reference pixel candidate block is unavailable, the pixel at the corresponding position obtained by the reference pixel generating process can be included in the reference pixel memory of the current block.
[0265] Next, as an example of the reference pixel generating process, a method of generating reference pixels at unusable positions will be described.
[0266] For example, a reference pixel may be generated using an arbitrary pixel value. Here, the arbitrary pixel value may be one pixel value (e.g., a minimum value, a maximum value, a median value, etc. of a pixel value range) belonging to a pixel value range (e.g., a pixel value range based on a bit depth or a pixel value range according to pixel distribution in a corresponding image). In particular, this may be an example that is applicable when all of the reference pixel candidate blocks are unavailable, but is not limited thereto, and may also be applicable when only a portion of the reference pixel candidate blocks are unavailable.
[0267] Alternatively, the reference pixels may be generated from an area where the encoding / decoding of the image has been completed. In particular, the reference pixels may be generated based on at least one usable block (or usable reference pixel) adjacent to the unusable block. Here, at least one of methods such as extrapolation, interpolation, and copying may be used.
[0268] (B) When the reference pixel candidate block is unavailable, the use of the prediction mode using the pixel at the corresponding position can be restricted. For example, when the reference pixel at the TR position in Fig. 9 is unavailable, the use of the 51st to 66th modes (Fig. 4) that perform prediction using the pixel at the corresponding position can be restricted, and the use of the 2nd to 50th modes (vertical modes) that perform prediction using the reference pixels at the T, TL, L, and BL positions that are not the reference pixel at the TR position can be permitted (in this example, the explanation is limited to the directional modes).
[0269] As another example, if reference pixels at all positions are unavailable, there may be no allowed prediction modes. In this case, a prediction block may be generated using any pixel value, as in some configurations of setting A, and a previously set prediction mode (e.g., DC mode, etc.) may be set as the prediction mode of the corresponding block for reference in the prediction mode encoding / decoding process of the next block. In other words, the prediction mode encoding / decoding process may be implicitly omitted.
[0270] The above example may be related to a prediction mode encoding / decoding process. The prediction mode encoding / decoding unit of the present invention will be described below assuming that it supports setting A. If setting B is supported, a part of the configuration of the prediction mode encoding / decoding unit may be changed. The case where reference pixels at all positions are unavailable has already been described above, so the case where reference pixels at some positions are unavailable will be described below.
[0271] For example, assume that the MPM candidate group includes a prediction mode of a neighboring block. If the prediction mode of the neighboring block is a prediction mode that uses reference pixels at a block position that is unavailable in the current block, a process of removing the corresponding mode from the MPM candidate group may be added. That is, a process of checking unavailable modes may be added to a process of checking redundancy, which will be described later, in the prediction mode encoding / decoding unit. Here, the unavailable mode may be specified by various definitions, but in this example, it is assumed to be a prediction mode that uses reference pixels at a block position that is unavailable. Therefore, a process of constructing an MPM candidate group of prediction modes may be performed according to a priority order, and whether or not to include the prediction mode in the MPM candidate group may be determined through a process of checking redundancy and / or a process of checking unavailable modes. Here, if the prediction mode of the corresponding order cannot pass the checking process, the prediction mode of the next priority order may be a candidate for the MPM candidate group construction process.
[0272] Alternatively, assume that when reference pixels at TR, T, and TL positions in FIG. 9 are unavailable, the use of modes 19 to 66 that perform prediction using pixels at the corresponding positions is restricted. In this case, modes 2 to 18 that perform prediction using reference pixels at L and BL positions may be available prediction modes. In this case, assuming that the number of candidates included in the MPM candidate group is six, the number of candidates included in the non-MPM candidate group may be 12. Here, just as it is inefficient to maintain the number of MPM candidates at six (a larger number than the total mode), when prediction mode use is restricted due to unavailable reference pixels, entropy encoding / decoding setting changes such as adjustment of the number of MPM candidates (e.g., p→q, p>q) and binarization (e.g., variable length binarization A→variable length binarization B, etc.) may also occur. In other words, this may be a situation in which adaptive prediction mode encoding / decoding is supported, and detailed description thereof will be omitted.
[0273] In addition, since a prediction mode whose use is restricted by unavailable reference pixels is also unlikely to occur in the non-MPM candidate group, it may not be necessary to include the corresponding mode in the candidate group. This may be a situation in which the entropy encoding / decoding settings such as number adjustment (e.g., s to t, s>t) and binarization (e.g., fixed length binarization to variable length binarization) of the non-MPM candidate group support adaptive prediction mode encoding / decoding that can be changed like the MPM candidate group.
[0274] In the above examples, various processing examples have been described for the case where reference pixels are unavailable, which may occur not only in the general prediction mode but also in the color copy mode.
[0275] Next, when a color copy mode is supported, reference pixels are classified into usable reference pixels and unusable reference pixels based on the availability of the reference pixels, and various processing examples for the classification will be described.
[0276] 11 is an example diagram for explaining the availability of reference pixels in a color copy mode according to an embodiment of the present invention, assuming that the left and upper blocks of a current block (current color space) and a corresponding block (other color space) represent areas used to obtain correlation information and are in some color format (YCbCr4:4:4).
[0277] It has already been mentioned that reference pixel availability can be determined based on the position of the current block (e.g., whether it is located outside the picture boundary, etc.) Figure 11 shows various examples of reference pixel availability that can be had based on the position of the current block.
[0278] In Figure 11, the reference pixel availability of the current block means the same as that of the corresponding block (result). However, it is assumed that the current color space is divided into the same division unit (tile, slice, etc.) as the current color space in another color space (however, it is necessary to consider the component ratio of the color format).
[0279] FIG. 11A shows the case where all reference pixels are usable, FIG. 11B and FIG. 11C show the case where some of the reference pixels are usable (the upper and left blocks, respectively), and FIG. 11D shows the case where none of the reference pixels are usable.
[0280] Except for case a in Figure 11, at least one reference pixel belongs to the case of being unusable, and therefore a process for this is necessary. In the case of setting reference pixel process A, a process of filling the unusable area can be performed by a reference pixel generation process. Here, the reference pixels of the current block can be processed by the reference pixel generation process (general intra-frame prediction) already described above.
[0281] The reference pixels of the corresponding block may be processed in the same way as the current block or differently. For example, if the reference pixel at L2 position (see FIG. 10) among the reference pixels of the current block is unavailable, it may be generated by an arbitrary pixel value or by available reference pixels. In detail, the available reference pixels may be located to the left / right or above / below the unavailable reference pixels (L1, L3, etc. in this example), or may be located on the same reference pixel line (R3, etc. in this example).
[0282] Meanwhile, the reference pixels of the corresponding block may be generated by any pixel value or by usable reference pixels. However, the position of the usable reference pixels may be the same as or different from that of the current block. In particular, the usable reference pixels may be located not only to the left / right / upper / lower directions of the unusable reference pixels but also in various directions such as the upper left, upper right, lower left, and lower right. In the case of the current block, since the encoding / decoding has not yet been completed, the pixels a to p in FIG. 10 do not belong to the usable reference pixels, but in the case of the corresponding block, since the encoding / decoding has been completed, the pixels aa to pp in FIG. 10 can also belong to the usable reference pixels. Therefore, the reference pixels at the unusable positions can be generated by various methods such as interpolation, extrapolation, copying, and filtering of the usable reference pixels.
[0283] Through the above process, reference pixels at unusable positions such as b to d in FIG. 11 can be generated and stored in the reference pixel memory, and the reference pixels at the corresponding positions can be used to obtain correlation information as in a in FIG. 11.
[0284] Next, the case of setting the reference pixel processing B will be described. It is possible to restrict the use of reference pixels at unusable positions. In addition, it is possible to restrict the use of prediction modes that perform prediction from reference pixels at unusable positions (applying adaptive prediction mode encoding / decoding, etc.), or to perform other processing.
[0285] First, a case where the use of reference pixels at unusable positions is restricted will be described. As shown in a of Figure 11, the left and upper blocks of the current block and the corresponding block must be used to obtain correlation information, while b and c of Figure 11 correspond to cases where some reference pixels are unusable. Here, correlation information can be obtained using usable reference pixels without using unusable reference pixels. Meanwhile, consideration must be given to cases where there is insufficient data for obtaining correlation information.
[0286] For example, if the left (N) and upper (M) blocks of the current block (M×N) are correlation information acquisition areas, the number of available reference pixels is k (k is a number greater than 0).<M+N> If the number of available reference pixels is less than k, the reference pixel cannot be used in the correlation information acquisition process.
[0287] Alternatively, if the number of available reference pixels in the left block is greater than / exceeds p (p is greater than 0 but less than N) and the number of available reference pixels in the upper block is greater than / exceeds q (q is greater than 0 but less than M), the corresponding reference pixels can be used in the correlation information acquisition process. If the number of available reference pixels in the left block is less than / equal to p or the number of available reference pixels in the upper block is less than / equal to q, the corresponding reference pixels cannot be used in the correlation information acquisition process.
[0288] In the former case, it may be classification according to boundary value conditions in the whole region for obtaining correlation information, and in the latter case, it may be classification according to boundary value conditions in a part (partial) region for obtaining correlation information. In the latter case, it may be an example of a case where adjacent regions for obtaining correlation information are classified into left, upper, upper left, upper right, and lower left positions, but it may also be an example applicable to a case where classification is based on various adjacent region divisions (e.g., division into left and upper positions. Here, the upper + upper right block is classified into upper *, and the left + lower left block is classified into left *).
[0289] In the latter case, the boundary values can be set to be the same or different for each region. For example, the left block can use the corresponding reference pixels in the correlation information acquisition process when all reference pixels (N pixels) are available, and the upper block can use the corresponding reference pixels in the correlation information acquisition process when at least one reference pixel is available (i.e., when some reference pixels are available).
[0290] Also, the above example may support the same or different settings depending on the color copy mode (e.g., a, b, c, etc. in FIG. 8). Related settings may be defined differently depending on other encoding / decoding settings (e.g., image type, block size, shape, position, block division type, etc.).
[0291] When the correlation information acquisition process is not performed (i.e., when even one reference pixel is not used in the correlation information acquisition process), correlation information can be acquired implicitly. For example, a and b can be set to 1 and 0, respectively, in the color copy mode formula (i.e., the data of the corresponding block can be used as the predicted value of the current block as is) or correlation information of a block that has been encoded / decoded in the color copy mode or pre-set correlation information can be used.
[0292] Alternatively, the prediction value of the current block may be filled with any value (e.g., the minimum, median, or maximum value of the bit depth or pixel value range), which is similar to the method used when all reference pixels are not available in general intra prediction.
[0293] The above example of the case where correlation information is implicitly acquired or where the predicted value is filled with an arbitrary value may be applicable to the case shown in FIG 11d, because not even one reference pixel is used in the correlation information acquisition process.
[0294] Next, a case will be described where the use of a prediction mode in which prediction is performed from reference pixels at unusable positions is restricted.
[0295] 12 is an exemplary diagram for explaining the possibility of using reference pixels in a color copy mode according to an embodiment of the present invention. In the following example, it is assumed that the left, upper, upper right, and lower left blocks of a current block and a corresponding block are used to obtain correlation information. It is also assumed that three color copy modes (mode_A, mode_B, and mode_C) are supported, and each mode is classified into a mode for obtaining correlation information in the left+upper, upper+upper right, and left+lower left blocks of each block.
[0296] 12, a of Fig. 12 shows a case where the current block is located inside an image (picture, slice, tile, etc.), b of Fig. 12 shows a case where the current block is located at the left boundary of an image, c of Fig. 12 shows a case where the current block is located at the top boundary of an image, and d of Fig. 12 shows a case where the current block is located at the top left boundary of an image. That is, it is assumed that the reference pixel availability is determined based on the position of the current block.
[0297] Referring to Fig. 12, mode_A, mode_B, and mode_C can be supported in Fig. 12a, mode_B can be supported in Fig. 12b, mode_C can be supported in Fig. 12c, and no mode can be supported in Fig. 12d. That is, if even one reference pixel used for obtaining correlation information is unavailable, the corresponding mode is not supported.
[0298] For example, when configuring an intra prediction mode candidate group for a chrominance component, the candidate group may include directional and non-directional modes, a color mode, and a color copy mode. Here, it is assumed that the candidate group includes a total of seven prediction modes, including four prediction modes such as DC, planar, vertical, horizontal, and diagonal modes, one color mode, and three color copy modes.
[0299] As in the above example, the color copy mode in which correlation information is obtained from unavailable reference pixels can be excluded. In the situation of c in Fig. 12, except for mode_c, the remaining six prediction modes can form a candidate group for the chrominance component. That is, it is possible to adjust from a group of m prediction mode candidates to n (m>n, n is an integer equal to or greater than 1). This may also require changing the prediction mode index setting and entropy encoding / decoding such as binarization.
[0300] In the situation of FIG. 12d, mode_A, mode_B, and mode_C can be excluded, and a total of four candidate groups can be configured including the remaining prediction modes, directional and non-directional modes, and color modes.
[0301] As in the above example, it is possible to configure an intra-frame prediction mode candidate group by applying a prediction mode use restriction setting based on unusable reference pixels.
[0302] As in the above example, when the reference pixel processing B is set, various processing methods can be supported. Depending on the encoding / decoding setting, the reference pixel processing and intra prediction can be performed based on one processing method either implicitly or explicitly.
[0303] The above example describes a case where the reference pixel availability of an adjacent area is determined based on the position of the current block. That is, the current block and the corresponding block are located at the same position in an image (picture, slice, tile, largest coding block, etc.), and therefore, if a certain block (current block or corresponding block) is adjacent to an image boundary, the corresponding block is also located at the image boundary. Therefore, when the reference pixel availability is determined based on the position of each block, the same result is produced.
[0304] In addition, the restricted intra prediction has been described as a criterion for determining the possibility of reference pixels, which may result in a possibility that the possibility of reference pixels in adjacent regions of a current block and a corresponding block may not be the same.
[0305] 13 is an exemplary diagram for explaining the possibility of using reference pixels in a color copy mode according to an embodiment of the present invention. In the following example, it is assumed that the left and upper blocks of the current block and the corresponding block are areas used for obtaining correlation information. That is, this may be a description of mode_A in FIG. 12, and it is assumed that the contents described below can be applied in the same or similar manner to other color copy modes.
[0306] 13, the cases can be classified into (i) when all adjacent areas of both blocks (current block and corresponding block) are usable, (ii) when only a portion of the area is usable, and (iii) when none of the areas are usable. Here, the case where only a portion of the area is used (ii) can be classified into (ii-1) when an area that can be used in common exists in both blocks and (ii-2) when it does not. Here, the case where an area that can be used in common exists (ii-1) can be classified into (ii-1-1) when the areas in both blocks completely match and (ii-1-2) when they only partially match.
[0307] Referring to Figure 13, in the above classification, i and iii correspond to a and f in Figure 13, ii-2 corresponds to d and e in Figure 13, ii-1-1 corresponds to c in Figure 13, and ii-1-2 corresponds to b in Figure 13. Here, compared to the case where the reference pixel possibility is determined based on the position of the current block (or the corresponding block), ii-2 and ii-1-2 may be new situations that must be considered.
[0308] The above process may be configured to include a step of determining the availability of reference pixels in a color copy mode. An intra-frame prediction setting of the color copy mode including processing related to the reference pixels may be determined according to the result of the determination. Next, an example of reference pixel processing and intra-frame prediction based on the availability of adjacent regions of a current block and a corresponding block will be described.
[0309] In the case of the reference pixel process A setting, the reference pixels at the unavailable positions can be filled by the various methods already described. However, the detailed setting can be made different depending on the classification according to the availability of the adjacent areas of the two blocks. The following mainly describes the cases of ii-2 and ii-1-2, and the detailed description of the other classifications will be omitted since they may overlap with the above-mentioned contents of the present invention.
[0310] For example, d(ii-2) in Fig. 13 may be a case where there is no area that can be used in common in the current block and the corresponding block. Also, b(ii-1-2) in Fig. 13 may be a case where some adjacent areas that can be used in the current block and the corresponding block overlap. In this case, the data of the adjacent area of the unavailable position can be filled (e.g., filled from the available area) and used in the correlation information acquisition process.
[0311] Alternatively, e(ii-2) in Figure 13 may be a case where either the current block or the corresponding block has an adjacent region that is unavailable, i.e., one of the two blocks has no data to be used in the correlation information acquisition process, and the corresponding region can be filled with various data.
[0312] Here, the corresponding region can be filled using any value, such as the minimum, median, or maximum value based on the pixel value range (or bit depth) of the image.
[0313] Here, the corresponding area can be filled with a neighboring area available in another color space by copying, etc. In this case, since the neighboring areas of both blocks have the same image characteristics (i.e., the same data) through the above process, it is possible to obtain pre-defined correlation information. For example, by setting a to 1 and b to 0 in the correlation related equation, it corresponds to copying the data of the corresponding block directly to the predicted value of the current block, and various other correlation information settings are possible.
[0314] In the case of the reference pixel processing B setting, the use of reference pixels at unusable positions or the use of the corresponding color copy mode can be restricted by various methods already described. However, detailed settings can be made different depending on the classification of the usability. The following mainly describes the cases of ii-2 and ii-1-2, and detailed descriptions of the other classifications will be omitted since they may overlap with the above-described contents of the present invention.
[0315] For example, d and e (ii-2) in FIG. 13 may be cases where there is no area that can be used in common in the current block and the corresponding block. Since there is no overlapping area that can be compared to obtain correlation information in both blocks, the use of the corresponding color copy mode may be restricted. Alternatively, the predicted value of the current block may be filled with an arbitrary value, etc. In other words, this may mean that the correlation information obtaining process is not performed.
[0316] Alternatively, b(ii-1-2) in Figure 13 may be a case where a portion of the adjacent usable areas of the current block and the corresponding block overlap, so the correlation information acquisition process can be performed even if it is limited to a portion of the adjacent usable areas.
[0317] 14 is a flow chart illustrating an intra-frame prediction method in a color copy mode according to an embodiment of the present invention. Referring to FIG. 14, a designated reference pixel area can be confirmed to obtain correlation information (S1400). Then, a processing setting for the reference pixels can be determined based on a usability determination of the designated reference pixel area (S1410). Then, intra-frame prediction can be performed according to the processing setting for the determined reference pixels (S1420). Here, correlation information can be obtained based on data of a reference pixel area that is usable according to the reference pixel processing setting, and a prediction block according to a color copy mode can be generated, or a prediction block filled with an arbitrary value can be generated.
[0318] In summary, when a color copy mode is supported for intra-frame prediction of chrominance components, a comparison area for obtaining correlation information designated by the color copy mode can be checked. Unlike a general intra-frame prediction mode, the color copy mode can check not only adjacent areas of the current block (particularly, areas used for correlation information comparison) but also the availability of reference pixels of the corresponding block. Depending on a preset reference pixel processing setting or one of multiple reference pixel processing settings and the availability of the reference pixels, reference pixel processing and intra-frame prediction according to the various examples described above can be performed.
[0319] Here, the reference pixel processing setting may be implicitly determined by the image type, color components, block size / position / shape, block width-to-length ratio, coding mode, intra-frame prediction mode (e.g., the range, position, number of pixels, etc. of the area to be compared for obtaining correlation information in the color copy mode), limited intra-frame prediction setting, etc., or related information may be explicitly generated in units of sequences, pictures, slices, tiles, etc. Here, the reference pixel processing setting may be defined by limiting the state information of the current block (or current image) or limiting the state information of the corresponding block (or other color image), or may be defined by combining multiple state information.
[0320] Although the reference pixel processing settings A and B have been described separately in the above example, the two settings may be used alone or in combination, which may also be determined based on the state information or explicit information.
[0321] After the reference pixels are constructed in the reference pixel interpolation unit, a small number of reference pixels may be generated by linearly interpolating the reference pixels, or the reference pixel interpolation process may be performed after a reference pixel filtering process, which will be described later.
[0322] Here, in the case of horizontal, vertical, some diagonal modes (e.g., modes with a 45 degree difference between vertical and horizontal, such as diagonal up right, diagonal down right, and diagonal down left, which correspond to modes 2, 34, and 66 in FIG. 4), non-directional mode, and color copy mode, the interpolation process is not performed, but in the case of other modes (other diagonal modes), the interpolation process can be performed.
[0323] A pixel position to be interpolated (i.e., which decimal unit is to be interpolated, determined from 1 / 2 to 1 / 64, etc.) can be determined according to a prediction mode (e.g., directionality of the prediction mode, such as dy / dx) and the positions of the reference pixel and the predicted pixel. Here, regardless of the decimal precision, one filter (e.g., assuming a filter in which the mathematical formula used to determine the length of the filter coefficients or filter taps is the same, but assuming a filter in which only the coefficients are adjusted according to the decimal precision <e.g., 1 / 32, 7 / 32, 19 / 32>) can be applied, or one of multiple filters (e.g., assuming a filter in which the mathematical formula used to determine the length of the filter coefficients or filter taps is differentiated) can be selected and applied according to the decimal precision.
[0324] In the former case, integer unit pixels may be used as inputs for the interpolation of fractional unit pixels, and in the latter case, input pixels may be different for each stage (for example, integer pixels are used in the case of 1 / 2 unit, integer and 1 / 2 unit pixels are used in the case of 1 / 4 unit, etc.), but are not limited thereto. In the present invention, the former case will be mainly described.
[0325] Fixed or adaptive filtering can be performed for reference pixel interpolation, which can be determined based on the encoding / decoding settings (e.g., one or a combination of two or more of the image type, color components, block position / size / shape, block width-to-height ratio, prediction mode, etc.).
[0326] Fixed filtering can perform reference pixel interpolation using one filter, while adaptive filtering can perform reference pixel interpolation using one of a plurality of filters.
[0327] In the case of adaptive filtering, one of the filters can be implicitly or explicitly determined according to the encoding / decoding settings. Here, the types of filters can be configured as a 4-tap DCT-IF filter, a 4-tap cubic filter, a 4-tap Gaussian filter, a 6-tap Wiener filter, an 8-tap Kalman filter, etc., and the filter candidate group supported by the color components can be defined differently (for example, some of the filter types are the same or different, or the length of the filter taps is short or long, etc.).
[0328] The reference pixel filter unit can perform filtering on the reference pixels in order to improve the accuracy of prediction by reducing degradation remaining due to the encoding / decoding process. The filter used at this time can be, but is not limited to, a low-pass filter. Whether or not filtering is applied can be determined according to the encoding / decoding setting (which can be derived from the above description). In addition, when filtering is applied, fixed filtering or adaptive filtering can be applied.
[0329] Fixed filtering means that reference pixel filtering is not performed or that reference pixel filtering is applied using one filter. Adaptive filtering means that whether or not filtering is applied is determined by the encoding / decoding settings, and if there are two or more supported types of filters, one of them can be selected.
[0330] Here, the filter type may support a plurality of filters classified according to various filter coefficients, filter tap lengths, etc., such as a 3-tap filter such as [1, 2, 1] / 4 and a 5-tap filter such as [2, 3, 6, 3, 2] / 16.
[0331] The reference pixel interpolation unit and the reference pixel filter unit introduced in the reference pixel configuration step may be necessary components for improving prediction accuracy. The above two processes may be performed independently, but a configuration in which the two processes are combined (i.e., processed in one filtering) is also possible.
[0332] The prediction block generator may generate a prediction block according to at least one prediction mode and may use reference pixels according to the prediction mode, where the reference pixels may be used in a method such as extrapolation (directional mode) or in a method such as interpolation, DC, or copy (non-directional mode) according to the prediction mode.
[0333] Next, reference pixels used depending on the prediction mode will be described.
[0334] In the case of directional modes, modes between horizontal and some diagonal modes (modes 2 to 17 in Figure 4) can use the reference pixels in the lower left + left block (Ref_BL, Ref_L in Figure 10), horizontal modes can use the reference pixels in the left block, modes between horizontal and vertical (modes 19 to 49 in Figure 4) can use the reference pixels in the left + upper left + upper blocks (Ref_L, Ref_TL, Ref_T in Figure 10), vertical modes can use the reference pixels in the upper block (Ref_L in Figure 10), and modes between vertical and some diagonal modes (diagonal down left) (modes 51 to 66 in Figure 4) can use the reference pixels in the upper + upper right block (Ref_T, Ref_TR in Figure 10).
[0335] In addition, in the case of a non-directional mode, reference pixels located in one or more of the lower left, left, upper left, upper, and upper right blocks (Ref_BL, Ref_L, Ref_TL, Ref_T, and Ref_TR in FIG. 10) can be used. For example, various combinations of reference pixels such as the left side, upper side, left side + upper side, left side + upper side + upper left side, left side + upper side + upper left side + upper right side + lower left side can be used for intra-screen prediction, which can be determined according to the non-directional mode (DC, Planar, etc.). In the example described below, it is assumed that the DC mode uses the left side + upper block, and the planar mode uses the left side + upper + lower left + upper right block as reference pixels for prediction.
[0336] In addition, in the case of the color copy mode, a restored block in another color space (Ref_C in FIG. 10) can be used as a reference pixel. In the example described below, a case where the current block and the corresponding block are used as reference pixels for prediction will be mainly described.
[0337] Here, the reference pixels used in the intra prediction may be divided into a plurality of concepts (or units). For example, the reference pixels used in the intra prediction may be divided into one or more categories, such as a first reference pixel and a second reference pixel. For convenience of explanation, the reference pixels are divided into a first reference pixel and a second reference pixel, but it may be understood that other additional reference pixels are also supported.
[0338] Here, the first reference pixel may be a pixel directly used in generating a predicted value of the current block, and the second reference pixel may be a pixel indirectly used in generating a predicted value of the current block. Alternatively, the first reference pixel may be a pixel used in generating a predicted value of all pixels in the current block, and the second reference pixel may be a pixel used in generating a predicted value of some pixels in the current block. Alternatively, the first reference pixel may be a pixel used in generating a first predicted value of the current block, and the second reference pixel may be a pixel used in generating a second predicted value of the current block. Alternatively, the first reference pixel may be a pixel located based on the starting point (or origin) of the prediction direction of the current block, and the second reference pixel may be a pixel located regardless of the prediction direction of the current block.
[0339] As described above, the prediction using the second reference pixels can be called a prediction block (or correction) process. That is, the prediction block generator of the present invention can be configured to add a prediction block correction unit.
[0340] Here, the reference pixels used in the prediction block generator and the prediction block corrector are not limited to the first and second reference pixels in each configuration. That is, the prediction block generator can perform prediction using the first reference pixels or the first and second reference pixels, and the prediction block corrector can perform prediction (or correction) using the second reference pixels or the first and second reference pixels. It should be understood that the present invention will be described by dividing it into a plurality of pixel concepts for convenience of description.
[0341] In the intra prediction of the present invention, prediction can be performed using the first reference pixel as well as the second reference pixel (i.e., correction can be performed), which can be determined by the encoding / decoding setting. First, information on whether the second reference pixel is supported (i.e., whether prediction value correction is supported) can be generated in units of a sequence, a picture, a slice, a tile, etc. Even if it is explicitly or implicitly determined that the second reference pixel is supported, whether it is supported in all blocks or in some blocks, and detailed settings related to the second reference pixel in the supported block (related contents can be referred to in the contents to be described later) can be determined by the encoding / decoding setting defined by the image type, color components, block size / shape / position, block width-to-length ratio, encoding mode, intra prediction mode, restricted intra prediction setting, etc. Alternatively, related setting information can be explicitly determined in units of a sequence, a picture, a slice, a tile, etc.
[0342] Here, the numbers of first and second reference pixels used for the predicted value of one pixel may be m and n, respectively, and m and n may have a pixel number ratio of (1:1), (1:2 or more), (2 or more:1), or (2 or more:2 or more), which may be determined according to the prediction mode, the size / shape / position of the current block, pixel position, etc. That is, m may be an integer of 1 or more, such as 1, 2, 3, etc., and n may be an integer of 1 or more, such as 1, 2, 3, 4, 5, 8, etc.
[0343] When the weights applied to the first and second reference pixels (assuming one pixel each is used in this example) are p and q, p can be greater than or equal to q, p can have a positive value, and q can have a positive or negative value.
[0344] Next, a case will be described in which, in addition to generating a prediction block using the first reference pixels, the second reference pixels are also used to perform prediction.
[0345] For example, (see Figure 10) the Diagonal up right direction mode is<Ref_BL+Ref_L> Block or<Ref_BL+Ref_L+Ref_T+Ref_TR> The prediction can be performed using the Ref_L block or<Ref_L+Ref_T+Ref_TL> The prediction can be performed using blocks. Also, the diagonal down right direction mode is<Ref_TL+Ref_T+Ref_L> Block or<Ref_TL+Ref_T+Ref_L+Ref_TR+Ref_BL> The prediction can be performed using the Ref_T block or<Ref_T+Ref_L+Ref_TL> You can use blocks to perform prediction. Also, the diagonal down left direction mode is<Ref_TR+Ref_T> Block or<Ref_TR+Ref_T+Ref_L+Ref_BL> Blocks can be used to perform predictions.
[0346] As another example, DC mode is<Ref_T+Ref_L> Block or<Ref_T+Ref_L+Ref_TL+Ref_TR+Ref_BL> Prediction can be performed using blocks. Planar mode is also<Ref_T+Ref_L+Ref_TR+Ref_BL> Block or<Ref_T+Ref_L+Ref_TR+Ref_BL+Ref_TL> Blocks can be used to perform predictions.
[0347] Also, the color copy mode is Ref_C block or<Ref_C+(Ref_T or Ref_L or Ref_TL or Ref_TR or Ref_BL)> You can use blocks to perform predictions, or<Ref_C+(Def_T or Def_B or Def_L or Def_R or Def_TL or Def_TR or Def_BL or Def_BR)> Here, Def is a term used to indicate adjacent blocks of Ref_C (current block and corresponding block) not shown in FIG. 10, and Def_T to Def_BR can be adjacent blocks in the upper, lower, left, right, upper left, upper right, lower left, and lower right directions of Ref_F. That is, the color copy mode can perform prediction using reference pixels adjacent to the current block or reference pixels adjacent to Ref_C (corresponding block) for the Ref_C block.
[0348] The above examples show some examples in which prediction is performed using the first reference pixels or the first and second reference pixels, but the present invention is not limited thereto and various modifications are possible.
[0349] The use of multiple reference pixels to generate or correct a prediction block may be performed in order to compensate for defects in an existing prediction mode.
[0350] For example, in the case of the directional mode, it may be a mode supported for the purpose of increasing the accuracy of prediction by assuming the existence of an edge in a specific direction of the current block, but the accuracy of prediction may be reduced because the change in the block may not be accurately reflected only by the reference pixel located at the start point of the prediction direction. Alternatively, in the case of the color copy mode, prediction is performed by reflecting correlation information from another color image at the same time, but the accuracy of prediction may be reduced because the deterioration remaining at the block boundary in the other color image may be reflected. To solve the above problem, the accuracy of prediction may be increased by additionally using a second reference pixel.
[0351] Next, a case where prediction is performed using a plurality of reference pixels in a color copy mode will be described. Parts not shown in the drawings of the following example can be guided with reference to FIG. 10. Here, the following description will focus on the conceptual part, assuming that the correlation information acquisition process of the color copy mode can be performed in a previous process or a subsequent process. In addition, except for obtaining a predicted value from another color space, the intra prediction in the color copy mode described below can be applied in the same or similar manner to other intra prediction modes.
[0352] 15 is an example diagram for explaining prediction in a color copy mode according to an embodiment of the present invention, in which a predicted block can be generated using reference pixels of a corresponding block in the color copy mode.
[0353] In detail, correlation information can be obtained from adjacent regions (specified by the color copy mode) of the current block and the corresponding block (p1). Then, data is obtained from the corresponding block (p2), and the previously obtained correlation information is applied (p3) to obtain a predicted block (pred_t), which can be compensated for by the predicted block (pred_f) of the current block (p4).
[0354] Since the above example can be derived from the color copy mode described above, detailed description will be omitted.
[0355] 16 is an example diagram for explaining prediction in a color copy mode according to an embodiment of the present invention. Referring to FIG 16, in the color copy mode, a predicted block can be generated and corrected using a corresponding block and its adjacent reference pixels.
[0356] In detail, correlation information can be obtained from adjacent regions of the current block and the corresponding block (p1). Then, correction can be performed on the data of the corresponding block. Here, the correction can be limited to the inside of the corresponding block (d5), or limited to the boundary of the block adjacent to the corresponding block (d1 to d9 excluding d5), or performed across the boundary between the inside and outside of the corresponding block (d1 to d9). That is, data of the corresponding block and the adjacent region of the corresponding block can be used for correction. Here, the outer boundary to which the correction is performed can be one or more (including all) of the upper, lower, left, right, upper left, upper right, lower left, and lower right directions (d1 to d9 respectively, excluding d5).
[0357] Then, data is obtained through the correction process of the corresponding block (p2), and the previously obtained correlation information is applied (p3) to obtain a predicted block (pred_t), which can be compensated for by the predicted block (pred_f) of the current block (p4).
[0358] 17 is an example diagram for explaining prediction in a color copy mode according to an embodiment of the present invention. Referring to FIG 17, in the color copy mode, a predicted block can be generated and corrected using adjacent reference pixels of a corresponding block and a current block.
[0359] In detail, correlation information can be obtained from adjacent areas of the current block and the corresponding block (p1). Then, data can be obtained from the corresponding block (p2), and a predicted block (pred_t) can be obtained by applying the previously obtained correlation information (p3). This can be compensated for by a first predicted block of the current block (p4), and correction can be performed on the first predicted block (referred to as a predicted block). Here, the correction can be limited to the inside of the predicted block (c5), or limited to the boundary between the predicted block and an adjacent block (c1 to c6 except for c5), or performed across the inside and outside boundaries of the predicted block (c1 to c6). That is, the predicted block (i.e., data based on the corresponding block; here, the expression "based on" means that correlation information has been applied) and adjacent data of the predicted block (data of an adjacent area of the current block) can be used for correction. Here, the outer boundary on which the correction is performed may be one or more (including all) of the upper, left, upper left, upper right, and lower left directions (c1 to c6, respectively, excluding c5).
[0360] The data obtained by the correction process of the predicted block can be compensated for with the (secondary or final) predicted block (pred_f) of the current block (p5).
[0361] 18 is a flow chart of a process for performing correction in a color copy mode according to an embodiment of the present invention, in which one of the correction processes described with reference to FIGS. 16 and 17 is selected and performed.
[0362] 18, correlation information can be obtained from adjacent regions of a current block and a corresponding block (S1800). Then, data of the corresponding block can be obtained (S1800). In the following description, it is assumed that it is implicitly or explicitly determined that compensation of the predicted block is to be performed. It can be determined whether compensation of the predicted block is to be performed in the current color space or in another color space (S1820).
[0363] If the predicted block is to be corrected in the current color space, correlation information can be applied to the data of the corresponding block to generate a predicted block (S1830). This can be the same as the process of a general color copy mode. Then, the predicted block can be corrected using adjacent areas of the current block (S1840). Here, not only the adjacent areas of the current block but also the obtained internal data of the predicted block can be used.
[0364] If the prediction block is to be compensated in another color space, the compensation of the corresponding block can be performed using adjacent regions of the corresponding block (S1850). Here, not only the adjacent regions of the corresponding block but also the internal data of the corresponding block can be used. Then, the prediction block can be generated by applying correlation information to the compensated corresponding block (S1860).
[0365] The data obtained by the above process can be compensated for by the predicted block of the current block (S1870).
[0366] The classification according to color space in the above example does not mean that the application examples of Figures 16 and 17 cannot be used together. That is, the application examples of Figures 16 and 17 can be combined. For example, as shown in Figure 16, a predicted block of a current block can be obtained by a correction process of a corresponding block, and as shown in Figure 17, a final predicted block can be obtained by a correction process of the obtained predicted block.
[0367] In the intra prediction of the present invention, performing correction may mean applying filtering to a pixel to be corrected and other pixels (adjacent pixels in this example). Here, filtering may be performed according to one pre-set filtering setting, or one of a plurality of filtering settings may be selected and the corresponding filtering setting may be performed. Here, the filtering setting may be included in the above-mentioned second reference pixel-related detailed setting.
[0368] The filtering setting may include whether filtering is applicable, the type of filter, the coefficient of the filter, the pixel position used in the filter, etc. Here, the unit to which the filtering setting is applied may be a block or a pixel unit. Here, the filtering setting may be related to the filtering setting by defining the encoding / decoding setting according to the image type, color component, color format (i.e., the composition ratio between color components), coding mode (Intra / Inter), block size / shape / position, block width-to-length ratio, pixel position in the block, intra-picture prediction mode, constrained intra-picture prediction, etc. Here, the block is described assuming the current block, but it may also be understood as a concept including adjacent blocks of the current block or adjacent blocks of the corresponding block in the color copy mode. That is, it means that state information of the current block and other blocks can act as input variables for the filtering setting. In addition, information on the filtering setting may be explicitly included in units of a sequence, a picture, a slice, a tile, a block, etc.
[0369] Next, the types of filters in the filtering settings will be described. Filtering can be applied to adjacent pixels on one line, such as a horizontal line, a vertical line, or a diagonal line, with the pixel to be corrected as the center (i.e., one-dimensional). Alternatively, filtering can be applied to spatially adjacent pixels in the left, right, upper, lower, upper left, upper right, lower left, and lower right directions with the pixel to be corrected as the center (i.e., two-dimensional). That is, filtering can be applied to adjacent pixels within M×N with the pixel to be corrected as the center. In the example described below, it is assumed that both M and N are 3 or less, but M or N can have a value greater than 3. In general, filtering can be applied to adjacent pixels symmetrically with respect to the pixel to be corrected as the center, but an asymmetric configuration is also possible.
[0370] 19 is an example diagram for explaining types of filters applied to a pixel to be corrected according to an embodiment of the present invention, in particular, a case where filter application positions are determined symmetrically around a pixel to be corrected (thick line in the figure).
[0371] Referring to Fig. 19, Fig. 19a and Fig. 19b refer to horizontal and vertical 3-tap filters, Fig. 19c and Fig. 19d refer to diagonal 3-tap filters (angles inclined at -45 and +45 degrees from the vertical), Fig. 19e and Fig. 19f refer to 5-tap filters having (+) or (x) shapes, and Fig. 19g refers to a square 9-tap filter.
[0372] As an example for applying filtering, a-g in FIG. 19 can be applied to the inside or outside (or boundary) of a block.
[0373] As another example, e to g in Fig. 19 can be applied to the inside of a block. Or a to d in Fig. 19 can be applied to the boundaries of a block. In particular, a in Fig. 19 can be applied to the left boundary of the block, and b in Fig. 19 can be applied to the top boundary of the block. c in Fig. 19 can be applied to the top left boundary of the block, and d in Fig. 19 can be applied to the top right and bottom left boundaries of the block.
[0374] The above examples are just some of the cases for filter selection based on the position of the pixel to be corrected, and are not limited thereto, and various application examples including the opposite case are possible.
[0375] 20 is an exemplary diagram illustrating a type of filter applied to a pixel to be corrected according to an embodiment of the present invention, in particular, a case where a filter application position is determined asymmetrically around a pixel to be corrected.
[0376] Referring to Fig. 20, a and b in Fig. 20 respectively refer to a 2-tap filter using only the left and upper pixels, and c in Fig. 20 refers to a 3-tap filter using the left and upper pixels. d to f in Fig. 20 respectively refer to 4-tap filters in the upper left direction, upper direction, and left direction. G to j in Fig. 20 respectively refer to 6-tap filters in the upper left and right directions, upper left and upper left directions, lower left and lower right directions, and lower right and lower right directions.
[0377] As an example for applying filtering, a to j in FIG. 20 can be applied inside or outside the block.
[0378] As another example, g to j in Fig. 20 can be applied to the inside of a block. Or a to f in Fig. 20 can be applied to the boundaries of a block. In particular, a and f in Fig. 20 can be applied to the left boundary of a block, and b and e in Fig. 20 can be applied to the top boundary of a block. C and d in Fig. 20 can be applied to the top left boundary of a block.
[0379] The above examples are just some of the cases for filter selection based on the position of the pixel to be corrected, and are not limited thereto, and various application examples including the opposite case are possible.
[0380] 19 and 20 can be set in various ways. For example, in the case of a 2-tap filter, weights are set in a ratio of 1:1 or 1:3 (where 1 and 3 are the weights of the pixel to be corrected), in the case of a 3-tap filter, weights are set in a ratio of 1:1:2 (where 2 is the weight of the pixel to be corrected), in the case of a 4-tap filter, weights are set in a ratio of 1:1:1:5 or 1:1:2:4 (where 4 and 5 are the weights of the pixel to be corrected), and in the case of a 5-tap filter, weights are set in a ratio of 1: Weighting values in a ratio of 1:1:1:4 (where 4 is the weighting value of the pixel to be corrected), weighting values in a ratio of 1:1:1:1:2:2 in the case of a 6-tap filter (where 2 is the weighting value of the pixel to be corrected), and weighting values in a ratio of 1:1:1:1:1:1:1:1:8 or 1:1:1:1:1:2:2:2:4 (where 8 and 4 are the weighting values of the pixel to be corrected) in the case of a 9-tap filter may be applied. In the above examples, the pixel to which the weighting value next to that of the pixel to be corrected is applied may be a pixel that is close to the pixel to be corrected (vertically and horizontally adjacent) or a pixel that is located at the center like the pixel to be corrected in a symmetrical structure. The above examples are only some examples of weighting value setting, and are not limited thereto, and various modifications are possible.
[0381] The filter used for compensation may be implicitly determined by the encoding / decoding setting, or may be explicitly included in units of a sequence, a picture, a slice, a tile, etc. Here, an explanation of defining the encoding / decoding setting may be derived from the various examples of the present invention described above.
[0382] Here, information indicating whether multiple filters are supported may be generated in the unit. If multiple filters are not supported, a pre-defined filter may be used, and if multiple filters are supported, filter selection information may be additionally generated. Here, the filters may include the filters shown in Figures 19 and 20 or other filters in the candidate group.
[0383] Next, assume that compensation is performed in another color space as shown in Figure 16, and refer to Figure 10 for this purpose (i.e., the block to be compensated is Ref_C). In the example described below, assume that the filter to be used in the corresponding block is determined by the previous process, and the corresponding filter is applied to all or most of the pixels in the corresponding block.
[0384] Although a collective filter is supported based on state information of adjacent regions of the correction target block, an adaptive filter based on the position of the correction target pixel can also be supported.
[0385] For example, a 5-tap filter (e in FIG. 19) may be applied to aa to pp of Ref_C centered on the pixel to be corrected. In this case, the same filter may be applied regardless of the position of the pixel to be corrected. In addition, there may be no restrictions on applying filtering to adjacent areas of the block to be corrected.
[0386] Alternatively, a 5-tap filter may be applied to ff, gg, jj, and kk of Ref_C (i.e., inside the block) centered on the pixel to be corrected, and filtering may be applied to other pixels (block boundaries) based on the state of adjacent blocks. In particular, if the left block of Ref_C is the outer edge of the picture or the coding mode is Inter (i.e., when the restricted intra-frame prediction setting is activated) and it is determined that the usability of the corresponding pixel is not possible (related explanation can be derived from the part of the present invention that checks the reference pixel usability), a vertical filter (e.g., b in FIG. 19, etc.) may be applied to aa, ee, ii, and mm of Ref_C. Alternatively, if the upper block of Ref_C is not usable, a horizontal filter (e.g., a in FIG. 19, etc.) may be applied to aa to dd of Ref_C.
[0387] Although a collective filter is supported based on an intra-frame prediction mode, an adaptive filter based on a pixel position to be corrected can also be supported.
[0388] For example, in the case of a mode in which correlation information is obtained from the left and upper blocks in the color copy mode, a 9-tap filter (g in FIG. 19) can be applied to aa to pp of Ref_C. In this case, it can be an example in which the same filter is applied regardless of the position of the correction target pixel.
[0389] Or, in the case of a mode in which correlation information is obtained from the left and lower left blocks in the color copy mode, a 9-tap filter can be applied to ee to pp of Ref_C, and other pixels (upper boundary) can be filtered based on the prediction mode setting. In this example, since correlation information is obtained from the left and lower left blocks, it can be estimated that the correlation with the upper block is low. Therefore, a filter in the left and lower left direction (for example, i in FIG. 20) can be applied to aa to dd of Ref_C.
[0390] In the present invention, the color copy mode has been described mainly in the case of a certain color format 4:4:4. There may be differences in the details of the correction depending on the color format. This example assumes that the correction is performed in another color space.
[0391] In color copy mode, the correlation information and prediction value can be obtained by directly obtaining the relevant pixel from a pixel in the current color space, for example, in the case of a 4:4:4 color format, where a pixel in the current color space corresponds to a pixel in another color space.
[0392] Meanwhile, in the case of some color formats (4:2:0), one pixel in the current color space corresponds to one or more pixels (four in this example) in another color space. If a single predefined pixel is not selected and associated data is not obtained from the pixel, a downsampling process may be required to obtain associated data from multiple corresponding pixels.
[0393] Next, the prediction and correction process of the color copy mode according to each color format will be explained. In the example below, it is assumed that a downsampling process is performed for some formats (4:2:0). Also, the current block and the corresponding block are called blocks A and B, respectively.
[0394] <1> In-screen prediction in 4:4:4 format <1-1> Obtain pixel values from adjacent areas of block A <1-2> Obtain pixel values corresponding to the pixels in <1-1> from the adjacent area of block B <1-3> Acquire correlation information based on pixel values of adjacent areas in each color space <1-4> Extraction of pixels in block B and adjacent pixels Apply filtering to pixels <1-5> and <1-4> to correct B block <1-6> Obtain pixel values of block B (M×N) corresponding to pixels of block A (M×N) <1-7> Apply correlation information to pixel values of <1-6> to generate predicted pixels
[0395] <2> In-frame prediction in 4:2:0 format (1) <2-1> Obtain pixel values from adjacent areas of block A <2-2> Extract pixels corresponding to the pixels in <2-1> and adjacent pixels from the adjacent area of block B <2-3> Apply downsampling to the pixels in <2-2> to obtain pixel values corresponding to the pixels in <2-1> from the adjacent area of block B. <2-4> Acquire correlation information based on pixel values of adjacent regions in each color space <2-5> Extraction of pixels in block B and adjacent pixels Apply filtering to pixels <2-6> and <2-5> to correct B block <2-7> Extraction of pixels in block B (2M x 2N) corresponding to pixels in block A (M x N) and adjacent pixels Apply downsampling to the pixels in <2-8> and <2-7> to obtain the pixel values of block B. Apply correlation information to the pixel values in <2-9> and <2-8> to generate predicted pixels.
[0396] <1> and <2> To explain the process, <1> applies filtering once in <1-5>, whereas <2> It can be seen that multiple filtering is applied in <2-6> and <2-8>. <2-6> can be a process for correcting data used to obtain predicted pixels, and <2-8> can be a downsampling process for obtaining predicted pixels, and the filters in each process can be configured differently. Of course, the encoding performance can be improved depending on each process, but overlapping filtering effects can occur. In addition, the increased number of filtering times can increase the complexity, which may be a configuration that is not suitable for some profiles. For this reason, in the case of some color formats, it may be necessary to support filtering that integrates this.
[0397] <3> In-frame prediction in 4:2:0 format (2) <3-1> Obtain pixel values from adjacent areas of block A <3-2> Extract pixels corresponding to the pixels in <3-1> and adjacent pixels from the adjacent area of block B Apply downsampling to the pixels in <3-3> and <3-2> to obtain pixel values corresponding to the pixels in <3-1> in the adjacent area of block B. <3-4> Acquire correlation information based on pixel values of adjacent regions in each color space <3-5> Extraction of pixels in block B and adjacent pixels Apply filtering to the pixels in <3-6> and <3-5> to correct the B block. <3-7> Obtain pixel values of block B (2M x 2N) corresponding to pixels of block A (M x N) <3-8> <3-7> Apply correlation information to pixel values to generate predicted pixels
[0398] <4> In-frame prediction in 4:2:0 format (3) <4-1> Obtain pixel values from adjacent areas of block A <4-2> Extract pixels corresponding to the pixels in <4-1> and adjacent pixels from the adjacent area of block B <4-3> Apply downsampling to the pixels in <4-2> to obtain pixel values corresponding to the pixels in <4-1> from the adjacent area of block B. <4-4> Acquire correlation information based on pixel values of adjacent regions in each color space <4-5> Extract pixels in block B (2M x 2N) corresponding to pixels in block A (M x N) and adjacent pixels <4-6> Apply downsampling to the pixels in <4-5> to obtain pixel values for block B. <4-7> Apply correlation information to the pixel values in <4-6> to generate predicted pixels.
[0399] <3> To explain the process, the downsampling process of block B is omitted and correlation information is applied to one pixel at a preset position. Instead, a correction process is performed, which has the effect of eliminating defects in the downsampling process.
[0400] on the other hand, <4> The above process is a case where the compensation process is omitted and the downsampling process of block B is performed. Instead, a filter capable of achieving the compensation effect during the downsampling process can be used. Although modifications may occur due to the configuration in which the compensation process is included in the downsampling, the settings during the compensation process described above can be applied as is. That is, filter selection information for downsampling can be explicitly determined in a higher order or implicitly determined by encoding / decoding settings. In addition, it can be implicitly or explicitly determined whether downsampling is performed using only internal data of block B or whether downsampling is performed using external data in one or more directions, such as block B and the left side, right side, upper side, lower side, upper left side, upper right side, lower left side, and lower right side.
[0401] The prediction mode decision unit performs a process for selecting an optimal mode from a group of prediction mode candidates. In general, the optimal mode can be determined in terms of coding cost using a rate-distortion technique that considers block distortion (e.g., distortion between a current block and a reconstructed block, SAD (Sum of Absolute Difference), SSD (Sum of Square Difference, etc.)) and the amount of bits generated by the corresponding mode. A prediction block generated based on the prediction mode determined by the above process can be transmitted to a subtraction unit and an addition unit.
[0402] In order to determine the optimal prediction mode, all prediction modes present in the prediction mode candidate group may be searched, or the optimal prediction mode may be selected through another determination process for the purpose of reducing the amount of calculation / complexity. For example, in the first step, some modes exhibiting good performance in terms of image quality degradation may be selected from all intra-frame prediction mode candidates, and in the second step, the optimal prediction mode may be selected from the modes selected in the first step, taking into consideration not only image quality degradation but also the amount of generated bits. In addition to the above methods, various methods may be applied in terms of reducing the amount of calculation / complexity.
[0403] In addition, the prediction mode determination unit may be a configuration that is generally included only in the encoder, but may also be a configuration that is included in the decoder depending on the encoding / decoding setting. For example, when the prediction method includes template matching or includes a method of inducing an intra-frame prediction mode in a neighboring region of a current block, in the latter case, it can be understood that a method of implicitly acquiring a prediction mode in a decoder is used.
[0404] The prediction mode encoding unit may encode the prediction mode selected by the prediction mode determination unit. It may encode index information corresponding to the prediction mode from a group of prediction mode candidates, or it may predict the prediction mode and encode information about the prediction mode. In the former case, the method may be applied to the luminance component, and in the latter case, the method may be applied to the chrominance component, but is not limited thereto.
[0405] When predicting and encoding a prediction mode, the predicted value (or prediction information) of the prediction mode may be referred to as an MPM (Most Probable Mode). The MPM may be configured with one prediction mode or multiple prediction modes, and the number of MPMs (k, where k is an integer equal to or greater than 1, such as 1, 2, 3, 6, etc.) may be determined according to the number of prediction mode candidate groups. When the MPM is configured with multiple prediction modes, it may be referred to as an MPM candidate group.
[0406] MPM is a concept that supports efficient coding of prediction modes, and can actually configure a group of candidates for prediction modes that are highly likely to occur in the prediction mode of a current block.
[0407] For example, the MPM candidate group may be configured with preset prediction modes (or statistically frequently occurring prediction modes, such as DC, Plainr, vertical, horizontal, and some diagonal modes), prediction modes of adjacent blocks (left, upper, upper left, upper right, lower left blocks, etc.), etc. Here, the prediction modes of adjacent blocks may be obtained from L0 to L3 (left blocks), T0 to T3 (upper blocks), TL (upper left blocks), R0 to R3 (upper right blocks), and B0 to B3 (lower left blocks) in FIG.
[0408] When an MPM candidate group can be configured from two or more subblock positions (e.g., L0, L2, etc.) in an adjacent block (e.g., a left block), the prediction mode of the corresponding block can be configured as a candidate group according to a predefined priority order (e.g., L0-L1-L2, etc.). Alternatively, when an MPM candidate group cannot be configured from two or more subblock positions, the prediction mode of a subblock corresponding to a predefined position (e.g., L0, etc.) can be configured as a candidate group. In detail, the prediction modes of the L3, T3, TL, R0, and B0 positions in the adjacent block can be selected as the prediction modes of the corresponding adjacent block and included in the MPM candidate group. The above description is a part of the case where the prediction modes of the adjacent blocks are configured as a candidate group, and is not limited thereto. In the following example, a case where prediction modes of predefined positions are configured as a candidate group is assumed.
[0409] In addition, when one or more prediction modes are configured in the MPM candidate group, modes derived from one or more prediction modes already included can also be configured as additional MPM candidates. In particular, when the kth mode (directional mode) is included in the MPM candidate group, modes derivable from the corresponding mode (modes having intervals of +a, -b based on k, where a and b are integers equal to or greater than 1, such as 1, 2, and 3) can be additionally included in the MPM candidate group.
[0410] There may be a priority for constructing MPM candidates, and the MPM candidates may be constructed in the order of prediction mode of adjacent blocks - preset prediction mode - induced prediction mode, etc. The process of constructing MPM candidates may be completed when the maximum number of MPM candidates is filled according to the priority. In the process, if a prediction mode matches a prediction mode already included, the corresponding prediction mode is not configured in the candidate group, and a redundancy confirmation process may be included in which the procedure moves to a candidate with the next priority.
[0411] Next, it is assumed that the MPM candidate set consists of six prediction modes.
[0412] For example, the candidate group may be configured in the order of LT-TL-TR-BL-Planar-DC-Vertical-Horizontal-Diagonal mode, etc. In this case, the prediction modes of adjacent blocks are preferentially configured in the candidate group, and an already-set prediction mode may be additionally configured.
[0413] Or LT-Planar-DC-<L+1> - <l-1> -<T+1>- <t-1>- A candidate group may be configured in the order of vertical-horizontal-diagonal mode, etc. A prediction mode of some adjacent blocks and a part of the preset prediction modes may be configured preferentially, and an induced mode and a part of the preset prediction modes may be additionally configured under the assumption that a prediction mode in a similar direction to the prediction mode of the adjacent block occurs.
[0414] The above examples are only a few of the cases regarding the configuration of MPM candidate groups, and various modifications are possible without being limited thereto.
[0415] The MPM candidate set can use binarization such as unary binarization and truncated rice binarization based on the index in the candidate set. That is, the mode bits can be represented by assigning short bits to candidates with small indices and long bits to candidates with large indices.
[0416] Modes that cannot be included in the MPM candidate group can be classified into the non-MPM candidate group. In addition, the non-MPM candidate group can be classified into two or more candidate groups depending on the encoding / decoding settings.
[0417] Next, it is assumed that there are 67 modes including directional modes and non-directional modes in the prediction mode candidate group, and 6 MPM candidates are supported, resulting in a non-MPM candidate group consisting of 61 prediction modes.
[0418] When the non-MPM candidate group is composed of one, since the prediction modes that could not be included in the MPM candidate group composition process remain, an additional candidate group composition process is not required. Therefore, binarization such as fixed length binarization and truncated unary binarization can be used based on the index in the non-MPM candidate group.
[0419] Assuming that the non-MPM candidate group is composed of two or more candidate groups, in this example, the non-MPM candidate group is classified into non-MPM_A (hereinafter, candidate group A) and non-MPM_B (hereinafter, candidate group B). It is assumed that candidate group A (p candidates, equal to or greater than the number of MPM candidate groups) constitutes the candidate group based on prediction modes that are more likely to occur in the prediction mode of the current block than candidate group B (q candidates, equal to or greater than the number of A candidates). Here, a process of constructing candidate group A can be added.
[0420] For example, some prediction modes with equal intervals (e.g., modes 2, 4, and 6) among the directional modes can be configured as candidate group A, or pre-defined prediction modes (e.g., modes derived from prediction modes included in the MPM candidate group) can be configured. The remaining prediction modes from the MPM candidate group configuration and candidate group A configuration can be configured as candidate group B, and no additional candidate group configuration process is required. Binary coding, such as fixed-length binarization and truncated unary binarization, can be used based on indexes within candidate group A and candidate group B.
[0421] The above examples are just some of the cases where the non-MPM candidate group is composed of two or more non-MPM candidates, but the present invention is not limited to these examples and various modifications are possible.
[0422] Next, a process for predicting and encoding a prediction mode will be described.
[0423] Information (mpm_flag) on whether the prediction mode of the current block matches the MPM (or some mode in the MPM candidate group) can be checked.
[0424] If it matches the MPM, the MPM index information (mpm_idx) can be additionally checked according to the MPM configuration (1 or more), and then the encoding process of the current block is completed.
[0425] If the MPM does not match, and the non-MPM candidate group consists of one, the non-MPM index information (remaining_idx) can be checked, and then the encoding process of the current block is completed.
[0426] If there are multiple non-MPM candidate groups (two in this example), information (non_mpm_flag) can be checked to see whether the prediction mode of the current block matches some of the prediction modes in candidate group A.
[0427] If it matches candidate group A, it can check candidate group A index information (non_mpm_A_idx), and if it does not match candidate group A, it can check candidate group B index information (remaining_idx). Then, the encoding process of the current block is completed.
[0428] When the prediction mode candidate group configuration is fixed, a prediction mode supported in a current block, a prediction mode supported in an adjacent block, and a previously set prediction mode may use the same prediction number index.
[0429] Meanwhile, when the prediction mode candidate group configuration is adaptive, the prediction mode supported by the current block, the prediction mode supported by the neighboring block, and the preset prediction mode may use the same prediction number index or different prediction number indexes. For the following description, refer to FIG. 4.
[0430] In the prediction mode encoding process, a prediction mode candidate group unification (or adjustment) process for configuring MPM candidates etc. may be performed. For example, the prediction mode of the current block may be one of the prediction mode candidate group of modes -5 to 61, and the prediction mode of the adjacent block may be one of the prediction mode candidate group of modes 2 to 66. In this case, a part of the prediction mode of the adjacent block (mode 66) may be a mode that is not supported by the prediction mode of the current block, so a process of unifying them may be performed in the prediction mode encoding process. That is, this process may not be required when supporting a fixed intra-frame prediction mode candidate group configuration, and may be required when supporting an adaptive intra-frame prediction mode candidate group configuration, and a detailed description thereof will be omitted.
[0431] Unlike the method using the MPM, the coding method can perform coding by assigning an index to a prediction mode belonging to a prediction mode candidate group.
[0432] For example, a method of assigning an index to a prediction mode according to a predefined priority and encoding the corresponding index when a prediction mode of a current block is selected corresponds to this method, which means that a set of prediction mode candidates is fixedly configured and a fixed index is assigned to a prediction mode.
[0433] Alternatively, when a prediction mode candidate group is adaptively configured, the fixed index allocation method may not be suitable. For this reason, a method of assigning an index to a prediction mode according to an adaptive priority and encoding the corresponding index when a prediction mode of a current block is selected can be applied. This allows the prediction mode to be efficiently encoded by varying the index assigned to the prediction mode according to the adaptive configuration of the prediction mode candidate group. That is, the adaptive priority is for assigning a candidate that is likely to be selected as a prediction mode of a current block to an index that generates a short mode bit.
[0434] Next, it is assumed that eight prediction modes including the preset prediction modes (directional mode and non-directional mode), the color copy mode, and the color mode (chrominance component) are supported in the prediction mode candidate group.
[0435] For example, assume that the preset prediction modes support four of planar, DC, horizontal, vertical, and diagonal modes (Diagonal down left in this example), one color mode (C), and three color copy modes (CP1, CP2, CP3). The basic order of indexes assigned to prediction modes can be given as preset prediction mode-color copy mode-color mode, etc.
[0436] Here, the preset prediction modes, directional mode, non-directional mode, and color copy mode, can be easily classified into prediction modes with different prediction methods. However, in the case of a color mode, it may be a directional mode or a non-directional mode, which may overlap with a preset prediction mode. For example, when the color mode is a vertical mode, it may overlap with a vertical mode, which is one of the preset prediction modes.
[0437] When the number of prediction mode candidates is adaptively adjusted according to the encoding / decoding setting, if the overlap occurs, the number of candidates can be adjusted (8 to 7). Alternatively, when the number of prediction mode candidates is kept fixed, if the overlap occurs, an index can be assigned by adding and considering other candidates, and this setting will be assumed below. In addition, the adaptive prediction mode candidates can be configured to be supported even when a variable mode such as a color mode is included. Therefore, the case of adaptive index allocation can be regarded as an example of an adaptive prediction mode candidate configuration.
[0438] Next, a case where indexes are assigned adaptively according to a color mode will be described. It is assumed that basic indexes are assigned in the following order: Planar(0)-Vertical(1)-Horizontal(2)-DC(3)-CP1(4)-CP2(5)-CP3(6)-C(7). In addition, if the color mode does not match the preset prediction mode, index assignment is performed in the above order.
[0439] For example, if the color mode matches one of the preset prediction modes (planar, vertical, horizontal, DC mode), the matching prediction mode is filled in the color mode index (7). The matching prediction mode index (one of 0 to 3) is filled in with the preset prediction mode (diagoanal down left). In particular, if the color mode is a horizontal mode, index assignment may be performed as follows: planar (0)-vertical (1)-diagoanal down left (2)-DC (3)-CP1 (4)-CP2 (5)-CP3 (6)-horizontal (7).
[0440] Alternatively, if the color mode matches one of the preset prediction modes, the prediction mode that matches the 0th index is filled. Then, the preset prediction mode (Diagonal down left) is filled in index (7) of the color mode. Here, if the filled prediction mode is not the existing 0th index (i.e., not the planar mode), the existing index configuration can be adjusted. In particular, if the color mode is DC mode, index assignment such as DC(0)-Planar(1)-Vertical(2)-Horizontal(3)-CP1(4)-CP2(5)-CP3(6)-Diagonal down left(7) can be performed.
[0441] The above examples are only some examples of adaptive index allocation, and various modifications are possible. Also, binarization such as fixed length binarization, unary binarization, truncated unary binarization, and truncated rice binarization can be used based on the indexes in the candidate set.
[0442] Next, another example of performing encoding by assigning indexes to prediction modes belonging to a prediction mode candidate group will be described.
[0443] For example, a method of classifying prediction mode candidates into a plurality of groups according to prediction mode, prediction method, etc., and assigning an index to a prediction mode belonging to the corresponding candidate group and encoding it corresponds to this method. In this case, the candidate group selection information encoding can be performed prior to the index encoding. As an example, the directional mode, the non-directional mode, and the color mode, which are prediction modes that perform prediction in the same color space, can belong to one candidate group (hereinafter, S candidate group), and the color copy mode, which is a prediction mode that performs prediction in another color space, can belong to one candidate group (hereinafter, D candidate group).
[0444] Next, it is assumed that nine prediction modes including a preset prediction mode, a color copy mode, and a color mode are supported in the prediction mode candidate group (color difference component).
[0445] For example, assume that the preset prediction modes support four of the planar, DC, horizontal, vertical, and diagonal modes, one color mode (C), and four color copy modes (CP1, CP2, CP3, and CP4). The S candidate group may have five candidates consisting of the preset prediction modes and color modes, and the D candidate group may have four candidates consisting of the color copy modes.
[0446] The S candidate group is an example of an adaptively configured prediction mode candidate group, and an example of adaptive index allocation has been described above, so a detailed description will be omitted. The D candidate group is an example of a fixedly configured prediction mode candidate group, so a fixed index allocation method can be used. For example, index allocation such as CP1(0)-CP2(1)-CP3(2)-CP4(3) can be performed.
[0447] Based on the index in the candidate set, binarization such as fixed length binarization, unary binarization, truncated unary binarization, truncated rice binarization, etc. may be used, and is not limited to the above examples. Various modifications are possible.
[0448] The prediction related information generated by the prediction mode encoding unit can be transmitted to an encoding unit and included in a bitstream.
[0449] In the video decoding method according to an embodiment of the present invention, the intra prediction may be configured as follows. The intra prediction of the predictor may include a prediction mode decoding step, a reference pixel configuration step, and a prediction block generation step. Also, the video decoding apparatus may be configured to include a prediction mode decoding unit, a reference pixel configuration unit, and a prediction block generation unit, which implement the prediction mode decoding step, the reference pixel configuration step, and the prediction block generation step. Some of the above-described steps may be omitted or other steps may be added, and the steps may be changed to an order other than the above-described order.
[0450] The reference pixel construction unit and prediction block generation unit of the video decoding device perform the same functions as the corresponding components of the video encoding device, so detailed descriptions are omitted, and the prediction mode decoding unit can be performed by using the same method used in the prediction mode encoding unit in reverse.
[0451] FIG. 21 is a diagram showing an inter prediction method according to an embodiment to which the present invention is applied.
[0452] Referring to FIG. 21, a candidate list for predicting motion information of a current block can be generated (S2100).
[0453] The candidate list may include one or more candidates based on an affine model (hereinafter, referred to as affine candidates). An affine candidate may refer to a candidate having a control point vector. The control point vector refers to a motion vector of a control point for an affine model, and may be defined with respect to a corner position of a block (e.g., at least one position of the upper left corner, the upper right corner, the lower left corner, or the lower right corner).
[0454] The affine candidates may include at least one of spatial candidates, temporal candidates, and configuration candidates. Here, the spatial candidates may be derived from vectors of neighboring blocks spatially adjacent to the current block, and the temporal candidates may be derived from vectors of neighboring blocks temporally adjacent to the current block. Here, the neighboring blocks may refer to blocks coded using an affine model. The vectors may refer to motion vectors or control point vectors.
[0455] A method for deriving spatial / temporal candidates based on vectors of spatial / temporal neighboring blocks will be described in detail with reference to FIG.
[0456] Meanwhile, the configuration candidates can be derived based on a combination between motion vectors of spatial / temporal neighboring blocks to the current block, which will be described in detail with reference to FIG.
[0457] The plurality of affine candidates may be arranged in the candidate list based on a predetermined priority. For example, the plurality of affine candidates may be arranged in the candidate list in the order of spatial candidates, temporal candidates, and compositional candidates. Alternatively, the plurality of affine candidates may be arranged in the candidate list in the order of temporal candidates, spatial candidates, and compositional candidates. However, without being limited thereto, the temporal candidates may be arranged next to the compositional candidates. Alternatively, some of the compositional candidates may be arranged before the spatial candidates, and the rest may be arranged after the spatial candidates.
[0458] A control point vector of the current block can be derived based on the candidate list and the candidate index (S2110).
[0459] The candidate index may refer to an index coded to derive a control point vector of a current block. The candidate index may identify one of a plurality of affine candidates belonging to a candidate list. The control point vector of the current block may be derived using the control point vector of the affine candidate identified by the candidate index.
[0460] For example, assume that the type of the affine model of the current block is 4-parameter (i.e., the current block is determined to use two control point vectors). Here, if the affine candidate identified by the candidate index has three control point vectors, only two control point vectors (e.g., the control point vectors with Idx=0, 1) can be selected from the three control point vectors and set as the control point vectors of the current block. Alternatively, the three control point vectors of the identified affine candidate can be set as the control point vectors of the current block. In this case, the type of the affine model of the current block can be updated to 6-parameter.
[0461] Conversely, assume that the type of the affine model of the current block is 6-parameter (i.e., when the current block is determined to use three control point vectors). Here, if the affine candidate specified by the candidate index has two control point vectors, one additional control point vector can be generated, and the two control point vectors of the affine candidate and the additional control point vector can be set as the control point vectors of the current block. The additional control point vector can be derived based on at least one of the two control point vectors of the affine candidate, the size or position information of the current / peripheral block. Alternatively, the two control point vectors of the specified affine candidate can be set as the control point vectors of the current block. In this case, the type of the affine model of the current block can be updated to 4-parameter.
[0462] The motion vector of the current block can be derived based on the control point vectors of the current block (S2120).
[0463] The motion vector can be derived in units of sub-blocks of the current block. Here, the N×M sub-blocks can be in the form of a rectangle (N>M or N<M) or a square (N = M). The values of N and M can be 4, 8, 16, 32 or more. The size / form of the sub-block can be a fixed size / form predefined in the decoder.
[0464] Alternatively, the size / form of the sub-block can also be variably derived based on the attributes of the aforementioned block. For example, if the size of the current block is larger than or equal to a predetermined threshold, the current block can be divided into units of the first sub-block (e.g., 8×8, 16×16), and otherwise, the current block can be divided into units of the second sub-block (e.g., 4×4). Alternatively, information about the size / form of the sub-block can also be encoded and signaled by the encoder.
[0465] Using the derived motion vector, inter prediction can be performed on the current block (S2130).
[0466] Specifically, the reference block may be identified using a motion vector of the current block. The reference block may be identified for each sub-block of the current block. The reference block of each sub-block may belong to one reference picture. That is, the sub-blocks belonging to the current block may share one reference picture. Alternatively, a reference picture index may be set independently for each sub-block of the current block.
[0467] The identified reference block may be set as a predicted block of the current block. The above-described embodiment may be applied equally / similarly to not only merge mode but also general inter modes (e.g., AMVP mode). The above-described embodiment may be performed only if the size of the current block is greater than or equal to a predetermined threshold. Here, the threshold may be 8x8, 8x16, 16x8, 16x16, or more.
[0468] FIG. 22 is a diagram illustrating a method for deriving affine candidates from spatial / temporal neighboring blocks according to one embodiment of the present invention.
[0469] For convenience of explanation, in this embodiment, a method for deriving affine candidates from spatially surrounding blocks is described.
[0470] 5, the width and height of the current block 2200 are cbW and cbH, respectively, and the position of the current block is (xCb, yCb). The width and height of the spatially surrounding block 2210 are nbW and nbH, respectively, and the position of the spatially surrounding block is (xNb, yNb). FIG. 22 illustrates the spatially surrounding block as the top left block of the current block, but is not limited thereto. That is, the spatially surrounding block may include at least one of the left block, the bottom left block, the top right block, the top block, or the top left block of the current block.
[0471] A spatial candidate may have n control point vectors (cpMVs), where the value of n may be an integer of 1, 2, 3, or more. The value of n may be determined based on at least one of information on whether the block is decoded in a subblock unit, information on whether the block is coded in an affine model, or information on the type of affine model (4-parameter or 6-parameter).
[0472] The information may be coded and signaled by the coding device. Alternatively, all or a part of the information may be derived by the decoding device based on the attributes of the block. Here, the block may mean the current block or the spatial / temporal neighboring blocks of the current block. The attributes may mean size, shape, position, partition type, inter mode, parameters related to residual coefficients, etc. The inter mode is a mode predefined in the decoding device, and may mean merge mode, skip mode, AMVP mode, affine model, intra / inter combined mode, current picture reference mode, etc. Alternatively, the n value may be derived by the decoding device based on the attributes of the block.
[0473] In this embodiment, the n control point vectors may be expressed as a first control point vector (cpMV[0]), a second control point vector (cpMV[1]), a third control point vector (cpMV[2]), ..., an nth control point vector (cpMV[n-1]). As an example, the first control point vector (cpMV[0]), the second control point vector (cpMV[1]), the third control point vector (cpMV[2]), and the fourth control point vector (cpMV[3]) may be vectors corresponding to the positions of the upper left sample, the upper right sample, the lower left sample, and the lower right sample of the block, respectively. Here, it is assumed that the spatial candidate has three control point vectors, and the three control point vectors may be any control point vectors selected from the first to nth control point vectors. However, without being limited thereto, the spatial candidate may have two control point vectors, and the two control point vectors may be any control point vectors selected from the first to nth control point vectors.
[0474] Meanwhile, the control point vector of the spatial candidate may be derived differently depending on whether the boundary 2220 shown in FIG. 22 is a coding tree block boundary (CTU boundary).
[0475] 1. When the boundary 2220 of the current block does not border the coding tree block boundary (CTU boundary):
[0476] The first control point vector can be derived based on at least one of the first control point vector of the spatially surrounding block, a predetermined difference value, position information (xCb, yCb) of the current block, or position information (xNb, yNb) of the spatially surrounding block.
[0477] The number of the difference values may be one, two, three or more. The number of the difference values may be variably determined in consideration of the above-mentioned block attributes, or may be a fixed value pre-assigned to the decoding device. The difference value may be defined as a difference value between any one of a plurality of control point vectors and another one. For example, the difference value may include at least one of a first difference value between a second control point vector and a first control point vector, a second difference value between a third control point vector and a first control point vector, a third difference value between a fourth control point vector and a third control point vector, or a fourth difference value between a fourth control point vector and a second control point vector.
[0478] For example, the first control point vector can be derived by the following mathematical formula 1:
number
[0479] In Equation 1, the variables mvScaleHor and mvScaleVer may represent the first control point vector of the spatial peripheral block, or may represent a value derived by applying a shift operation of k to the first control point vector, where k may be an integer of 1, 2, 3, 4, 5, 6, 7, 8, 9 or more. The variables dHorX and dVerX correspond to the x and y components of the first difference value between the second control point vector and the first control point vector, respectively. The variables dHorY and dVerY correspond to the x and y components of the second difference value between the third control point vector and the first control point vector, respectively. The above variables may be derived according to Equation 2 below.
[0480]
number
[0481] The second control point vector may be derived based on at least one of the first control point vector of the spatial neighboring block, a predetermined difference value, position information (xCb, yCb) of the current block, block size (width or height), or position information (xNb, yNb) of the spatial neighboring block. Here, the block size may refer to the size of the current block and / or the spatial neighboring block. The difference value is the same as that described in the first control point vector, so a detailed description will be omitted here. However, the range and / or number of difference values used in the process of deriving the second control point vector may be different from that of the first control point vector.
[0482] For example, the second control point vector can be derived by the following mathematical formula 3.
[0483]
number
[0484] In Equation 3, the variables mvScaleHor, mvScaleVer, dHorX, dVerX, dHorY, and dVerY are the same as those described in Equation 1, so detailed description thereof will be omitted here.
[0485] The third control point vector may be derived based on at least one of the first control point vector of the spatial neighboring block, a predetermined difference value, position information (xCb, yCb) of the current block, block size (width or height), or position information (xNb, yNb) of the spatial neighboring block. Here, the block size may refer to the size of the current block and / or the spatial neighboring block. The difference value is the same as that described in the first control point vector, so a detailed description will be omitted here. However, the range and / or number of difference values used in the process of deriving the third control point vector may be different from that of the first control point vector or the second control point vector.
[0486] For example, the third control point vector can be derived by the following mathematical formula 4.
number
[0487] In Equation 4, variables mvScaleHor, mvScaleVer, dHorX, dVerX, dHorY, and dVerY are the same as those described in Equation 1, so detailed description will be omitted here. Meanwhile, the n-th control point vector of the spatial candidate can be derived through the above-mentioned process.
[0488] 2. When the boundary 2220 of the current block is adjacent to the boundary of the coding tree block (CTU boundary),
[0489] The first control point vector can be derived based on at least one of a motion vector (MV) of a spatially surrounding block, a predetermined difference value, position information (xCb, yCb) of the current block, or position information (xNb, yNb) of a spatially surrounding block.
[0490] The motion vector may be a motion vector of a sub-block located at the bottom of the spatially peripheral block. The sub-block may be located at the leftmost, center, or rightmost of a plurality of sub-blocks located at the bottom of the spatially peripheral block. Alternatively, the motion vector may represent an average value, a maximum value, or a minimum value of the motion vectors of the sub-blocks.
[0491] The number of the difference values may be 1, 2, 3 or more. The number of the difference values may be variably determined in consideration of the above-mentioned block attributes, or may be a fixed value pre-defined to the decoding device. The difference value may be defined as a difference value between any one of a plurality of motion vectors stored in sub-block units in the spatially neighboring block and another one. For example, the difference value may mean a difference value between the motion vector of the lower right sub-block and the motion vector of the lower left sub-block of the spatially neighboring block.
[0492] For example, the first control point vector can be derived by the following mathematical formula 5.
number
[0493] In Equation 5, the variables mvScaleHor and mvScaleVer may represent the motion vector (MV) of the spatially surrounding block or a value derived by applying a k-shift operation to the motion vector, where k may be an integer of 1, 2, 3, 4, 5, 6, 7, 8, 9, or more.
[0494] The variables dHorX and dVerX correspond to the x and y components of a predetermined difference value, respectively. Here, the difference value means a difference value between a motion vector of a rightmost sub-block and a motion vector of a leftmost sub-block in a spatially peripheral block. The variables dHorY and dVerY can be derived based on the variables dHorX and dVerX. The above variables can be derived according to the following Equation 6.
[0495]
number
[0496] The second control point vector may be derived based on at least one of the motion vector (MV) of the spatial surrounding block, a predetermined difference value, position information (xCb, yCb) of the current block, block size (width or height) or position information (xNb, yNb) of the spatial surrounding block. Here, the block size may refer to the size of the current block and / or the spatial surrounding block. The motion vector and difference value are the same as those described in the first control point vector, so a detailed description will be omitted here. However, the position of the motion vector, the range of the difference value and / or the number of the difference values used in the process of deriving the second control point vector may be different from that of the first control point vector.
[0497] For example, the second control point vector can be derived by the following mathematical formula 7.
number
[0498] In Equation 7, the variables mvScaleHor, mvScaleVer, dHorX, dVerX, dHorY, and dVerY are the same as those described in Equation 5, so detailed description thereof will be omitted here.
[0499] The third control point vector may be derived based on at least one of the motion vector (MV) of the spatial surrounding block, a predetermined difference value, position information (xCb, yCb) of the current block, block size (width or height) or position information (xNb, yNb) of the spatial surrounding block. Here, the block size may refer to the size of the current block and / or the spatial surrounding block. The motion vector and difference value are the same as those described in the first control point vector, so a detailed description will be omitted here. However, the position of the motion vector, the range of the difference value and / or the number of the motion vectors used in the process of deriving the third control point vector may be different from those of the first control point vector or the second control point vector.
[0500] For example, the third control point vector can be derived by the following mathematical formula 8.
number
[0501] In Equation 8, variables mvScaleHor, mvScaleVer, dHorX, dVerX, dHorY, and dVerY are the same as those described in Equation 5, so detailed description will be omitted here. Meanwhile, the n-th control point vector of the spatial candidate can be derived through the above-mentioned process.
[0502] The above-described affine candidate derivation process may be performed for each of predefined spatially neighboring blocks, which may include at least one of a left block, a lower left block, a upper right block, a top block, or an upper left block of a current block.
[0503] Alternatively, the affine candidate derivation process may be performed for each group of the spatially neighboring blocks, where the spatially neighboring blocks may be classified into a first group including a left block and a bottom left block, and a second group including a top right block, a top block, and a top left block.
[0504] For example, one affine candidate may be derived from the spatially neighboring blocks belonging to the first group. The derivation may be performed until a usable affine candidate is found according to a predetermined priority. The priority may be in the order of the left block→the bottom left block, or in the reverse order.
[0505] Similarly, one affine candidate can be derived from the spatially neighboring blocks belonging to the second group. The derivation can be performed until a usable affine candidate is found according to a predetermined priority. The priority can be in the order of the top right block→top block→top left block, or in the reverse order.
[0506] The above embodiment can be applied to the temporal neighboring blocks in the same manner. Here, the temporal neighboring blocks may be blocks that belong to a different picture from the current block but are located at the same position as the current block. The block at the same position may be a block including the position of the top left sample of the current block, the center position, or the position of a sample adjacent to the bottom right sample of the current block.
[0507] Alternatively, the temporally neighboring block may refer to a block shifted by a predetermined displacement vector from the block at the same position, where the displacement vector may be determined based on the motion vector of any one of the spatially neighboring blocks of the current block.
[0508] FIG. 23 is a diagram illustrating a method for deriving configuration candidates based on a combination of motion vectors of spatial / temporal neighboring blocks according to an embodiment of the present invention.
[0509] The configuration candidates of the present invention can be derived based on at least two combinations of control point vectors (hereinafter referred to as control point vectors (cpMVCorner[n])) corresponding to each corner of the current block, where n can be 0, 1, 2, or 3.
[0510] The control point vector may be derived based on the motion vector of a spatially neighboring block and / or a temporally neighboring block. Here, the spatially neighboring block may include at least one of a first neighboring block (C, D, or E) adjacent to the top left sample of the current block, a second neighboring block (F or G) adjacent to the top right sample of the current block, or a third neighboring block (A or B) adjacent to the bottom left sample of the current block. The temporally neighboring block may refer to a fourth neighboring block (Col) adjacent to the bottom right sample of the current block, which belongs to a different picture from the current block.
[0511] The first neighboring block may refer to the neighboring block at the top left (D), top (E), or left (C) of the current block. It may be determined whether the motion vectors of the neighboring blocks C, D, and E are available according to a predetermined priority order, and the control point vector may be determined using the motion vectors of the available neighboring blocks. The availability determination may be performed until a neighboring block having an available motion vector is found. Here, the priority order may be D→E→C. However, the priority order is not limited thereto, and may be D→C→E, C→D→E, or E→D→C.
[0512] The second neighboring block may refer to a neighboring block at the top (F) or top right (G) of the current block. Similarly, it may be determined whether the motion vectors of neighboring blocks F and G are available according to a predetermined priority, and a control point vector may be determined using the motion vectors of the available neighboring blocks. The availability determination may be performed until a neighboring block having an available motion vector is found. Here, the priority may be F→G or G→F.
[0513] The third neighboring block may refer to a neighboring block on the left side (B) or the bottom left end (A) of the current block. Similarly, it may be determined whether the motion vector of the neighboring block is available according to a predetermined priority, and the control point vector may be determined using the motion vector of the available neighboring block. The availability determination may be performed until a neighboring block having an available motion vector is found. Here, the priority may be in the order of A→B or B→A.
[0514] For example, the first control point vector (cpMVCorner[0]) can be set to the motion vector of the first surrounding block, the second control point vector (cpMVCorner[1]) can be set to the motion vector of the second surrounding block, the third control point vector (cpMVCorner[2]) can be set to the motion vector of the third surrounding block, and the fourth control point vector (cpMVCorner[3]) can be set to the motion vector of the fourth surrounding block.
[0515] Alternatively, any one of the first to fourth control point vectors may be derived based on the other one. For example, the second control point vector may be derived by applying a predetermined offset vector to the first control point vector. The offset vector may be a difference vector between the third control point vector and the first control point vector, or may be derived by applying a predetermined scaling factor to the difference vector. The scaling factor may be determined based on at least one of the width and height of the current block and / or the surrounding block.
[0516] According to the present invention, K configuration candidates (ConstK) can be determined by a combination of at least two of the first to fourth control point vectors. The value of K can be an integer of 1, 2, 3, 4, 5, 6, 7 or more. The value of K can be derived based on information signaled by the encoding device, or can be a value already promised to the decoding device. The information can include information indicating the maximum number of configuration candidates included in a candidate list.
[0517] Specifically, the first configuration candidate (Const1) may be derived by combining the first to third control point vectors. For example, the first configuration candidate (Const1) may have a control point vector as shown in Table 1 below. Meanwhile, only when the reference picture information of the first peripheral block is the same as the reference picture information of the second and third peripheral blocks, the control point vector may be restricted to be configured as shown in Table 1. Here, the reference picture information may refer to a reference picture index indicating the position of the corresponding reference picture in a reference picture list, or may refer to a POC (picture order count) value indicating the output order.
[0518] [Table 1] The second configuration candidate (Const2) can be derived by combining the first, second and fourth control point vectors. For example, the second configuration candidate (Const2) can have a control point vector as shown in Table 2 below. Meanwhile, only when the reference picture information of the first peripheral block is the same as the reference picture information of the second and fourth peripheral blocks, the control point vector can be restricted to be configured as shown in Table 2. Here, the reference picture information is as described above.
[0519] [Table 2] The third configuration candidate (Const3) can be derived by combining the first, third and fourth control point vectors. For example, the third configuration candidate (Const3) can have a control point vector as shown in Table 3 below. Meanwhile, only when the reference picture information of the first peripheral block is the same as the reference picture information of the third and fourth peripheral blocks, the control point vector can be restricted to be configured as shown in Table 2. Here, the reference picture information is as described above.
[0520] [Table 3] The fourth configuration candidate (Const4) can be derived by combining the second, third and fourth control point vectors. For example, the fourth configuration candidate (Const4) can have a control point vector as shown in Table 4 below. Meanwhile, only when the reference picture information of the second peripheral block is the same as the reference picture information of the third and fourth peripheral blocks, it can be restricted to be configured as shown in Table 4. Here, the reference picture information is as described above.
[0521] [Table 4] The fifth configuration candidate (Const5) can be derived by combining the first and second control point vectors. For example, the fifth configuration candidate (Const5) can have a control point vector as shown in Table 5 below. Meanwhile, only when the reference picture information of the first peripheral block is the same as the reference picture information of the second peripheral block, the control point vector can be restricted to be configured as shown in Table 5. Here, the reference picture information is as described above.
[0522] [Table 5] The sixth configuration candidate (Const6) can be derived by combining the first and third control point vectors. For example, the sixth configuration candidate (Const6) can have a control point vector as shown in Table 6 below. Meanwhile, only when the reference picture information of the first peripheral block is the same as the reference picture information of the third peripheral block, the control point vector can be restricted to be configured as shown in Table 6. Here, the reference picture information is as described above.
[0523] [Table 6] In Table 6, cpMvCorner[1] may be a second control point vector derived based on the first and third control point vectors. The second control point vector may be derived based on at least one of the first control point vector, a predetermined difference value, or the size of the current / neighboring block. For example, the second control point vector may be derived according to the following mathematical formula 9.
[0524]
number
[0525] The first to sixth configuration candidates described above may be all or only some of them may be included in the candidate list.
[0526] The methods according to the present invention may be embodied in the form of program instructions executable by various computer means and stored on a computer readable medium. The computer readable medium may include, alone or in combination with program instructions, data files, data structures, and the like. The program instructions stored on the computer readable medium may be those specially designed and constructed for the present invention, or they may be of the kind commonly known and available to those skilled in the art of computer software.
[0527] Examples of computer readable media include hardware devices specially configured to store and execute program instructions, such as Read Only Memory (ROM), RAM, flash memory, etc. Examples of program instructions include high level language code executable by a computer using an interpreter, etc., as well as machine code, such as that produced by a compiler. The above-mentioned hardware devices may be configured to operate as at least one software module to perform the operations of the present invention, and vice versa.
[0528] Furthermore, the above-described methods or apparatuses may be embodied in such a manner that all or part of their components or functions are combined or separated.
[0529] Although the present invention has been described above with reference to preferred embodiments thereof, those skilled in the art will understand that the present invention can be modified and changed in various ways without departing from the spirit and scope of the present invention as set forth in the following claims. [Industrial Applicability]
[0530] The present invention can be used to encode / decode video signals. < / l-1> < / array> < / qt>
Claims
1. obtaining, from the bitstream, first information indicating whether a color copy mode is supported; determining whether the color copy mode is applied to a current chrominance block of an image if the first information indicates that the color copy mode is supported; downsampling a luma block corresponding to the current chrominance block in response to determining that the color copy mode is applied to the current chrominance block; deriving correlation information of the color copy mode based on a predetermined reference area; predicting the current chrominance block by applying the correlation information to the downsampled luminance block; Equipped with Whether the color copy mode is supported is further determined in different ways depending on the image type; the image type indicates a slice type of the current chrominance block; Video decoding method.
2. the downsampling of the luma block is performed using a first pixel of the luma block corresponding to a current pixel of the current chroma block and a number of neighboring pixels adjacent to the first pixel; The video decoding method according to claim 1.
3. The adjacent pixel includes at least one of a right adjacent pixel, a lower adjacent pixel, or a lower right adjacent pixel; The video decoding method according to claim 2 .
4. The reference region includes at least one of an adjacent region of the current chrominance block or an adjacent region of the luminance block; The video decoding method according to claim 1.
5. a number of reference pixel lines belonging to the reference region used to derive the correlation information is adaptively determined according to whether the current chrominance block is located on a boundary; The video decoding method according to claim 4.
6. the correlation information is derived using some, but not all, pixels of the reference region; The video decoding method according to claim 1.
7. the correlation information comprises at least one weight parameter and at least one offset parameter. The video decoding method according to claim 6.
8. the number of correlation information applied to the downsampled luminance block is less than or equal to three; The video decoding method according to claim 7.
9. obtaining, from the bitstream, first information indicating whether a color copy mode is supported; determining whether the color copy mode is applied to a current chrominance block of an image if the color copy mode is supported; downsampling a luminance block corresponding to the current chrominance block in response to determining that the color copy mode is applied to the current chrominance block; determining correlation information of the color copy mode based on a predetermined reference area; predicting the current chrominance block by applying the correlation information to the downsampled luminance block; Equipped with Whether the color copy mode is supported is further determined in different ways depending on the image type; the image type indicates a slice type of the current chrominance block; Video decoding method.
10. When executed by a processor, obtaining, from the bitstream, first information indicating whether a color copy mode is supported; determining whether the color copy mode is applied to a current chrominance block of an image if the color copy mode is supported; downsampling a luminance block corresponding to the current chrominance block in response to determining that the color copy mode is applied to the current chrominance block; deriving correlation information of the color copy mode based on a predetermined reference area; predicting the current chrominance block by applying the correlation information to the downsampled luminance block; A method comprising: Whether the color copy mode is supported is further determined in different ways depending on the image type; the image type indicates a slice type of the current chrominance block; method, A computer-readable medium having stored thereon instructions for carrying out the steps of:
11. determining whether a color copy mode is supported and encoding first information indicative of whether the color copy mode is supported; determining whether the color copy mode is applied to a current chrominance block of an image if the color copy mode is supported; downsampling a luma block corresponding to the current chrominance block in response to determining that the color copy mode is applied to the current chrominance block; determining correlation information of the color copy mode based on a predetermined reference area; predicting the current chrominance block by applying the correlation information to the downsampled luminance block; Equipped with Whether the color copy mode is supported is further determined in different ways depending on the image type; the image type indicates a slice type of the current chrominance block; A method for transmitting a bitstream generated by a method for encoding video.
Citation Information
Patent Citations
Linear model chroma intra prediction for video coding
WO2018053293A1
Cited By
Image encoding / decoding method and device
JP2025163176A