Method and apparatus for video coding with intra prediction of chrominance blocks based on geometric segmentation
By performing intra prediction of chromaticity blocks based on the reconstruction area and geometric segmentation mode in the video decoding device, the problems of low encoding efficiency and image quality in the prior art are solved, and more efficient video encoding and better video quality are achieved.
Patent Information
- Application Number
- CN202380080425.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-01
- Filing Date
- 2023-09-06
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, it is difficult to effectively use the geometric segmentation mode to perform intra prediction of chroma blocks in video encoding, resulting in insufficient encoding efficiency and image quality.
By establishing a corresponding brightness area from the reconstruction area of the brightness component in the video decoding device, and determining the prediction mode of the current chromaticity block is determined based on whether the spatial geometric segmentation mode is applied to the area, the intra prediction mode and the block segmentation structure, and finally generating the prediction block of the current chromaticity block.
Improve video encoding efficiency and enhance video quality, especially in the intra prediction process of chroma blocks, the geometric segmentation mode is more efficiently utilized.
Smart Images

Figure CN120226349A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a video coding method and apparatus for intra prediction of chrominance blocks based on geometric partitioning. Background Art
[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.
[0003] Since video data has a large amount of data compared to audio or still image data, video data requires a large amount of hardware resources (including memory) to store or transmit video data without processing for compression.
[0004] Thus, an encoder is typically used to compress and store or transmit video data. A decoder receives the compressed video data, decompresses the received compressed video data, and plays the decompressed video data. Video compression techniques include H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC), which have about 30% or more improved coding efficiency compared to HEVC.
[0005] However, since the image size, resolution, and frame rate are gradually increasing, the amount of data to be encoded also increases. Thus, there is a need to provide a new compression technique with higher coding efficiency and improved image enhancement effects than existing compression techniques.
[0006] The VVC technology employs a Geometric Partitioning Mode (GPM)-inter prediction technique for prediction based on a partitioning more flexible than the square and rectangular partitioning of a Quadtree Plus Multi-Type Tree (QT+MTT) partitioning structure. The GPM performs prediction of the current block based on the mode index and motion vector information of the partitioning region. Here, the mode index indicates one of the predefined geometric partitioning modes for partitioning the current block, and the motion vector information can be derived to predict sub-regions by partitioning the current block.
[0007] The encoder transmits the mode index and motion vector information to the decoder. The decoder divides the current block into two regions according to the parsed geometric partitioning mode. The decoder uses the motion vector information to generate a prediction signal for each sub-region, and then performs a weighted sum of the generated prediction signals to generate a final prediction block. At this time, the weight for the weighted sum can be determined based on the geometric partitioning mode, and the mixing region can be determined as a fixed region based on the size of the current block and the geometric partitioning mode.
[0008] Meanwhile, the Spatial Geometry Partitioning Mode (SGPM) is suitable for intra prediction of the current block. SGPM is an intra prediction technique that is equally applied to the luminance component and the chrominance component. Therefore, in order to improve video coding efficiency and enhance video quality, a method for effectively utilizing geometric partitioning in the intra prediction of chrominance blocks is needed. Summary of the Invention
[0009] Technical Problem
[0010] The present disclosure attempts to provide a video coding method and apparatus for intra predicting a current chrominance block by using geometric partitioning when predicting a chrominance component after predicting and reconstructing a luminance component for the current block.
[0011] Technical Solution
[0012] At least one aspect of the present disclosure provides a method for reconstructing a current chrominance block by a video decoding apparatus. The method includes establishing a corresponding luminance region of the current chrominance block from a reconstructed region of the luminance component. The corresponding luminance region is an intra prediction region and represents a luminance block or a luminance region located at the same position as the current chrominance block. The method further includes obtaining a prediction mode of the current chrominance block based on whether the Spatial Geometry Partitioning Mode is applied to the corresponding luminance region, the intra prediction mode of the corresponding luminance region, and the block partitioning structures of the luminance component and the chrominance component. The method further includes generating a prediction block of the current chrominance block based on the prediction mode of the current chrominance block.
[0013] Another aspect of the present disclosure provides a method for encoding a current chrominance block by a video coding apparatus. The method includes establishing a corresponding luminance region of the current chrominance block from a reconstructed region of the luminance component. The corresponding luminance region is an intra prediction region and represents a luminance block or a luminance region located at the same position as the current chrominance block. The method further includes determining a prediction mode of the current chrominance block based on whether the Spatial Geometry Partitioning Mode is applied to the corresponding luminance region, the intra prediction mode of the corresponding luminance region, and the block partitioning structures of the luminance component and the chrominance component. The method further includes generating a prediction block of the current chrominance block based on the prediction mode of the current chrominance block.
[0014] Yet another aspect of the present disclosure provides a computer-readable recording medium storing a bitstream generated by a video coding method. The video coding method includes establishing a corresponding luminance region of the current chrominance block from a reconstructed region of the luminance component. The corresponding luminance region is an intra prediction region and represents a luminance block or a luminance region located at the same position as the current chrominance block. The video coding method further includes determining a prediction mode of the current chrominance block based on whether the Spatial Geometry Partitioning Mode is applied to the corresponding luminance region, the intra prediction mode of the corresponding luminance region, and the block partitioning structures of the luminance component and the chrominance component. The video coding method further includes generating a prediction block of the current chrominance block based on the prediction mode of the current chrominance block.
[0015] Beneficial effects
[0016] As described above, the present disclosure provides a video encoding method and apparatus for searching for a plurality of candidate prediction blocks by using template matching and generating a prediction block of a current block from the searched plurality of candidate prediction blocks. Accordingly, the video encoding method and apparatus improve video encoding efficiency and improve video quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a block diagram of a video encoding apparatus that can implement the technology of the present disclosure.
[0018] Figure 2 illustrates a method for partitioning a block by using a quadtree plus binary tree plus ternary tree (QTBTTT) structure.
[0019] Figure 3a and Figure 3b illustrates a plurality of intra prediction modes including a wide-angle intra prediction mode.
[0020] Figure 4 illustrates neighboring blocks of a current block.
[0021] Figure 5 is a block diagram of a video decoding apparatus that can implement the technology of the present disclosure.
[0022] Figure 6 is a block diagram of a detailed portion of a video decoding apparatus according to at least one embodiment of the present disclosure.
[0023] Figure 7 is a diagram illustrating a partitioning pattern of a corresponding luminance block according to at least one embodiment of the present disclosure.
[0024] Figure 8a and Figure 8b is a diagram illustrating the determination of a prediction mode of a current chrominance block according to some embodiments of the present disclosure.
[0025] Figure 9 is a diagram illustrating samples of a luminance component and a chrominance component for deriving a formula for cross-component prediction (CC-based prediction) according to at least one embodiment of the present disclosure.
[0026] Figures 10a to 10c is a diagram illustrating the determination of a prediction mode of a current chrominance block according to some embodiments of the present disclosure.
[0027] Figure 11 is a diagram illustrating the determination of a prediction mode of a current chrominance block according to at least one embodiment of the present disclosure.
[0028] Figure 12Is a diagram showing samples of a luminance component and a chrominance component for deriving a formula for CC-based prediction according to at least one embodiment of the present disclosure.
[0029] Figure 13 Is a diagram showing the determination of a prediction mode of a current chrominance block according to at least one embodiment of the present disclosure.
[0030] Figure 14 Is a diagram showing samples of a luminance component and a chrominance component for deriving a formula for CC-based prediction according to at least one embodiment of the present disclosure.
[0031] Figure 15 Is a diagram showing a plurality of predefined downsampling filters according to at least one embodiment of the present disclosure.
[0032] Figure 16 Is a flowchart of a method for predicting a current chrominance block by a video encoding device according to at least one embodiment of the present disclosure.
[0033] Figure 17 Is a flowchart of a method for predicting a current chrominance block by a video decoding device according to at least one embodiment of the present disclosure. Detailed Description of Specific Embodiments
[0034] Hereinafter, some embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In the following description, although elements are shown in different drawings, the same reference numerals denote the same elements. In addition, in the following description of some embodiments, for the purpose of clarity and conciseness, detailed descriptions of related known components and functions may be omitted when it is considered that they may obscure the subject matter of the present disclosure.
[0035] Figure 1 Is a block diagram of a video encoding device that can implement the technology of the present disclosure. Hereinafter, with reference to Figure 1 the illustration, the video encoding device and the components of the device will be described.
[0036] The encoding device may include an image splitter 110, a predictor 120, a subtractor 130, a transformer 140, a quantizer 145, a rearrangement unit 150, an entropy encoder 155, an inverse quantizer 160, an inverse transformer 165, an adder 170, a loop filter unit 180, and a memory 190.
[0037] Each component of the encoding device may be implemented as hardware or software or as a combination of hardware and software. In addition, the function of each component may be implemented as software, and the microprocessor may also be implemented to execute the functions of the software corresponding to each component.
[0038] A video is composed of one or more sequences including a plurality of pictures. Each picture is divided into a plurality of regions, and encoding is performed on each region. For example, a picture is divided into one or more tiles and / or slices. Here, one or more tiles can be defined as a tile group. Each tile and / or slice is divided into one or more coding tree units (CTUs). Additionally, each CTU is divided into one or more coding units (CUs) through a tree structure. The information applied to each coding unit (CU) is encoded as the syntax of the CU, and the information commonly applied to the CUs included in one CTU is encoded as the syntax of the CTU. Furthermore, the information commonly applied to all blocks in a slice is encoded as the syntax of the slice header, and the information applied to all blocks constituting one or more pictures is encoded as a picture parameter set (PPS) or a picture header. Additionally, the information commonly referred to by multiple pictures is encoded into a sequence parameter set (SPS). Moreover, the information commonly referred to by one or more SPSs is encoded into a video parameter set (VPS). Additionally, the information commonly applied to a tile or a tile group can also be encoded as the syntax of a tile or tile group header. The syntax included in the SPS, PPS, slice header, tile or tile group header can be referred to as high-level syntax.
[0039] The image splitter 110 determines the size of a coding tree unit (CTU). The information regarding the size of the CTU (CTU size) is encoded as the syntax of the SPS or PPS and is delivered to the video decoding device.
[0040] The image splitter 110 divides each picture constituting the video into a plurality of coding tree units (CTUs) having a predetermined size, and then recursively divides the CTUs by using a tree structure. The leaf nodes in the tree structure become coding units (CUs), which are the basic units of encoding.
[0041] The tree structure can be a quadtree (QT), where a higher node (or parent node) is divided into four lower nodes (or child nodes) of the same size. The tree structure can also be a binary tree (BT), where a higher node is divided into two lower nodes. The tree structure can also be a ternary tree (TT), where a higher node is divided into three lower nodes in a 1:2:1 ratio. The tree structure can also be a structure in which two or more of the QT structure, BT structure, and TT structure are mixed. For example, a quadtree plus binary tree (QTBT) structure can be used or a quadtree plus binary tree ternary tree (QTBTTT) structure can be used. Here, adding a binary tree ternary tree (BTTT) to the tree structure is referred to as a multi-type tree (MTT).
[0042] Figure 2 is a diagram for describing a method of dividing blocks by using the QTBTTT structure.
[0043] As Figure 2 shown, the CTU can first be split into QT structures. The quadtree splitting can be recursive until the size of the split blocks reaches the minimum block size (MinQTSize) of the leaf nodes allowed in the QT. The first flag (QT_split_flag) indicating whether each node of the QT structure is split into four lower-level nodes is encoded by the entropy encoder 155 and signaled to the video decoding device. When the leaf nodes of the QT are not larger than the maximum block size (MaxBTSize) of the root nodes allowed in the BT, the leaf nodes can be further split into at least one of the BT structure and the TT structure. There can be multiple splitting directions in the BT structure and / or the TT structure. For example, there can be two directions, i.e., the direction in which the block of the corresponding node is split horizontally and the direction in which the block of the corresponding node is split vertically. As Figure 2 shown, when the MTT splitting starts, the second flag (mtt_split_flag) indicating whether the node is split, and additionally the flag indicating the splitting direction (vertical or horizontal) and / or the flag indicating the splitting type (binary or ternary) if the node is split are encoded by the entropy encoder 155 and signaled to the video decoding device.
[0044] Alternatively, before encoding the first flag (QT_split_flag) indicating whether each node is split into four lower-level nodes, the CU splitting flag (split_cu_flag) indicating whether the node is split can also be encoded. When the value of the CU splitting flag (split_cu_flag) indicates that each node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU as the basic unit of encoding. When the value of the CU splitting flag (split_cu_flag) indicates that each node is split, the video encoding device first starts encoding the first flag by the above scheme.
[0045] When QTBT is used as another example of the tree structure, there can be two types, i.e., the type in which the block of the corresponding node is split horizontally into two blocks of the same size (i.e., symmetric horizontal splitting) and the type in which the block of the corresponding node is split vertically into two blocks of the same size (i.e., symmetric vertical splitting). The splitting flag (split_flag) indicating whether each node of the BT structure is split into lower-level blocks and the splitting type information indicating the splitting type are encoded by the entropy encoder 155 and delivered to the video decoding device. At the same time, the type in which the block of the corresponding node is split into two asymmetric blocks can be additionally presented. The asymmetric form can include the form in which the block of the corresponding node is split into two rectangular blocks with a size ratio of 1:3, or can also include the form in which the block of the corresponding node is split along the diagonal direction.
[0046] The CU can have different sizes according to the QTBT or QTBTTT partitioning from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) is referred to as the "current block". Due to the adoption of QTBTTT partitioning, the shape of the current block can be a rectangular shape in addition to the square shape.
[0047] The predictor 120 predicts the current block to generate a predicted block. The predictor 120 includes an intra predictor 122 and an inter predictor 124.
[0048] Generally, each of the current blocks in a picture can be predictively encoded. Generally, the prediction of the current block can be performed by using intra prediction techniques (using data from the picture including the current block) or inter prediction techniques (using data from the pictures encoded before the picture including the current block). Inter prediction includes both uni-directional prediction and bi-directional prediction.
[0049] The intra predictor 122 predicts the pixels in the current block by using the pixels (reference pixels) located around the current block in the current picture including the current block. According to the prediction direction, there are multiple intra prediction modes. For example, as Figure 3a shown, the multiple intra prediction modes can include 2 non-directional modes including the planar mode and the DC mode, and can include 65 directional modes. The neighboring pixels to be used and the arithmetic formula are differently defined according to each prediction mode.
[0050] To perform effective directional prediction on the current block having a rectangular shape, directional modes (intra prediction modes #67 to #80, #-1 to #-14) as shown by the dashed arrows in Figure 3b can be additionally used. The directional modes can be referred to as "wide-angle intra prediction modes". In Figure 3b the arrow indicates the corresponding reference sample for prediction, rather than the prediction direction. The prediction direction is opposite to the direction shown by the arrow. When the current block has a rectangular shape, the wide-angle intra prediction mode is a mode that performs the prediction of a specific directional mode in the opposite direction without additional bit transmission. In this case, in the wide-angle intra prediction mode, some wide-angle intra prediction modes available for the current block can be determined by the ratio of the width and height of the current block having a rectangular shape. For example, when the current block has a rectangular shape with a height less than the width, the wide-angle intra prediction modes with an angle less than 45 degrees (intra prediction modes #67 to #80) are available. When the current block has a rectangular shape with a width greater than the height, the wide-angle intra prediction modes with an angle greater than -135 degrees are available.
[0051] The intra predictor 122 may determine an intra prediction to be used for encoding a current block. In some examples, the intra predictor 122 may encode the current block by using multiple intra prediction modes, and may also select an appropriate intra prediction mode to be used from the tested modes. For example, the intra predictor 122 may calculate rate-distortion values by using rate-distortion analysis for multiple tested intra prediction modes, and may also select an intra prediction mode having the best rate-distortion characteristics from the tested modes.
[0052] The intra predictor 122 selects one intra prediction mode from multiple intra prediction modes, and predicts the current block by using neighboring pixels (reference pixels) determined according to the selected intra prediction mode and arithmetic formulas. Information about the selected intra prediction mode is encoded by the entropy encoder 155 and is delivered to the video decoding device.
[0053] The inter predictor 124 generates a prediction block for the current block by using a motion compensation process. The inter predictor 124 searches for a block most similar to the current block in a reference picture encoded and decoded earlier than the current picture, and generates a prediction block for the current block by using the searched block. Additionally, a motion vector (MV) is generated, which corresponds to the displacement between the current block in the current picture and the prediction block in the reference picture. Generally, motion estimation is performed on the luminance component, and the motion vector calculated based on the luminance component is used for both the luminance component and the chrominance component. Motion information including information about the reference picture and information about the motion vector used for predicting the current block is encoded by the entropy encoder 155 and is delivered to the video decoding device.
[0054] The inter predictor 124 may also perform interpolation on the reference picture or reference block to improve prediction accuracy. In other words, subsamples between two consecutive integer samples are interpolated by applying filter coefficients to multiple consecutive integer samples including the two integer samples. When performing the process of searching for a block most similar to the current block for the interpolated reference picture, the motion vector may be represented not with integer sampling unit precision but with fractional unit precision. The precision or resolution of the motion vector may be set differently for each target region to be encoded (e.g., units such as slices, tiles, CTUs, CUs, etc.). When such an adaptive motion vector resolution (AMVR) is applied, information about the motion vector resolution to be applied to each target region should be signaled. For example, when the target region is a CU, information about the motion vector resolution applied to each CU is signaled. The information about the motion vector resolution may be information representing the precision of the motion vector difference described below.
[0055] Meanwhile, the inter-frame predictor 124 may perform inter-frame prediction by using bidirectional prediction. In the case of bidirectional prediction, two reference pictures and two motion vectors representing the positions of the blocks most similar to the current block in each reference picture are used. The inter-frame predictor 124 selects a first reference picture and a second reference picture from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively. The inter-frame predictor 124 also searches for the block most similar to the current block in the corresponding reference picture to generate a first reference block and a second reference block. Additionally, a predicted block for the current block is generated by averaging or weighted averaging the first reference block and the second reference block. Additionally, motion information including information about the two reference pictures used for predicting the current block and information about the two motion vectors is delivered to the entropy encoder 155. Here, reference picture list 0 may be composed of pictures that are in the display order before the current picture among the pre-reconstructed pictures, and reference picture list 1 may be composed of pictures that are in the display order after the current picture among the pre-reconstructed pictures. However, although not particularly limited thereto, pre-reconstructed pictures that are in the display order after the current picture may be additionally included in reference picture list 0. In contrast, pre-reconstructed pictures that are before the current picture may also be additionally included in reference picture list 1.
[0056] To minimize the amount of bits consumed for encoding motion information, various methods may be used.
[0057] For example, when the reference picture and motion vector of the current block are the same as those of a neighboring block, information identifying the neighboring block is encoded to deliver the motion information of the current block to the video decoding device. This method is referred to as the merge mode.
[0058] In the merge mode, the inter-frame predictor 124 selects a predetermined number of merge candidate blocks (hereinafter referred to as "merge candidates") from the neighboring blocks of the current block.
[0059] As Figure 4 shown, all or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current picture may be used as neighboring blocks for deriving merge candidates. Additionally, blocks other than the current picture in which the current block is located within the reference picture (which may be the same as or different from the reference picture used for predicting the current block) may also be used as merge candidates. For example, a block at the same position as the current block within the reference picture or a block adjacent to the block at the same position may be additionally used as a merge candidate. If the number of merge candidates selected by the method described above is less than the preset number, zero vectors are added to the merge candidates.
[0060] The inter-frame predictor 124 configures a merge list including a predetermined number of merge candidates by using neighboring blocks. A merge candidate to be used as motion information of the current block is selected from the merge candidates included in the merge list, and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoder 155 and delivered to the video decoding device.
[0061] Merge skip mode is a special case of merge mode. After quantization, when all transform coefficients for entropy coding are close to zero, only neighbor block selection information is transmitted without transmitting residual signals. By using merge skip mode, relatively high coding efficiency can be achieved for images with slight motion, still images, screen content images, etc.
[0062] Hereinafter, the merge mode and the merge skip mode are collectively referred to as the merge / skip mode.
[0063] Another method for encoding motion information is the Advanced Motion Vector Prediction (AMVP) mode.
[0064] In the AMVP mode, the inter-frame predictor 124 derives a motion vector predictor candidate for a motion vector of the current block by using neighboring blocks of the current block. Figure 4 All or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current picture shown are used as neighboring blocks for deriving motion vector prediction value candidates. In addition, blocks other than the current picture where the current block is located, which are located in a reference picture (which may be the same as or different from the reference picture used to predict the current block), may also be used as neighboring blocks for deriving motion vector prediction value candidates. For example, a block located at the same position as the current block in the reference picture or a block adjacent to a block located at the same position may be used. If the number of motion vector candidates selected by the method described above is less than the preset number, a zero vector is added to the motion vector candidates.
[0065] The inter predictor 124 derives motion vector predictor candidates by using motion vectors of neighboring blocks, and determines a motion vector predictor for a motion vector of a current block by using the motion vector predictor candidates. In addition, a motion vector difference is calculated by subtracting the motion vector predictor from the motion vector of the current block.
[0066] A motion vector prediction value can be obtained by applying a predefined function (e.g., median and average operations, etc.) to motion vector prediction value candidates. In this case, the video decoding device also knows the predefined function. Additionally, since the neighboring blocks used to derive the motion vector prediction value candidates are blocks that have already been encoded and decoded, the video decoding device may also already know the motion vectors of the neighboring blocks. Therefore, the video encoding device does not need to encode the information for identifying the motion vector prediction value candidates. Thus, in this case, the information about the motion vector difference and the information about the reference picture used to predict the current block are encoded.
[0067] Meanwhile, the motion vector prediction value can also be determined by a scheme of selecting any one of the motion vector prediction value candidates. In this case, the information for identifying the selected motion vector prediction value candidate is additionally encoded jointly with the information about the motion vector difference and the information about the reference picture used to predict the current block.
[0068] The subtractor 130 generates a residual block by subtracting the predicted block generated by the intra predictor 122 or the inter predictor 124 from the current block.
[0069] The transformer 140 converts the residual signal in the residual block having pixel values in the spatial domain into transform coefficients in the frequency domain. The transformer 140 can perform the transformation on the residual signal in the residual block by using the total size of the residual block as the transform unit, or can also divide the residual block into multiple sub - blocks and perform the transformation by using these sub - blocks as the transform unit. Alternatively, the residual block is divided into two sub - blocks, a transform region and a non - transform region, to perform the transformation on the residual signal by using only the transform region sub - block as the transform unit. Here, the transform region sub - block can be one of two rectangular blocks having a size ratio of 1:1 based on the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_flag) indicates that only the sub - block is transformed, and the direction (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or the position information (cu_sbt_pos_flag) are encoded by the entropy encoder 155 and signaled to the video decoding device. Additionally, the size of the transform region sub - block can have a size ratio of 1:3 based on the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_quad_flag) distinguishing the corresponding segmentation is additionally encoded by the entropy encoder 155 and signaled to the video decoding device.
[0070] Meanwhile, the transformer 140 can perform transformations on the residual blocks separately in the horizontal and vertical directions. For the transformation, various types of transformation functions or transformation matrices can be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation can be defined as a multiple transformation set (MTS). The transformer 140 can select a pair of transformation functions in the MTS with the highest transformation efficiency, and can transform the residual blocks in each of the horizontal and vertical directions. Information (mts_idx) about the pair of transformation functions in the MTS is encoded by the entropy encoder 155 and signaled to the video decoding device.
[0071] The quantizer 145 quantizes the transform coefficients output from the transformer 140 using quantization parameters, and outputs the quantized transform coefficients to the entropy encoder 155. The quantizer 145 can also immediately quantize the relevant residual blocks without transformation for any block or frame. The quantizer 145 can also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients in the transform block. The quantization matrix applied to the quantized transform coefficients arranged in a two-dimensional layout can be encoded and signaled to the video decoding device.
[0072] The rearrangement unit 150 can perform rearrangement of coefficient values for the quantized residual values.
[0073] The rearrangement unit 150 can change the 2D coefficient array into a 1D coefficient sequence by using coefficient scanning. For example, the rearrangement unit 150 can scan the DC coefficient to the high-frequency domain coefficients by using a zig-zag scan or a diagonal scan to output the 1D coefficient sequence. According to the size of the transform unit and the intra prediction mode, vertical scanning that scans the 2D coefficient array along the column direction and horizontal scanning that scans the 2D block type coefficients along the row direction can also be used instead of the zig-zag scan. In other words, according to the size of the transform unit and the intra prediction mode, the scanning method to be used can be determined among the zig-zag scan, diagonal scan, vertical scan, and horizontal scan.
[0074] The entropy encoder 155 encodes the sequence of 1D quantized transform coefficients output from the rearrangement unit 150 by using various coding schemes including context-based adaptive binary arithmetic coding (CABAC), exponential Golomb, etc. to generate a bitstream.
[0075] In addition, the entropy encoder 155 encodes information related to block partitioning (such as CTU size, CTU partition flag, QT partition flag, MTT partition type, MTT partition direction, etc.) to allow a video decoding device to partition blocks equivalently to a video encoding device. In addition, the entropy encoder 155 encodes information regarding a prediction type indicating whether the current block is encoded by intra prediction or inter prediction. The entropy encoder 155 encodes intra prediction information (i.e., information regarding an intra prediction mode) or inter prediction information (in the case of the merge mode, merge index, and in the case of the AMVP mode, information regarding a reference picture index and a motion vector difference) according to the prediction type. In addition, the entropy encoder 155 encodes information related to quantization (i.e., information regarding a quantization parameter and information regarding a quantization matrix).
[0076] The inverse quantizer 160 inverse quantizes the quantized transform coefficients output from the quantizer 145 to generate transform coefficients. The inverse transformer 165 transforms the transform coefficients output from the inverse quantizer 160 from the frequency domain to the spatial domain to reconstruct a residual block.
[0077] The adder 170 adds the reconstructed residual block and the prediction block generated by the predictor 120 to reconstruct the current block. When performing intra prediction on the next sequence block, pixels in the reconstructed current block can be used as reference pixels.
[0078] The loop filter unit 180 performs filtering on the reconstructed pixels in order to reduce block effects, ringing effects, blurring effects, etc. that occur due to block-based prediction and transform / quantization. The loop filter unit 180 as a loop filter may include all or some of a deblocking filter 182, a sample adaptive offset (SAO) filter 184, and an adaptive loop filter (ALF) 186.
[0079] The deblocking filter 182 filters the boundaries between the reconstructed blocks in order to remove block effects that occur due to block unit encoding / decoding, and the SAO filter 184 and the ALF 186 perform additional filtering on the deblocked filtered video. The SAO filter 184 and the ALF 186 are filters for compensating for the difference between the reconstructed pixels and the original pixels that occurs due to lossy encoding. The SAO filter 184 applies an offset in units of CTUs to enhance subjective image quality and coding efficiency. On the other hand, the ALF 186 performs block unit filtering and applies different filters by distinguishing the boundaries of the corresponding blocks and the degree of change amount to compensate for distortion. Information regarding filter coefficients to be used for the ALF may be encoded and signaled to a video decoding device.
[0080] The reconstructed blocks filtered by the deblocking filter 182, the SAO filter 184, and the ALF 186 are stored in the memory 190. When all the blocks in a picture are reconstructed, the reconstructed picture can be used as a reference picture for inter prediction of blocks in pictures to be encoded later.
[0081] The video encoding device may store the bitstream of the encoded video data in a non - transitory storage medium or transmit the bitstream to a video decoding device via a communication network.
[0082] Figure 5 is a functional block diagram of a video decoding device that can implement the technology of the present disclosure. Hereinafter, with reference to Figure 5 , the video decoding device and the components of the device will be described.
[0083] The video decoding device may include an entropy decoder 510, a rearrangement unit 515, an inverse quantizer 520, an inverse transformer 530, a predictor 540, an adder 550, a loop filter unit 560, and a memory 570.
[0084] Similar to Figure 1 the video encoding device, each component of the video decoding device may be implemented as hardware or software or implemented as a combination of hardware and software. In addition, the functions of each component may be implemented as software, and the microprocessor may also be implemented to execute the functions of the software corresponding to each component.
[0085] The entropy decoder 510 decodes the bitstream generated by the video encoding device to extract information related to block partitioning to determine the current block to be decoded, and extracts the prediction information and information about the residual signal required for reconstructing the current block.
[0086] The entropy decoder 510 determines the size of the CTU by extracting information about the CTU size from the sequence parameter set (SPS) or the picture parameter set (PPS), and divides the picture into CTUs with the determined size. Additionally, the CTU is determined as the highest layer of the tree structure, i.e., the root node, and the partitioning information of the CTU can be extracted to partition the CTU using the tree structure.
[0087] For example, when partitioning the CTU using the QTBTTT structure, first, the first flag (QT_split_flag) related to the partitioning of the QT is extracted to divide each node into four lower - layer nodes. Additionally, for the nodes corresponding to the leaf nodes of the QT, the second flag (mtt_split_flag) related to the partitioning of the MTT, the partitioning direction (vertical / horizontal), and / or the partitioning type (binary / ternary) are extracted to divide the corresponding leaf nodes into the MTT structure. As a result, each node below the leaf nodes of the QT is recursively divided into the BT or TT structure.
[0088] As another example, when a CTU is split by using the QTBTTT structure, a CU split flag (split_cu_flag) indicating whether the CU is split is extracted. When the corresponding block is split, a first flag (QT_split_flag) can also be extracted. During the splitting process, for each node, zero or more recursive MTT splits can occur after zero or more recursive QT splits. For example, for a CTU, the MTT split can occur immediately, or conversely, only multiple QT splits may occur.
[0089] As another example, when a CTU is split by using the QTBT structure, a first flag (QT_split_flag) related to the split of the QT is extracted to split each node into four lower-layer nodes. In addition, a split flag (split_flag) indicating whether the node corresponding to the leaf node of the QT is further split into a BT and split direction information are extracted.
[0090] Meanwhile, when the entropy decoder 510 determines the current block to be decoded by using the splitting of the tree structure, the entropy decoder 510 extracts information about the prediction type indicating whether the current block is intra prediction or inter prediction. When the prediction type information indicates intra prediction, the entropy decoder 510 extracts the syntax element for the intra prediction information (intra prediction mode) of the current block. When the prediction type information indicates inter prediction, the entropy decoder 510 extracts the information representing the syntax element for the inter prediction information (i.e., the motion vector and the reference picture referenced by the motion vector).
[0091] In addition, the entropy decoder 510 extracts quantization-related information and extracts information about the quantized transform coefficients of the current block as information about the residual signal.
[0092] The rearrangement unit 515 can change the sequence of 1D quantized transform coefficients entropy decoded by the entropy decoder 510 back into a 2D coefficient array (i.e., a block) in an order opposite to the coefficient scan order performed by the video coding device.
[0093] The inverse quantizer 520 inverse quantizes the quantized transform coefficients and inverse quantizes the quantized transform coefficients by using the quantization parameter. The inverse quantizer 520 can also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in 2D. The inverse quantizer 520 can perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video coding device to the 2D array of quantized transform coefficients.
[0094] The inverse transformer 530 generates a residual block for the current block by reconstructing the residual signal by inverse-transforming the inverse quantized transform coefficients from the frequency domain to the spatial domain.
[0095] In addition, when the inverse transform unit 530 performs inverse transform on a partial region (sub-block) of the transform block, the inverse transform unit 530 extracts a flag (cu_sbt_flag) indicating that only the sub-block of the transform block is transformed, direction (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block. The inverse transform unit 530 also inverse-transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to reconstruct the residual signal, and fills the un-inverse-transformed region with the value "0" as the residual signal to generate a final residual block for the current block.
[0096] In addition, when applying MTS, the inverse transform unit 530 determines a transform function or a transform matrix to be applied in each of the horizontal direction and the vertical direction by using MTS information (mts_idx) signaled from the video coding device. The inverse transform unit 530 also performs inverse transform on the transform coefficients in the transform block in the horizontal direction and the vertical direction by using the determined transform function.
[0097] The predictor 540 may include an intra predictor 542 and an inter predictor 544. The intra predictor 542 is activated when the prediction type of the current block is intra prediction, and the inter predictor 544 is activated when the prediction type of the current block is inter prediction.
[0098] The intra predictor 542 determines an intra prediction mode of the current block among a plurality of intra prediction modes according to a syntax element of the intra prediction mode extracted from the entropy decoder 510. The intra predictor 542 also predicts the current block by using neighboring reference pixels of the current block according to the intra prediction mode.
[0099] The inter predictor 544 determines a motion vector of the current block and a reference picture referred to by the motion vector by using syntax elements for the inter prediction mode extracted from the entropy decoder 510, and predicts the current block by using the motion vector and the reference picture.
[0100] The adder 550 reconstructs the current block by adding the residual block output from the inverse transform unit 530 and the prediction block output from the inter predictor 544 or the intra predictor 542. Pixels in the reconstructed current block are used as reference pixels for intra prediction of blocks to be decoded later.
[0101] The loop filter unit 560 serving as a loop filter may include a deblocking filter 562, a SAO filter 564, and an ALF 566. The deblocking filter 562 performs deblocking filtering on the boundaries between the reconstructed blocks to remove the blocking artifacts that occur due to block-based unit decoding. The SAO filter 564 and the ALF 566 perform additional filtering on the reconstructed blocks after deblocking filtering to compensate for the difference between the reconstructed pixels and the original pixels that occurs due to lossy coding. The filter coefficients of the ALF are determined by using the information about the filter coefficients decoded from the bitstream.
[0102] The reconstructed blocks filtered by the deblocking filter 562, the SAO filter 564, and the ALF 566 are stored in the memory 570. When all the blocks in a picture are reconstructed, the reconstructed picture can be used as a reference picture for inter prediction of the blocks within the pictures to be encoded later.
[0103] In some embodiments, the present disclosure relates to encoding and decoding video images as described above. More specifically, the present disclosure provides a video coding method and apparatus for intra prediction of a current chrominance block by using geometric partitioning when predicting a chrominance component after predicting and reconstructing a luminance component for a current block.
[0104] The following embodiments may be performed by the intra predictor 122 in a video coding apparatus. The following embodiments may also be performed by the intra predictor 542 in a video decoding apparatus.
[0105] When encoding a current block, the video coding apparatus may generate signaling information associated with the present embodiment in terms of optimizing the bitrate distortion. The video coding apparatus may encode the signaling information by using the entropy encoder 155 and transmit the encoded signaling information to the video decoding apparatus. The video decoding apparatus may decode the signaling information associated with the decoding of the current block from the bitstream by using the entropy decoder 510.
[0106] In the following description, the term "target block" may be used interchangeably with the current block or the coding unit (CU), or may refer to some regions of the coding unit.
[0107] In addition, the value of a flag being true indicates when the flag is set to 1. Additionally, the value of a flag being false indicates when the flag is set to 0.
[0108] I. Intra Prediction of Chrominance Component
[0109] Several techniques are introduced to improve the coding efficiency by using intra prediction.
[0110] In the Versatile Video Coding (VVC) technology, the luma block has 65 intra prediction modes with subdivided directional modes, i.e., modes 2 to 66 except for the non-directional modes (i.e., the planar mode and the DC mode), as Figure 3a shown. The 65 directional modes, the planar mode, and the DC mode are collectively referred to as 67 IPMs.
[0111] Depending on the prediction direction utilized by the luma block, the chroma block may also have restricted access to the intra prediction of these subdivided directional modes. However, except for the horizontal and vertical directions, the intra prediction of the chroma block may not always utilize the various directional modes available for the luma block. To be able to use these different directional modes, the prediction mode of the current chroma block needs to be set to the direct mode (DM). By setting the prediction mode to DM, the current chroma block can utilize the directional modes other than the horizontal and vertical modes of the luma block.
[0112] When encoding the chroma block, the intra prediction modes most frequently used or defaulted to maintain video quality include the planar mode, the DC mode, the vertical mode, the horizontal mode, and DM. In DM, the intra prediction mode of the luma block spatially corresponding to the current chroma block is used as the intra prediction mode of the chroma block.
[0113] The video coding device may signal to the video decoding device whether the intra prediction mode of the chroma block is DM. The video coding device may convey DM to the video decoding device in various ways. For example, the video coding device may indicate whether the intra prediction mode of the chroma block is DM by setting intra_chroma_pred_mode, which is the information for indicating the intra prediction mode of the chroma block, to a specific value and transmitting it to the video decoding device.
[0114] When encoding the chroma block in the intra prediction mode, the video coding device may set the intra prediction mode IntraPredModeC of the chroma block according to Table 1.
[0115] Hereinafter, to distinguish the information intra_chroma_pred_mode and IntraPredModeC related to the intra prediction mode of the chroma block, they are respectively represented as the chroma intra prediction mode indicator and the chroma intra prediction mode.
[0116] [Table 1]
[0117]
[0118] Here, lumaIntraPredMode is the intra prediction mode of the luma block corresponding to the current chroma block (hereinafter referred to as the "luma intra prediction mode"). lumaIntraPredMode representsFigure 3a One of the prediction modes shown. For example, in Table 1, lumaIntraPredMode = 0 refers to the planar prediction mode, and lumaIntraPredMode = 1 refers to the DC prediction mode. The cases where lumaIntraPredMode is 18, 50, and 66 indicate the directional modes called horizontal, vertical, and VDIA, respectively. On the other hand, the cases where intra_chroma_pred_mode = 0, 1, 2, and 3 indicate the planar, vertical, horizontal, and DC prediction modes, respectively. The case where intra_chroma_pred_mode = 4 is DM, in which the value of IntraPredModeC (chroma intra prediction mode) is set to be equal to the value of lumaIntraPredMode.
[0119] Meanwhile, the process of parsing the intra prediction mode of a chroma block performed by the video decoding device is shown in Table 2.
[0120] [Table 2]
[0121]
[0122] The video decoding device parses the cclm_mode_flag indicating whether to use the cross-component linear model (CCLM) mode. If cclm_mode_flag is 1 to enable the CCLM mode, the video decoding device parses the cclm_mode_idx indicating the CCLM mode. Depending on the value of cclm_mode_idx, the CCLM mode can indicate one of three modes. On the other hand, if cclm_mode_flag is 0 indicating that the CCLM mode is not used, the video decoding device parses the intra_chroma_pred_mode indicating the intra prediction mode as described above.
[0123] If the CCLM mode is applied to the intra prediction of the current chroma block, the video decoding device determines the corresponding region in the luminance image corresponding to the current chroma block (hereinafter, "corresponding luminance region"). For the prediction of the linear model, the left reference pixel and the upper reference pixel of the corresponding luminance region and the left reference pixel and the upper reference pixel of the target chroma block can be utilized. Hereinafter, the left reference pixel and the upper reference pixel are generally referred to as reference pixels, neighboring pixels, or adjacent pixels. In addition, the reference pixels in the chroma channel are called chroma reference pixels, and the reference pixels in the luminance channel are called luminance reference pixels.
[0124] In CCLM prediction, a prediction block serving as a predictor of a target chrominance block is generated by deriving a linear model between a reference pixel in a corresponding luminance region and a reference pixel of a chrominance block, and then applying the linear model to a reconstructed pixel in the corresponding luminance region. For example, four pairs of pixels (which are pixels in the peripheral pixel line of the current chrominance block and pixels in the combined corresponding luminance region) can be used to derive the linear model. In response to the four pairs of pixels, the video decoding device can derive α and β representing the linear model, as shown in Equation 1.
[0125] [Equation 1]
[0126] β = Y a -α·X a
[0127] Here, X a and X b represent the average of the minimum value and the second minimum value, and the average of the maximum value and the second maximum value of the corresponding luminance pixels in the four pairs of pixels, respectively. In addition, Y a and Y b represent the average of the minimum value and the second minimum value, and the average of the maximum value and the second maximum value of the chrominance pixels in the four pairs of pixels, respectively. Then, the video decoding device can use the linear model to generate a predictor pred L (i, j) of the current chrominance block from the pixel value rec' C (i, j) of the corresponding luminance region, as shown in Equation 2.
[0128] [Equation 2]
[0129] pred C (i,j) = α·rec′ L (i,j) + β
[0130] As described above, according to the positions of the neighboring pixels used in the derivation of the linear model, the CCLM mode is divided into three modes, CCLM_LT, CCLM_L, and CCLM_T. The CCLM_LT mode uses two pixels in each direction from the neighboring pixels adjacent to the left and upper sides of the current chrominance block. The CCLM_L mode uses four pixels from the neighboring pixels adjacent to the left side of the current chrominance block. Finally, the CCLM_T mode utilizes four pixels from the neighboring pixels adjacent to the upper side of the current chrominance block.
[0131] The following embodiments are described with respect to a video decoding device, but can be implemented in the same or similar manner in a video encoding device.
[0132] II. Embodiments According to the Present Disclosure
[0133] Figure 6 It is a block diagram showing a part of a video decoding device according to at least one embodiment of the present disclosure.
[0134] A video decoding device according to some embodiments can determine a prediction unit and a transform unit, and perform prediction and inverse transform with respect to a current block corresponding to the determined unit by using the determined prediction technique and prediction mode, so as to finally generate a reconstructed block of the current block. Figure 6 The operations shown can be performed by the inverse transformer 530, the predictor 540, and the adder 550 of the video decoding device. On the other hand, as Figure 6 shown, the same operations can be performed by the inverse transformer 165, the image splitter 110, the predictor 120, and the adder 170 of the video encoding device. In this case, the video decoding device uses the encoded information parsed from the bitstream, but the video encoding device can use the encoded information set from a higher layer in terms of minimizing bitrate distortion. Hereinafter, for ease of description, the embodiments are described centering on the video decoding device.
[0135] As Figure 5 shown, depending on the prediction technique, the predictor 540 includes an intra predictor 542 and an inter predictor 544, but as Figure 6 shown, the predictor 540 can include all or part of a prediction unit determiner 602, a prediction technique determiner 604, a prediction mode determiner 606, and a prediction executor 608.
[0136] When the color format of the input video is the YUV format (YUV420, YUV411, YUV422, YUV444, etc.), the video decoding device can perform prediction and reconstruction of the luminance component, and then can perform prediction and reconstruction of the chrominance component. In other words, the luminance component and the chrominance component can be Figure 6 reconstructed in the shown component order. On the other hand, when the color format of the input video is RGB, the video encoding device can perform a color format transformation from RGB to YUV, and then can encode the transformed video. Here, in the case of the YUV format, the color format represents the correspondence between the pixels in the luminance component and the pixels in the chrominance component.
[0137] The prediction unit determiner 602 determines a prediction unit (PU). The prediction technique determiner 604 determines the prediction technique with respect to the prediction unit, for example, intra prediction, inter prediction, or intra block copy (IBC) mode, palette mode, etc. The prediction mode determiner 606 determines the detailed prediction mode for the prediction technique. The prediction executor 608 generates a prediction block of the current block according to the determined prediction mode.
[0138] The inverse transformer 530 includes a transform unit determiner 610 and an inverse transform executor 612. The transform unit determiner 610 determines a transform unit (TU) for the inverse quantized signal with respect to the current block, and the inverse transform executor 612 performs an inverse transform on the transform unit represented by the inverse quantized signal to generate a residual signal.
[0139] An adder 550 adds the predicted block and the residual signal to generate a reconstructed block. The reconstructed block is stored in the memory and can be used to predict other future blocks.
[0140] The prediction unit determined by the prediction unit determiner 602 can be the current block or one of the sub-blocks divided from the current block. In this case, depending on the color format, the prediction unit of the chrominance component may correspond in size to the prediction unit of the luminance component. Alternatively, the prediction units of the luminance component and the chrominance component can be determined separately, and prediction can be performed for the prediction unit of the chrominance component.
[0141] The prediction technique determiner 604 determines the prediction technique for the prediction unit. As described above, the prediction technique can be one of inter prediction, intra prediction, IBC mode, and palette mode. In this case, the prediction technique of the chrominance component can be determined to be the same as that of the corresponding luminance component without signaling and parsing separate information.
[0142] The following describes the case where the prediction technique of the current chrominance block is intra prediction when the prediction mode determiner 606 determines the prediction mode of the current chrominance block and the prediction executor 608 predicts the current chrominance block.
[0143] In one example, the prediction mode determiner 606 can decode the intra prediction mode using the neighboring reconstructed chrominance samples of the current chrominance block, and then can determine the decoded prediction mode as the prediction mode of the current chrominance block. At this time, the prediction executor 608 can generate a predicted block of the current chrominance block by using the neighboring reconstructed chrominance samples according to the determined intra prediction mode.
[0144] As an example, the following describes the case where the luminance component and the chrominance component have the same block partitioning structure, and as Figure 7 shown, when predicting the corresponding luminance block according to the intra prediction mode, the corresponding luminance block is predicted according to the spatial geometric partitioning mode (SGPM). Here, the corresponding luminance block refers to the luminance block at the same position as the current chrominance block. SGPM applies a geometric partitioning mode to the intra-predicted luminance block. At this time, the divided luminance sub-regions can be predicted according to different intra prediction modes. The video coding device can signal a flag to the video decoding device to indicate whether SGPM is to be applied to the luminance block.
[0145] Refer toFigure 8a and Figure 8b For an example of Figure 8b , the prediction mode determiner 606 may determine the prediction mode of the current chrominance block as follows.
[0146] The prediction mode determiner 606 decodes a flag indicating whether SGPM is applied to the current chrominance block (hereinafter, "SGPM flag"). Based on the SGPM flag, the prediction mode determiner 606 may determine the prediction mode for the current chrominance block.
[0147] As Figure 8a shown, if the SGPM flag is false, the prediction mode determiner 606 may decode the prediction mode of the current chrominance block without dividing it into sub-regions. In this case, the decoded prediction mode of the current chrominance block may be one of prediction modes such as vertical, horizontal, DC, and planar modes, direct mode (DM), cross-component based prediction (CC-based prediction), etc. When the decoded prediction mode is DM, the prediction mode for the current chrominance block may be the prediction mode of a sub-region including the center position or the upper left position of the luminance block, or the prediction mode of a sub-region including the most luminance samples in the sub-region. The cross-component linear model (CCLM) as described above may be used as an example of the CC-based prediction for the current chrominance block.
[0148] If the SGPM flag is true, the prediction mode determiner 606 may determine that the prediction mode of the current chrominance block is SGPM, as Figure 8b shown. The prediction mode determiner 606 may divide the current chrominance block into sub-regions by using the sampling of the color format of the current picture and the geometric segmentation mode of the corresponding luminance block, and then may determine the prediction mode of each sub-region according to the following embodiments.
[0149] In one example, the prediction mode determiner 606 may decode the prediction mode of the sub-region. In this case, the prediction mode of the sub-region may be one of vertical, horizontal, DC, planar mode, and DM. When determined to be DM, the prediction mode used for the sub-region may be the prediction mode of the luminance sub-region corresponding to the chrominance sub-region (hereinafter, referred to as "corresponding luminance sub-region"). The corresponding luminance sub-region exists within the corresponding luminance block.
[0150] Alternatively, the prediction mode of the sub-region may be a prediction mode derived by using neighboring reconstructed samples of the current chrominance block. In this case, the direction information of the neighboring reconstructed samples may be used to derive the prediction mode.
[0151] Alternatively, the prediction mode of the sub-region can be determined to be CC-based prediction. In CC-based prediction, the current sub-region is predicted from the corresponding luma sub-region based on the relationship between the neighboring reconstructed samples of the current sub-region and the neighboring reconstructed samples of the corresponding luma sub-region.
[0152] At this time, CC-based prediction can be performed for each sub-region by signaling / parsing a separate cross-component prediction flag for each sub-region. Here, the cross-component prediction flag can indicate whether CC-based prediction is performed for the sub-region. For example, if the cross-component prediction flag is true, the prediction mode of the relevant sub-region can be determined to be CC-based prediction. On the other hand, if the cross-component prediction flag is false, the prediction mode of the relevant sub-region can be one of the vertical mode, horizontal mode, DC mode, planar mode, DM mode, and the prediction mode derived by using neighboring reconstructed samples.
[0153] As another example, by signaling / parsing the cross-component prediction flag for the current chrominance block, CC-based prediction can be performed for the sub-blocks within the current chrominance block. Here, the cross-component prediction flag can indicate whether CC-based prediction is to be performed for the sub-region. For example, if the cross-component prediction flag is true, the prediction mode of the sub-region can be determined to be CC-based prediction. On the other hand, if the cross-component prediction flag is false, the prediction mode of the sub-region can be one of the vertical mode, horizontal mode, DC mode, planar mode, DM, and the prediction mode derived by using neighboring reconstructed samples.
[0154] Alternatively, it can be determined whether to apply CC-based prediction without signaling / parsing separate information for each sub-region. If the sub-region is a region where CC-based prediction is not restricted, the prediction mode of the sub-region is determined to be CC-based prediction. On the other hand, if the sub-region is a region where CC-based prediction is restricted, the prediction mode of the sub-region can be one of the vertical mode, horizontal mode, DC mode, planar mode, DM, and the prediction mode derived by using neighboring reconstructed samples.
[0155] To derive the formula for CC-based prediction, the positions of the samples of the luma component and chroma component can be determined according to the geometric partitioning mode of the current chrominance block, as Figure 9 shown.
[0156] In one example, if the number of chroma reference samples for prediction in each sub-region is less than the threshold (t r ), then CC-based prediction can be restricted in that sub-region. In this case, the threshold (t r) It can be determined in advance by an agreement between a video encoding device and a video decoding device. Alternatively, the threshold can be determined implicitly based on the size of the current chrominance block and the size of the sub-region.
[0157] In addition, as described above, if the prediction based on CC is restricted by the number of reference samples, the prediction mode of the relevant sub-region can be determined according to DM without additionally signaling / parsing separate information.
[0158] As another example, the case where the luminance component and the chrominance component have different block partitioning structures and the corresponding luminance region is predicted according to the intra prediction mode is described below. Here, the corresponding luminance region indicates the luminance region at the same position as the current chrominance block. As Figures 10a to 10c shown, the prediction mode determiner 606 can determine the prediction mode of the current chrominance block based on the prediction mode of the luminance blocks in the corresponding luminance region as follows. Here, the corresponding luminance region indicates the luminance region at the corresponding position of the current chrominance block.
[0159] As an example, if the corresponding luminance region is intra-predicted and predicted according to SGPM, as Figure 10a shown, the prediction mode of the current chrominance block can be determined as shown in the examples of Figure 8a and Figure 8b .
[0160] As another example, the case where the corresponding luminance region includes luminance blocks predicted according to SGPM is described below, as Figure 10b shown. In this case, if a single luminance block is predicted according to SGPM, the geometric partitioning pattern of the block can be utilized. If multiple luminance blocks are predicted according to SGPM, the geometric partitioning pattern of the block with the largest size can be utilized. Alternatively, if the block predicted by SGPM to contain the samples at the center of the corresponding luminance region or the samples adjacent to the upper left boundary is used, the geometric partitioning pattern of the block can be used. When the SGPM flag is true, the following embodiments can be applied.
[0161] The prediction mode determiner 606 can apply the block partitioning information to the current chrominance block according to the sampling of the color format based on the geometric partitioning pattern of the block. As Figure 11 shown, the prediction mode of the current chrominance block can be determined to perform SGPM for predicting each sub-region.
[0162] In one example, the prediction mode determiner 606 can decode the prediction mode of the sub-region. In this case, the prediction mode of the sub-region can be one of a vertical mode, a horizontal mode, a DC mode, a planar mode, and DM. If the prediction mode of the sub-region is determined to be DM, the prediction mode of the corresponding luminance sub-region can be used as the prediction mode of the sub-region. The corresponding luminance sub-region exists within the corresponding luminance block.
[0163] Alternatively, the prediction mode of the sub-region can be derived by using neighboring reconstruction samples of the current chrominance block. In this case, the direction information of the neighboring reconstruction samples can be used to derive the prediction mode.
[0164] Alternatively, the prediction mode of the sub-region can be determined as CC-based prediction. In CC-based prediction, the current sub-region is predicted from the corresponding luma sub-region based on the relationship between the neighboring reconstruction samples of the current sub-region and the neighboring reconstruction samples of the corresponding luma sub-region.
[0165] In this case, CC-based prediction can be performed for each sub-region by signaling / parsing a separate cross-component prediction flag for each sub-region. Here, the cross-component prediction flag can indicate whether CC-based prediction is to be performed for a given sub-region. For example, if the cross-component prediction flag is true, the prediction mode of the relevant sub-region can be determined as CC-based prediction. On the other hand, if the cross-component prediction flag is false, the prediction mode of the relevant sub-region can be one of the vertical mode, horizontal mode, DC mode, planar mode, DM, and the prediction mode derived by using neighboring reconstruction samples.
[0166] As another example, by signaling / parsing the cross-component prediction flag for the current chrominance block, CC-based prediction can be performed for sub-blocks within the current chrominance block. Here, the cross-component prediction flag can indicate whether CC-based prediction is to be performed for the sub-region. For example, if the cross-component prediction flag is true, the prediction mode of the sub-region can be determined as CC-based prediction. On the other hand, if the cross-component prediction flag is false, the prediction mode of the sub-region can be one of the vertical mode, horizontal mode, DC mode, planar mode, DM, and the prediction mode derived by using neighboring reconstruction samples.
[0167] Alternatively, it can be determined whether to apply CC-based prediction without signaling / parsing separate information for each sub-region. If the sub-region is a region where CC-based prediction is not restricted, the prediction mode of the sub-region is determined as CC-based prediction. On the other hand, if the sub-region is a region where CC-based prediction is restricted, the prediction mode of the sub-region can be the vertical mode, horizontal mode, DC mode, planar mode, DM, and the prediction mode derived by using neighboring reconstruction samples.
[0168] To derive the formula for CC-based prediction, the positions of the samples of the luma component and chrominance component can be determined according to the geometric partitioning pattern of the current chrominance block, as Figure 12 shown.
[0169] In one example, if the number of chrominance reference samples used for prediction in each sub-region is less than a threshold (t r ), then the prediction based on CC can be restricted in that sub-region. In this case, the threshold (t r ) can be predetermined by agreement between the video encoding device and the video decoding device. Alternatively, the threshold can be implicitly determined based on the size of the current chrominance block and the size of the sub-region.
[0170] Furthermore, as described above, if the prediction based on CC is restricted based on the number of reference samples, the prediction mode of the relevant sub-region can be determined as the prediction mode according to DM without additionally signaling / parsing separate information.
[0171] As yet another example, the case where the corresponding luminance region is predicted and reconstructed while being divided into multiple blocks is described below, as Figure 10c shown. When the SGPM flag is true, the following embodiments can be applied. As Figure 13 shown, the prediction mode determiner 606 can apply the segmentation information of the corresponding luminance region to the current chrominance block according to the sampling of the color format. The prediction mode of the current chrominance block can be determined to perform SGPM for predicting each sub-region. In other words, the current chrominance block can be divided into sub-regions according to the segmentation information of the corresponding luminance region.
[0172] In one example, the prediction mode determiner 606 can decode the prediction mode of the sub-region. The prediction mode of the sub-region can be one of a vertical mode, a horizontal mode, a DC mode, a planar mode, and DM. If the prediction mode of the sub-region is determined to be DM, the prediction mode of the corresponding luminance sub-region can be used as the prediction mode of the sub-region. The corresponding luminance sub-region exists within the corresponding luminance block.
[0173] Alternatively, the prediction mode of the sub-region can be derived by using the neighboring reconstructed samples of the current chrominance block. In this case, in order to derive the prediction mode, the direction information of the sampled neighboring reconstruction can be used.
[0174] Alternatively, the prediction mode of the sub-region can be determined as a prediction based on CC. In the prediction based on CC, the current sub-region is predicted from the corresponding luminance sub-region based on the relationship between the neighboring reconstructed samples of the current sub-region and the neighboring reconstructed samples of the corresponding luminance sub-region.
[0175] At this time, for each sub-region, CC-based prediction can be performed by signaling / parsing a separate cross-component prediction flag. Here, the cross-component prediction flag can indicate whether CC-based prediction is to be performed for a given sub-region. For example, if the cross-component prediction flag is true, the prediction mode for the relevant sub-region can be determined as CC-based prediction. On the other hand, if the cross-component prediction flag is false, the prediction mode for the relevant sub-region can be a vertical mode, a horizontal mode, a DC mode, a planar mode, a DM, and a prediction mode derived by using neighboring reconstructed samples.
[0176] As another example, by signaling / parsing a cross-component prediction flag for the current chrominance block, CC-based prediction can be performed for sub-blocks within the current chrominance block. Here, the cross-component prediction flag can indicate whether CC-based prediction is to be performed for the sub-region. For example, if the cross-component prediction flag is true, the prediction mode for the sub-region can be determined as CC-based prediction. On the other hand, if the cross-component prediction flag is false, the prediction mode for the sub-region can be one of a vertical mode, a horizontal mode, a DC mode, a planar mode, a DM, and a prediction mode derived by using neighboring reconstructed samples.
[0177] Alternatively, it can be determined whether to apply CC-based prediction without signaling / parsing separate information for each sub-region. If the sub-region is a region where CC-based prediction is not restricted, the prediction mode for the sub-region is determined as CC-based prediction. On the other hand, if the sub-region is a region where CC-based prediction is restricted, the prediction mode for the sub-region can be one of a vertical mode, a horizontal mode, a DC mode, a planar mode, a DM, and a prediction mode derived by using neighboring reconstructed samples.
[0178] To derive a formula for CC-based prediction, the positions of samples of the luminance component and the chrominance component can be determined according to the geometric partitioning mode of the current chrominance block, as Figure 14 shown.
[0179] In one example, if the number of chrominance reference samples for prediction in each sub-region is less than a threshold (t r ), then CC-based prediction can be restricted in that sub-region. In this case, the threshold (t r ) can be predetermined by agreement between the video coding device and the video decoding device. Alternatively, the threshold can be implicitly determined based on the size of the current chrominance block and the size of the sub-region.
[0180] Furthermore, as described above, if CC-based prediction is restricted based on the number of reference samples, the prediction mode for the relevant sub-region can be determined according to DM without additional signaling / parsing of separate information.
[0181] On the other hand, if the current chrominance block is divided into multiple sub-regions, the prediction executor 608 predicts the sub-regions in sequence by using the prediction modes determined for each sub-region. After predicting the last sub-region, the prediction executor 608 determines the blending region. The prediction executor 608 can generate the final chrominance prediction block of the current chrominance block by weighted summing the sub-regions according to the blending region using weights.
[0182] The following describes the linear prediction relationship established for sub-regions, assuming that the prediction mode of the current chrominance block is determined as in the example of Figure 8b or Figure 11 , and one or more sub-regions are predicted based on CC according to the geometric partitioning pattern as in the examples of Figure 9 and Figure 12 ; or, assuming that the prediction mode of the current chrominance block is determined as in the example of Figure 13 , and one or more sub-regions are predicted based on CC according to the geometric partitioning pattern as in the example of Figure 14 . The video decoding device can use the samples in the determined reconstructed regions of each component of the sub-region to derive the parameters α and β representing the linear relationship between the components.
[0183] On the other hand, if, depending on the color format of the input video, the luminance component and the chrominance component have different resolutions, the video decoding device can downsample the region of the reconstructed luminance component and use the luminance samples in the downsampled region and the samples in the reconstructed region around the current chrominance block to model the linear relationship. The downsampling filter can be fixed, or the video encoding device can signal the coefficients of the downsampling filter to the video decoding device.
[0184] In one example, the video decoding device can derive the parameters by using all the samples in the determined reconstructed regions of each component. For example, after sorting the samples in the reconstructed region of the luminance component in descending order and sorting the samples in the reconstructed region of the chrominance component in descending order, the video decoding device can calculate the L corresponding to the average of the maximum value and the second maximum value and the average of the minimum value and the second minimum value for the corresponding component m-max 、L m-min 、C m-max and C m-min . Then, the video decoding device can derive α and β as shown in Equation 3.
[0185] [Equation 3]
[0186] β = L m-min - α · C m-min
[0187] As another example, the video decoding apparatus may derive the parameter by using only such samples corresponding to the predefined position based on the block size in the reconstruction region of each component. For example, after sorting the samples at the predetermined position in the reconstruction region of the luminance component in descending order and sorting the samples at the predetermined position in the reconstruction region of the chrominance component in descending order, the video decoding apparatus calculates the L of the corresponding component corresponding to the average of the maximum value and the second maximum value and the average of the minimum value and the second minimum value. m-max , L m-min , C m-max and C m-min Then, the video decoding device may derive α and β as shown in Formula 3.
[0188] As another example, the video decoding apparatus may adaptively use a down-sampling filter according to a prediction mode of a sub-region within a corresponding luminance block.
[0189] For example, Figure 15 As shown, a plurality of down-sampling filters may be predefined by agreement between a video encoding device and a video decoding device. The video decoding device may adaptively determine the down-sampling filter based on the prediction mode of a sub-region within a corresponding luminance block of a chrominance sub-region.
[0190] As another example, a downsampling filter corresponding to an intra prediction mode may be predefined by an agreement between a video encoding device and a video decoding device. The video decoding device may derive parameters for CC-based prediction by using a downsampling filter corresponding to a prediction mode of a sub-region within a corresponding luminance block of a chrominance sub-region.
[0191] Then, the video decoding apparatus may calculate the predicted sample Pred of the chrominance component by using the calculated parameters as shown in Formula 4. Chroma .
[0192] [Formula 4]
[0193] Pred Chroma =α·Rec′ L +β
[0194] In the prediction process, such as in Formula 4, Rec' L can be the sample value in the corresponding luminance block or the sample value in the downsampled corresponding luminance block, and Pred Chroma It can be the sample value in the current chrominance block. In the above parameter derivation process, Rec' L It can be the neighboring reconstructed sample value of the corresponding luminance block or the downsampled neighboring reconstructed sample value of the corresponding luminance block, and Pred Chroma It can be the neighboring reconstructed sample value of the chrominance component.
[0195] The following describes another linear prediction relationship established for sub-regions, assuming that the prediction mode of the current chrominance block is determined as in the example of Figure 8b or Figure 11 and one or more sub-regions are predicted based on CC according to the geometric segmentation mode as in the example of Figure 9 and Figure 12 ; or, assuming that the prediction mode of the current chrominance block is determined as in the example of Figure 13 and one or more sub-regions are predicted based on CC according to the geometric segmentation mode as in the example of Figure 14 . The video decoding device can use the samples in the determined reconstruction regions of each component of the relevant sub-regions to derive the parameters α T (T = 0, 1, 2) and β. The video decoding device can then calculate the predicted samples Pred of the chrominance component as shown in Equation 5 or Equation 6 Chroma .
[0196] [Equation 5]
[0197] Pred Chroma = α0·Rec′ L + α1·Rec″ L + β
[0198] [Equation 6]
[0199] Pred Chroma = α0·Rec′ L + α1·Rec″ L + α2·((Rec″′ L + midVal) >> bitDepth)+ β
[0200] In the prediction process such as Equation 5 or Equation 6, Rec L ', Rec L " and Rec L "' can be the sample values in the corresponding luma block or the sample values in the downsampled corresponding luma block, and Pred Chroma can be the sample values in the current chrominance block. During the parameter derivation process, Rec L ', Rec L " and Rec L "' can be the neighboring reconstruction sample values of the corresponding luma block or the neighboring reconstruction sample values of the downsampled corresponding luma block, and Pred Chroma can be the neighboring reconstruction sample values of the chrominance component. In addition, midVal can be 1 >> bitDepth, where bitDepth indicates the bit depth of the input video.
[0201] On the other hand, depending on the color format of the input video, if the luminance component and the chrominance component have different resolutions, the video decoding device may downsample the region of the reconstructed luminance component and model the linear relationship by using the luminance samples in the downsampled region and the samples in the neighboring reconstructed regions of the current chrominance block. The downsampling filter may be fixed, or the video encoding device may signal the coefficients of the downsampling filter to the video decoding device.
[0202] In one example, when performing the modeling by using Equation 5, Rec L ' and Rec L ” may be the luminance sample values reconstructed by different downsampling filters.
[0203] For example, Rec L ' may be obtained by the same downsampling filter based on a coding tree unit (CTU), slice, picture, or sequence. The downsampling filter for generating Rec L ” may be derived according to the prediction mode of the sub-region of the corresponding luminance block. As described in a later embodiment, the downsampling filter for generating Rec L ” may be derived by using a loss function. Alternatively, the downsampling filter for generating Rec L ” may be determined according to the signaling / parsing between the video encoding device and the video decoding device.
[0204] In addition, when performing the modeling by using Equation 6, Rec L ', Rec L ”, and Rec L ”' may be the luminance sample values reconstructed by different downsampling filters.
[0205] For example, Rec L ” may be obtained by the same downsampling filter based on a coding tree unit (CTU), slice, picture, or sequence. The downsampling filter for generating Rec L ” may be derived according to the prediction mode of the sub-region of the corresponding luminance block. As described in a later embodiment, the downsampling filter for generating Rec L ” may be derived by using a loss function. Alternatively, the downsampling filter for generating Rec L ” may be determined according to the signaling / parsing between the video encoding device and the video decoding device. RRec L ”' may be calculated by the same downsampling filter as the one used for calculating Rec L ' or Rec L ”.
[0206] Embodiments related to determining the hybrid region are described below.
[0207] After generating the prediction signals for each sub-region to generate the final prediction block of the current chrominance block, the video decoding device can explicitly or implicitly determine the final blending region to be blended during the weighted summation process. Next, the video decoding device can calculate the blending matrix to be used in the weighted summation process based on the determined final blending region.
[0208] In one example, the blending region can be determined as a fixed region based on the size of the current chrominance block and the derived geometric partitioning pattern.
[0209] Alternatively, the blending region can be explicitly determined between the video encoding device and the video decoding device by signaling / parsing the BlendingArea_idx indicating the blending region.
[0210] Alternatively, the blending region can be implicitly determined based on the process of determining the blending region. For example, after calculating the initial prediction signals for each sub-region, the video decoding device can implicitly determine the blending region by using the neighboring initial prediction values of the geometric partitioning boundary.
[0211] Meanwhile, the video decoding device can calculate the blending matrix W based on the implicitly or explicitly determined final blending region B . Each coefficient (W B (i, j)) in the blending matrix can be an integer value between 0 and 2 n (where n is an integer greater than or equal to 0). In this case, n can be determined based on the size of the current block and / or the geometric partitioning pattern.
[0212] The following describes embodiments related to weighted summation.
[0213] For example, if the current chrominance block includes two sub-regions, the video decoding device can perform weighted summation on the initial prediction signals P0 and P1 using the blending matrix W B to generate the final prediction block P G , as shown in Equation 7.
[0214] [Equation 7]
[0215] P G (i,j) = (W B (i,j) × P0 + (2 n - W B (i,j)) × P1 + (2 n-1 )) >> n
[0216] In Equation 7, each of the initial prediction signals P0 and P1 for the two sub-regions can be generated in the region of the current chrominance block by using, for example, the prediction mode of each sub-region.
[0217] In addition, if the current chrominance block includes two or more sub-regions, the video decoding device may use the mixing matrix W calculated during the process of determining the mixing region based on each geometric segmentation boundary B , to generate a prediction block (P G k , k = 0, 1, …, T−1) by performing a weighted sum of the initial prediction signals P0 and P1 on both sides, as shown in Equation 7. Here, T (a natural number) represents the number of geometric segmentation boundaries existing between two sub-regions, that is, the number of mixing matrices W B used to generate P G k The sequence can be a z-scan sequence using the top-left corner of the current chrominance block as a reference. The video decoding device may generate all prediction blocks P for two adjacent sub-regions for each geometric segmentation boundary G k , and then may sum the generated prediction blocks to generate a final prediction block, as shown in Equation 8.
[0218] [Equation 8]
[0219]
[0220] Hereinafter, a method for predicting the current chrominance block by using Figure 16 and Figure 17 is described. In this case, Figure 16 and Figure 17 assume that the SGPM flag is true.
[0221] Figure 16 is a flowchart of a method for predicting a current chrominance block by a video encoding device according to at least one embodiment of the present disclosure.
[0222] The video encoding device establishes a corresponding luminance region of the current chrominance block from a reconstructed region of the luminance component (S1600). Here, the corresponding luminance region represents a luminance block or a luminance region at the same position as the current chrominance block, and is an intra-prediction region.
[0223] The video encoding device determines a prediction mode of the current chrominance block based on whether SGPM is to be applied to the corresponding luminance region, the intra-prediction mode of the corresponding luminance region, and the block segmentation structure of the luminance component and the chrominance component (S1602).
[0224] The video encoding device generates a prediction block of the current chrominance block based on the prediction mode of the current chrominance block (S1602).
[0225] Meanwhile, when the luminance component and the chrominance component have the same block partitioning structure, and when predicting the corresponding luminance region according to the spatial geometry partitioning mode, steps S1602 and S1604 may include the following steps S1610 to S1618, and the video encoding device may also perform the steps after S1620.
[0226] The video encoding device divides the current chrominance block into chrominance sub-regions by using the color format of the current picture and the geometry partitioning mode of the corresponding luminance region (S1610).
[0227] The video encoding device determines the prediction mode of the chrominance sub-region as cross-component prediction (S1612).
[0228] The video encoding device generates a first prediction block of the chrominance sub-region based on cross-component prediction (S1614).
[0229] The video encoding device determines the prediction mode of the chrominance sub-region (S1616). For example, the video encoding device may determine the prediction mode of the chrominance sub-region as one of a vertical mode, a horizontal mode, DC, a planar mode, DM, and a prediction mode derived by using neighboring reconstructed samples of the current chrominance block. In this case, the video encoding device may determine the prediction mode of the chrominance sub-region in terms of optimizing the bitrate distortion.
[0230] The video encoding device generates a second prediction block of the chrominance sub-region according to the prediction mode of the chrominance sub-region (S1618).
[0231] The video encoding device determines a flag indicating whether to perform cross-component prediction on the chrominance sub-region based on the first prediction block and the second prediction block (S1620). In terms of optimizing the rate distortion, the video encoding device may determine the value of the flag based on the first prediction block and the second prediction block. For example, if the first prediction block is the best, the video encoding device may set the flag to true. On the other hand, if the second prediction block is the best, the video encoding device may set the flag to false.
[0232] The video encoding device encodes the determined flag (S1622).
[0233] In addition, if the flag is false, the video encoding device may encode the prediction mode of the chrominance sub-region determined in step S1616.
[0234] As another example, when the luminance component and the chrominance component have different block partitioning structures, and when the corresponding luminance region includes blocks predicted according to SGPM, step S1610 may be replaced as follows.
[0235] The video encoding device may divide a current chrominance block into chrominance sub-regions by using the color format of the current picture and the geometric partitioning pattern of a predicted block according to SGPM. The video encoding device may then perform steps S1612 to S1622.
[0236] As another example, when the luminance component and the chrominance component have different block partitioning structures and when the corresponding luminance region is divided into a plurality of blocks, step S1610 may be replaced as follows.
[0237] The video encoding device divides a current chrominance block into chrominance sub-regions by using the color format of the current picture and the partitioning information of the corresponding luminance region. The video encoding device may then perform steps S1612 to S1622.
[0238] Figure 17 is a flowchart of a method for a video decoding device to predict a current chrominance block according to at least one embodiment of the present disclosure.
[0239] The video decoding device establishes a corresponding luminance region of the current chrominance block from a reconstructed region of the luminance component (S1700). Here, the corresponding luminance region represents a luminance block or a luminance region at the same position as the current chrominance block and is an intra prediction region.
[0240] The video decoding device obtains a prediction mode of the current chrominance block based on whether SGPM is to be applied to the corresponding luminance region, the intra prediction mode of the corresponding luminance region, and the block partitioning structures of the luminance component and the chrominance component (S1702).
[0241] The video decoding device generates a predicted block of the current chrominance block based on the prediction mode of the current chrominance block (S1704).
[0242] In one example, when the luminance component and the chrominance component have the same block partitioning structure and when predicting the corresponding luminance region according to SGPM, step S1702 may include the following steps.
[0243] The video decoding device divides the current chrominance block into chrominance sub-regions by using the color format of the current picture and the geometric partitioning pattern of the corresponding luminance region (S1710).
[0244] The video decoding device decodes from the bitstream a flag indicating whether to perform cross-component based prediction (S1712).
[0245] The video decoding device checks the flag (S1714).
[0246] If the flag is true (yes in S1714), the video decoding device determines the prediction mode of the chrominance sub-region as cross-component based prediction (S1716).
[0247] On the other hand, if the flag is false (No in S1714), the video decoding device decodes the prediction mode of the chrominance sub-region (S1718). Here, the prediction mode of the chrominance sub-region can be one of a vertical mode, a horizontal mode, a DC mode, a planar mode, a direct mode (DM), and a prediction mode derived by using neighboring reconstructed samples of the current chrominance block.
[0248] As another example, when the luminance component and the chrominance component have different block partitioning structures, and when the corresponding luminance region includes blocks predicted according to SGPM, step S1710 can be replaced as follows.
[0249] The video decoding device divides the current chrominance block into chrominance sub-regions by using the color format of the current picture and the geometric partitioning mode of the predicted blocks according to SGPM. Then, the video decoding device can obtain the prediction mode of each chrominance sub-region according to steps S1712 to S1718.
[0250] As another example, when the luminance component and the chrominance component have different block partitioning structures, and when the corresponding luminance region is divided into multiple blocks, step S1710 can be replaced as follows.
[0251] The video decoding device divides the current chrominance block into chrominance sub-regions by using the color format of the current picture and the partitioning information of the corresponding luminance region. Then, the video decoding device can obtain the prediction mode of each chrominance sub-region according to steps S1712 to S1718.
[0252] Although the steps in the corresponding flowcharts are described as being executed sequentially, these steps merely illustrate the technical concepts of some embodiments of the present disclosure. Therefore, those of ordinary skill in the art to which the present disclosure pertains can execute the steps by changing the order described in the corresponding drawings or by executing two or more steps in parallel. Therefore, the steps in the corresponding flowcharts are not limited to the shown time series order.
[0253] It should be understood that the above description presents illustrative embodiments that can be implemented in various other ways. The functions described in some embodiments can be implemented by hardware, software, firmware, and / or combinations thereof. It should also be understood that the functional components described in the present disclosure are marked with “… unit” to strongly emphasize the possibility of their independent implementation.
[0254] Meanwhile, the various methods or functions described in some embodiments can be implemented as instructions stored in a non-transitory recording medium that can be read and executed by one or more processors. For example, the non-transitory recording medium can include various types of recording devices, where data is stored in a form readable by a computer system. For example, the non-transitory recording medium can include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical disk drives, magnetic hard disk drives, and solid-state drives (SSD), etc.
[0255] Although the embodiments of the present disclosure have been described for illustrative purposes, those of ordinary skill in the art to which the present disclosure pertains should recognize that various modifications, additions, and substitutions are possible without departing from the concept and scope of the present disclosure. Therefore, for the sake of brevity and clarity, the embodiments of the present disclosure have been described. The scope of the technical concept of the embodiments of the present disclosure is not limited by the illustrations. Therefore, those of ordinary skill in the art to which the present disclosure pertains should understand that the scope of the present disclosure should not be limited by the embodiments explicitly described above, but by the claims and their equivalents.
[0256] (Reference numeral)
[0257] 122: Intra predictor
[0258] 542: Intra predictor
[0259] 602: Prediction unit determiner
[0260] 604: Prediction technique determiner
[0261] 606: Prediction mode determiner
[0262] 608: Prediction executor.
[0263] Cross-reference to related applications
[0264] This application claims the priority and benefits of Korean Patent Application No. 10-2022-0156673, filed on November 21, 2022, and Korean Patent Application No. 10-2023-0116352, filed on September 1, 2023, the entire contents of each of which are incorporated herein by reference.
Claims
1. A method for reconstructing a current chrominance block by a video decoding device, the method comprising: establishing a corresponding luminance region of the current chrominance block from a reconstructed region of a luminance component, the corresponding luminance region being an intra prediction region and representing a luminance block or luminance region located at the same position as the current chrominance block; obtaining a prediction mode of the current chrominance block based on whether a spatial geometric segmentation mode (spatial geometric segmentation mode) is applied to the corresponding luminance region, an intra prediction mode of the corresponding luminance region, and a block segmentation structure of the luminance component and the chrominance component; and generating a prediction block of the current chrominance block based on the prediction mode of the current chrominance block.
2. The method according to claim 1, wherein When the block segmentation structures of the luminance component and the chrominance component are the same and the corresponding luminance region is predicted according to the spatial geometric segmentation mode, obtaining the prediction mode of the current chrominance block includes: segmenting the current chrominance block into chrominance sub-regions using a color format of a current picture and a geometric segmentation mode of the corresponding luminance region; and obtaining a prediction mode of each of the chrominance sub-regions.
3. The method according to claim 2, wherein, Obtaining a prediction mode of each of the chrominance sub-regions includes: decoding a flag indicating whether to perform cross-component (cross-component) based prediction from a bitstream; and checking the flag, wherein the method further includes: when the flag is true: determining the prediction mode of each of the chrominance sub-regions as the cross-component based prediction.
4. The method according to claim 3, wherein Obtaining a prediction mode of each of the chrominance sub-regions includes: when the flag is false: decoding a prediction mode of each of the chrominance sub-regions, wherein a prediction mode of each of the chrominance sub-regions is one of vertical (vertical), horizontal (horizontal), DC, planar, DM (direct mode), and a prediction mode derived using neighboring reconstructed samples of the current chrominance block.
5. The method according to claim 4, wherein, Obtaining a prediction mode of each of the chrominance sub-regions includes: when the decoding mode is the DM: determining the prediction mode of each of the chrominance sub-regions as a prediction mode of a luminance sub-region corresponding to each of the chrominance sub-regions.
6. The method according to claim 2, wherein, Obtaining a prediction mode of each of the chrominance sub-regions includes: determining whether to enable cross-component based prediction based on a count of chrominance reference samples used for prediction of each of the chrominance sub-regions, wherein the method further includes: when the cross-component based prediction is enabled: determining the prediction mode of each of the chrominance sub-regions as the cross-component based prediction.
7. The method according to claim 1, wherein, When the block segmentation structures of the luminance component and the chrominance component are different and the corresponding luminance region includes blocks predicted according to the spatial geometric segmentation mode, obtaining the prediction mode of the current chrominance block includes: segmenting the current chrominance block into chrominance sub-regions using a color format of a current picture and a geometric segmentation mode of the blocks predicted according to the spatial geometric segmentation mode; and obtaining a prediction mode of each of the chrominance sub-regions.
8. The method according to claim 1, wherein, When the block segmentation structures of the luminance component and the chrominance component are different and the corresponding luminance region is segmented into a plurality of blocks, obtaining the prediction mode of the current chrominance block includes: Partition the current chrominance block into chrominance sub-regions using the color format of the current picture and the segmentation information of the corresponding luminance region; and Obtain the prediction mode for each of the chrominance sub-regions.
9. The method according to claim 3, wherein Generating a prediction block for the current chrominance block includes: when the prediction mode for each of the chrominance sub-regions is the cross-component based prediction: Use samples in the reconstructed regions of the luminance component and the chrominance component to calculate parameters representing the linear relationship between the luminance component and the chrominance component; and Generate a prediction block for each of the chrominance sub-regions from samples in the corresponding luminance region based on the parameters.
10. The method according to claim 2, wherein, Generating a prediction block for the current chrominance block includes: when the current chrominance block includes two sub-regions: Generate initial prediction blocks for the two sub-regions respectively; and Perform weighted summation on the initial prediction blocks of the two sub-regions using a mixing matrix.
11. A method for encoding a current chrominance block by a video encoding device, the method includes: Establish a corresponding luminance region of the current chrominance block from the reconstructed region of the luminance component, the corresponding luminance region being an intra prediction region and representing a luminance block or luminance region at the same position as the current chrominance block; Determine the prediction mode of the current chrominance block based on whether a spatial geometric segmentation mode (spatial geometric segmentation mode) is applied to the corresponding luminance region, the intra prediction mode of the corresponding luminance region, and the block segmentation structures of the luminance component and the chrominance component; And Generate a prediction block for the current chrominance block based on the prediction mode of the current chrominance block.
12. The method according to claim 11, wherein, When the block segmentation structures of the luminance component and the chrominance component are the same and the corresponding luminance region is predicted according to the spatial geometric segmentation mode, determining the prediction mode of the current chrominance block includes: Partition the current chrominance block into chrominance sub-regions using the color format of the current picture and the geometric segmentation mode of the corresponding luminance region; and Determine the prediction mode for each of the chrominance sub-regions.
13. The method according to claim 12, wherein, Determining the prediction mode for each of the chrominance sub-regions includes: Determine the prediction mode for each of the chrominance sub-regions as cross-component (cross-component) based prediction, and Wherein, generating a prediction block for the current chrominance block includes: Generate a first prediction block for each of the chrominance sub-regions based on the cross-component based prediction.
14. The method according to claim 13, wherein, Determining the prediction mode for each of the chrominance sub-regions includes: Determine the prediction mode for each of the chrominance sub-regions as one of vertical, horizontal, DC, planar, DM, and a prediction mode derived using neighboring reconstructed samples of the current chrominance block, and Wherein, generating a prediction block for the current chrominance block includes: Generate a second prediction block for each of the chrominance sub-regions according to the prediction mode of each of the chrominance sub-regions.
15. The method according to claim 14, further includes: Based on the first prediction block and the second prediction block, determine a flag indicating whether to perform the cross-component based prediction for each of the chrominance sub-regions; And Encode the flag.
16. A computer-readable recording medium for storing a bitstream generated by a video encoding method, wherein, The video encoding method includes: Establish a corresponding luminance region of the current chrominance block from the reconstructed region of the luminance component, where the corresponding luminance region is an intra prediction region and represents a luminance block of the luminance region located at the same position as the current chrominance block or; Determine the prediction mode of the current chrominance block based on whether a spatial geometric segmentation mode (spatial geometric segmentation mode) is applied to the corresponding luminance region, the intra prediction mode of the corresponding luminance region, and the block segmentation structure of the luminance component and the chrominance component; and Generate a predicted block of the current chrominance block based on the prediction mode of the current chrominance block.
Citation Information
Patent Citations
Wrist architecture
KR1020220156673A
Encapsulation film
KR1020230116352A