Video coding method and apparatus for adaptively generating most probable mode (MPM) list
By changing the prediction mode of the basic block to the regular intra prediction mode, the MPM list is adaptively generated, which solves the problem of low MPM list generation efficiency in the non-rule intra mode, and improves the video encoding and decoding efficiency and quality.
Patent Information
- Application Number
- CN202380083576.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-24
- Filing Date
- 2023-10-27
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, when generating the most likely mode list, when the base block is in irregular intra prediction mode, it is impossible to efficiently generate an effective MPM list, resulting in inefficient video encoding and decoding and poor video quality.
By changing the prediction mode of the base block into a regular intra prediction mode, an MPM list is adaptively generated, and the MPM list is constructed using the changed base block.
Improve video encoding and decoding efficiency, enhance video quality, and reduce the number of bits in the transmission of non-regular intra prediction modes.
Smart Images

Figure CN120303928A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video encoding and decoding method and apparatus for adaptively generating a most probable mode list (MPM list). Background Art
[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Since video data has a large amount of data compared to audio data or still image data, video data requires a large amount of hardware resources (including memory) to store or transmit uncompressed video data.
[0004] Accordingly, an encoder is typically used to compress and store or transmit video data. A decoder receives the compressed video data, decompresses the received compressed video data, and plays the decompressed video data. Video compression technologies include H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC), and the Versatile Video Coding (VVC) has a coding and decoding efficiency improvement of approximately 30% or more compared to HEVC.
[0005] However, due to the gradual increase in image size, resolution, and frame rate, the amount of data to be encoded also increases. Accordingly, there is a need to provide a new compression technology with higher coding and decoding efficiency and improved image enhancement effects compared to existing compression technologies.
[0006] Intra prediction uses pixel information within a common picture to predict the pixel values of a current block to be encoded. With intra prediction, one of many intra prediction modes can be selected and used to predict the current block according to the characteristics of the video. The encoder selects and then uses one of many intra prediction modes to encode the current block. Thereafter, the encoder can convey information about the mode to the decoder.
[0007] The HEVC technology uses a total of 35 intra prediction modes for intra prediction, including 33 angular modes with directionality and 2 non-angular modes without directionality. However, as the spatial resolution of the video increases from 720×480 to 2048×1024 or 8192×4096, the size of the prediction block unit becomes larger and larger, thus requiring the addition of more diverse intra prediction modes. As Figure 3a shown, the VVC technology uses 65 prediction modes further subdivided for intra prediction, thus allowing more types of prediction directions than the prior art.
[0008] On the other hand, for the luminance channel, the most probable mode (MPM) is used for the efficient encoding / decoding of 67 intra-prediction modes (IPMs) on a per-CU basis, including the planar and DC modes as non-angular prediction modes. Based on the characteristic similarity between the prediction modes of adjacent blocks when a block is encoded in an intra-prediction mode, the MPM technique selects 6 MPM candidates based on the modes utilized by the adjacent blocks of the current block, their adjacent modes, and statistically frequently used modes. The group of these 6 MPM candidates is referred to as the MPM list. If the intra-prediction mode of the current block is included in the MPM list, the encoder encodes the MPM index indicating the prediction mode in the MPM list. On the other hand, if the intra-prediction mode of the current block is not included in the MPM list, the encoder encodes the intra-prediction mode of the current block by utilizing the MPM remainder consisting of the IPMs excluding the 6 MPM candidates.
[0009] When generating the MPM list, the conventional MPM technique only utilizes the prediction modes of the blocks including the pixels adjacent to the lower left side of the current block and the blocks including the pixels adjacent to the upper right side, which are hereinafter referred to as "base blocks". In this case, if the prediction mode of the base block is an irregular intra-prediction mode or an irregular intra-frame mode, the encoder may not be able to finally select one of the prediction modes in the MPM list as the prediction mode of the current block. Therefore, in order to improve the video encoding / decoding efficiency and enhance the video quality, a method for efficiently generating the MPM list when the prediction mode of the base block is an irregular intra-frame mode is needed. Summary of the Invention
[0010] Technical Problem
[0011] The present invention is dedicated to providing a video encoding / decoding method and apparatus that, when generating the MPM list of the current block and the base block is in an irregular intra-prediction mode or an irregular intra-frame mode, changes the base block into a block with a regular intra-prediction mode and then adaptively generates the MPM list by utilizing the changed base block.
[0012] Technical Solution
[0013] At least one aspect of the present invention provides a method for a video decoding device to reconstruct a current block. The method includes obtaining a prediction mode of a base block of the current block. The base block is a block including pixels to the left of the pixel at the lower left of the current block or a block including pixels above the pixel at the upper right of the current block. The method further includes checking the prediction mode of the base block. When the prediction mode of the base block is a non-regular intra mode, the method further includes: obtaining a transformation method for the base block, transforming the base block into a block with a regular intra prediction mode according to the transformation method, and generating a most probable mode list (MPM list) for the current block by using the regular intra prediction mode of the transformed base block.
[0014] Another aspect of the present invention provides a method for a video encoding device to encode a current block. The method includes obtaining a prediction mode of a base block of the current block. The base block is a block including pixels to the left of the pixel at the lower left of the current block or a block including pixels above the pixel at the upper right of the current block. The method further includes checking the prediction mode of the base block. When the prediction mode of the base block is a non-regular intra mode, the method further includes: determining a transformation method for the base block, transforming the base block into a block with a regular intra prediction mode according to the transformation method, and generating a most probable mode list (MPM list) for the current block by using the regular intra prediction mode of the transformed base block.
[0015] Still another aspect of the present invention provides a computer-readable recording medium storing a bitstream generated by a video encoding method. The video encoding method includes obtaining a prediction mode of a base block of the current block. The base block is a block including pixels to the left of the pixel at the lower left of the current block or a block including pixels above the pixel at the upper right of the current block. The video encoding method further includes checking the prediction mode of the base block. When the prediction mode of the base block is a non-regular intra mode, the video encoding method further includes: determining a transformation method for the base block, transforming the base block into a block with a regular intra prediction mode according to the transformation method, and generating a most probable mode list (MPM list) for the current block by using the regular intra prediction mode of the transformed base block.
[0016] Beneficial Effects
[0017] As described above, the present invention provides a video encoding and decoding method and apparatus, which, when generating the MPM list of the current block and the base block is in a non-regular intra prediction mode or a non-regular intra mode, transforms the base block into a block with a regular intra prediction mode, and then adaptively generates the MPM list by using the transformed base block. Therefore, the video encoding and decoding method and apparatus improve the video encoding and decoding efficiency and enhance the video quality. Description of the Drawings
[0018] Figure 1It is a block diagram of a video encoding device that can implement the technology of the present invention.
[0019] Figure 2 It shows a method of partitioning a block by using a quadtree plus binary tree plus ternary tree (QTBTTT) structure.
[0020] Figure 3a and Figure 3b It shows a plurality of intra prediction modes including a wide-angle intra prediction mode.
[0021] Figure 4 It shows adjacent blocks of the current block.
[0022] Figure 5 It is a block diagram of a video decoding device that can implement the technology of the present invention.
[0023] Figure 6 It is a schematic diagram showing pixels used in the construction of the most probable mode list (MPM list).
[0024] Figure 7 It is a schematic diagram showing the generation of the MPM list when the mode of the block including pixel L is an irregular intra mode.
[0025] Figure 8 It is a schematic diagram showing the change of the base block when the mode of the block including pixel L is an irregular intra mode according to at least one embodiment of the present invention.
[0026] Figure 9 It is a schematic diagram showing the construction of the MPM list when the mode of the block including pixel A is an irregular intra mode.
[0027] Figure 10 It is a schematic diagram showing the change of the base block when the mode of the block including pixel A is an irregular intra mode according to at least one embodiment of the present invention.
[0028] Figures 11a to 11d It is a schematic diagram showing a predetermined adjacent pixel search order according to at least one embodiment of the present invention.
[0029] Figure 12 It is a schematic diagram showing the change of the base block by using the adjacent pixel search order according to at least one embodiment of the present invention.
[0030] Figure 13 It is a schematic diagram showing the change of the base block by using the adjacent pixel search order according to another embodiment of the present invention.
[0031] Figure 14 It is a schematic diagram showing the change of the base block by using the proximity length according to at least one embodiment of the present invention.
[0032] Figure 15 is a schematic diagram showing a change by using a base block of an adjacent length according to another embodiment of the present invention.
[0033] Figure 16 is a schematic diagram showing a change by using a base block of an area according to at least one embodiment of the present invention.
[0034] Figure 17 is a schematic diagram showing a change by using a base block of an area according to another embodiment of the present invention.
[0035] Figure 18 is a schematic diagram showing a change by using a base block of an aspect ratio according to at least one embodiment of the present invention.
[0036] Figure 19 is a schematic diagram showing a change by using a base block of an aspect ratio according to another embodiment of the present invention.
[0037] Figure 20 is a flowchart of a method for encoding a current block by a video encoding device according to at least one embodiment of the present invention.
[0038] Figure 21 is a flowchart of a method for reconstructing a current block by a video decoding device according to at least one embodiment of the present invention.
[0039] Figure 22 is a flowchart of a method for encoding a current block by a video encoding device according to another embodiment of the present invention.
[0040] Figure 23 is a flowchart of a method for reconstructing a current block by a video decoding device according to another embodiment of the present invention. Detailed Description
[0041] Hereinafter, some embodiments of the present invention will be described in detail with reference to the accompanying illustrative drawings. In the following description, the same reference numerals denote the same elements, although the elements are shown in different drawings. In addition, in the following description of some embodiments, when the detailed description of related known components and functions is considered to obscure the subject matter of the present invention, the detailed description of the related known components and functions may be omitted for clarity and conciseness.
[0042] Figure 1 is a block diagram of a video encoding device that can implement the technology of the present invention. Hereinafter, with reference to Figure 1 illustrations, the video encoding device and the components of the device will be described.
[0043] The encoding device may include: an image splitter 110, a predictor 120, a subtractor 130, a transformer 140, a quantizer 145, a rearrangement unit 150, an entropy encoder 155, an inverse quantizer 160, an inverse transformer 165, an adder 170, a loop filter unit 180, and a memory 190.
[0044] Each component of the encoding device may be implemented as hardware or software, or as a combination of hardware and software. Additionally, the functions of each component may be implemented as software, and the microprocessor may also be implemented to execute the functions of the software corresponding to each component.
[0045] A video consists of one or more sequences including multiple images. Each image is divided into multiple regions, and encoding is performed on each region. For example, an image is divided into one or more tiles or / and slices. Here, one or more tiles may be defined as a tile group. Each tile or / and slice is divided into one or more coding tree units (CTUs). Additionally, each CTU is divided into one or more coding units (CUs) through a tree structure. The information applied to each coding unit (CU) is encoded into the syntax of the CU, and the information commonly applied to the CUs included in one CTU is encoded into the syntax of the CTU. Additionally, the information commonly applied to all blocks in a slice is encoded into the syntax of the slice header, while the information applied to all blocks constituting one or more images is encoded into the Picture Parameter Set (PPS) or the picture header. Furthermore, the information commonly referred to by multiple images is encoded into the Sequence Parameter Set (SPS). Additionally, the information commonly referred to by one or more SPSs is encoded into the Video Parameter Set (VPS). Moreover, the information commonly applied to a tile or a tile group may also be encoded into the syntax of the tile or tile group header. The syntax included in the SPS, PPS, slice header, tile or tile group header may be referred to as high-level syntax.
[0046] The image splitter 110 determines the size of the coding tree unit (CTU). The information regarding the size of the CTU (CTU size) is encoded into the syntax of the SPS or PPS and is transmitted to the video decoding device.
[0047] The image splitter 110 divides each image constituting the video into multiple coding tree units (CTUs) of a predetermined size, and then recursively divides the CTUs by using a tree structure. The leaf nodes in the tree structure become coding units (CUs), which are the basic units for encoding.
[0048] The tree structure can be a quadtree (QT), where a higher node (or parent node) is divided into four lower nodes (or child nodes) of the same size. The tree structure can also be a binary tree (BT), where a higher node is divided into two lower nodes. The tree structure can also be a ternary tree (TT), where a higher node is divided into three lower nodes in a 1:2:1 ratio. The tree structure can also be a structure that mixes two or more of the QT structure, BT structure, and TT structure. For example, a quadtree plus binarytree (QTBT) structure can be used, or a quadtree plus binarytreeternarytree (QTBTTT) structure can be used. Here, the binarytreeternarytree (BTTT) is added to the tree structure to form a multiple-type tree (MTT).
[0049] Figure 2 is a schematic diagram for describing a method of dividing a block by using the QTBTTT structure.
[0050] As Figure 2 shown, the CTU can first be divided into a QT structure. The quadtree division can be recursive until the size of the divided block reaches the minimum block size (MinQTSize) of the leaf nodes allowed in the QT. The entropy encoder 155 encodes a first flag (QT_split_flag) indicating whether each node of the QT structure is divided into four lower nodes and signals it to the video decoding device. When the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in the BT, the leaf node can be further divided into at least one of the BT structure or the TT structure. There can be multiple division directions in the BT structure and / or the TT structure. For example, there can be two directions, namely, the direction of horizontally dividing the block of the corresponding node and the direction of vertically dividing the block of the corresponding node. As Figure 2 shown, when the MTT division starts, the entropy encoder 155 encodes a second flag (mtt_split_flag) indicating whether the node is divided, and a flag indicating the division direction (vertical or horizontal) and / or a flag indicating the division type (binary or ternary) in the case where the node is divided, and signals it to the video decoding device.
[0051] Alternatively, before encoding a first flag (QT_split_flag) indicating whether each node is split into four lower-layer nodes, a CU split flag (split_cu_flag) indicating whether a node is split may also be encoded. When the value of the CU split flag (split_cu_flag) indicates that each node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU, which is a basic unit of encoding. When the value of the CU split flag (split_cu_flag) indicates that each node is split, the video encoding device starts encoding the first flag first with the above scheme.
[0052] When QTBT is used as another example of a tree structure, there may be two types, that is, a type in which the block of the corresponding node is horizontally split into two blocks of the same size (i.e., symmetric horizontal split) and a type in which the block of the corresponding node is vertically split into two blocks of the same size (i.e., symmetric vertical split). The entropy encoder 155 encodes a split flag (split_flag) indicating whether each node of the BT structure is split into lower-layer blocks and split type information indicating the split type, and transmits it to the video decoding device. On the other hand, there may additionally be a type in which the block of the corresponding node is split into two asymmetric blocks. The asymmetric form may include a form in which the block of the corresponding node is split into two rectangular blocks with a size ratio of 1:3, or may also include a form in which the block of the corresponding node is split in the diagonal direction.
[0053] A CU may have various sizes according to the QTBT or QTBTTT split from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) is referred to as the "current block". When QTBTTT split is adopted, in addition to the square shape, the shape of the current block may also be a rectangular shape.
[0054] The predictor 120 predicts the current block to generate a prediction block. The predictor 120 includes an intra predictor 122 and an inter predictor 124.
[0055] Generally, each of the current blocks in the image can be predictively encoded. Generally, the prediction of the current block can be performed by using an intra prediction technique (which uses data from the image including the current block) or an inter prediction technique (which uses data from the image encoded before the image including the current block). Inter prediction includes both uni-directional prediction and bi-directional prediction.
[0056] The intra predictor 122 predicts the pixels in the current block by using the pixels (reference pixels) adjacent to the current block in the current image including the current block. According to the prediction direction, there are multiple intra prediction modes. For example, as Figure 3aAs shown, multiple intra prediction modes may include two non-directional modes including a Planar mode and a DC mode, and may include 65 directional modes. Adjacent pixels to be used and algorithmic equations are defined differently according to each prediction mode.
[0057] For efficient directional prediction of a current block having a rectangular shape, the directional modes shown by the dashed arrows in Figure 3b (#67 to #80, intra prediction modes #-1 to #-14) may additionally be used. The directional modes may be referred to as "wide angle intra-prediction modes". In Figure 3b , the arrows indicate the corresponding reference samples for prediction, rather than representing the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. When the current block has a rectangular shape, the wide angle intra-prediction mode is a mode that performs prediction in the direction opposite to a specific directional mode without additional bit transmission. In this case, in the wide angle intra-prediction mode, some wide angle intra-prediction modes available for the current block may be determined by the ratio of the width to the height of the current block having a rectangular shape. For example, when the current block has a rectangular shape with a height less than the width, wide angle intra-prediction modes having an angle less than 45 degrees (intra prediction modes #67 to #80) are available. When the current block has a rectangular shape with a width greater than the height, wide angle intra-prediction modes having an angle greater than -135 degrees are available.
[0058] The intra predictor 122 may determine the intra prediction to be used for encoding the current block. In some examples, the intra predictor 122 may encode the current block by using multiple intra prediction modes, and may also select an appropriate intra prediction mode to be used from test modes. For example, the intra predictor 122 may calculate rate-distortion values by performing rate-distortion analysis on multiple tested intra prediction modes, and may also select an intra prediction mode having the best rate-distortion characteristics in the test modes.
[0059] The intra predictor 122 selects one intra prediction mode from multiple intra prediction modes, and predicts the current block by using adjacent pixels (reference pixels) and algorithmic equations determined according to the selected intra prediction mode. Information about the selected intra prediction mode is encoded by the entropy encoder 155 and transmitted to the video decoding device.
[0060] The inter - frame predictor 124 generates a predicted block of the current block by using motion - compensation processing. The inter - frame predictor 124 searches for the block most similar to the current block in a reference image that has been encoded and decoded earlier than the current image, and generates a predicted block of the current block by using the searched - for block. In addition, a motion vector (MV) is generated, which corresponds to the displacement between the current block in the current image and the predicted block in the reference image. Generally, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma component and the chroma component. The entropy encoder 155 encodes the motion information including the information of the reference image and the information about the motion vector used for predicting the current block, and transmits it to the video decoding device.
[0061] The inter - frame predictor 124 can also perform interpolation of the reference image or reference block to increase the prediction accuracy. In other words, sub - samples are interpolated between two consecutive integer samples by applying filter coefficients to a plurality of consecutive integer samples including two integer samples. When performing the process of searching for the block most similar to the current block on the interpolated reference image, the motion vector can represent fractional - unit precision instead of integer - sample - unit precision. For each target region to be encoded, such as units like slices, tiles, CTUs, CUs, etc., the precision or resolution of the motion vector can be set differently. When applying such adaptive motion vector resolution (AMVR), information about the motion vector resolution to be applied to each target region should be signaled. For example, when the target region is a CU, information about the motion vector resolution applied to each CU is signaled. The information about the motion vector resolution can be information representing the precision of the motion vector difference described below.
[0062] On the other hand, the inter-frame predictor 124 can perform inter-frame prediction by using bidirectional prediction. In the case of bidirectional prediction, two reference images and two motion vectors representing the positions of the blocks most similar to the current block in each reference image are used. The inter-frame predictor 124 selects a first reference image and a second reference image from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively. The inter-frame predictor 124 also searches for the block most similar to the current block in the corresponding reference image to generate a first reference block and a second reference block. In addition, a predicted block of the current block is generated by averaging or weighted averaging the first reference block and the second reference block. Further, motion information including information about the two reference images used for predicting the current block and information about the two motion vectors is transmitted to the entropy encoder 155. Here, reference picture list 0 may be composed of images in the pre-reconstructed images that are before the current image in the display order, and reference picture list 1 may be composed of images in the pre-reconstructed images that are after the current image in the display order. However, although not particularly limited thereto, pre-reconstructed images after the current image in the display order may be additionally included in reference picture list 0. Conversely, pre-reconstructed images before the current image may also be additionally included in reference picture list 1.
[0063] To minimize the amount of bits consumed for encoding motion information, various methods can be used.
[0064] For example, when the reference image and motion vector of the current block are the same as those of an adjacent block, information identifying the adjacent block is encoded to transmit the motion information of the current block to the video decoding device. This method is called merge mode.
[0065] In merge mode, the inter-frame predictor 124 selects a predetermined number of merge candidates (hereinafter referred to as "merge candidates") from the adjacent blocks of the current block.
[0066] As the adjacent blocks for deriving merge candidates, all or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current image can be used, as Figure 4 shown. In addition, in addition to the current image where the current block is located, blocks within the reference image (which may be the same as or different from the reference image used for predicting the current block) can also be used as merge candidates. For example, the co-located block of the current block within the reference image or a block adjacent to the co-located block can be additionally used as a merge candidate. If the number of merge candidates selected by the above method is less than the preset number, zero vectors are added to the merge candidates.
[0067] The inter-frame predictor 124 configures a merge list including a predetermined number of merge candidates by using adjacent blocks. A merge candidate to be used as the motion information of the current block is selected from among the merge candidates included in the merge list, and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoder 155 and transmitted to the video decoding device.
[0068] The merge skip mode is a special case of the merge mode. After quantization, when all transform coefficients for entropy coding are close to zero, only the adjacent block selection information is transmitted without transmitting the residual signal. By using the merge skip mode, relatively high coding efficiency can be achieved for images with slight motion, still images, screen content images, etc.
[0069] Hereafter, the merge mode and the merge skip mode are collectively referred to as the merge / skip mode.
[0070] Another method for encoding motion information is the advanced motion vector prediction (AMVP) mode.
[0071] In the AMVP mode, the inter-frame predictor 124 derives motion vector prediction candidates for the motion vector of the current block by using adjacent blocks of the current block. As the adjacent blocks for deriving the motion vector prediction candidates, all or some of the left block A0, the lower left block A1, the upper block B0, the upper right block B1, and the upper left block B2 adjacent to the current block in the current image shown in Figure 4 can be used. In addition, in addition to the current image where the current block is located, blocks located in a reference image (which may be the same as or different from the reference image used for predicting the current block) can also be used as adjacent blocks for deriving the motion vector prediction candidates. For example, the co-located block of the current block in the reference image or a block adjacent to the co-located block can be used. If the number of motion vector candidates selected by the above method is less than a preset number, a zero vector is added to the motion vector candidates.
[0072] The inter-frame predictor 124 derives motion vector prediction candidates by using the motion vectors of adjacent blocks, and determines the motion vector prediction of the motion vector of the current block by using the motion vector prediction candidates. In addition, the motion vector difference is calculated by subtracting the motion vector prediction from the motion vector of the current block.
[0073] Motion vector prediction can be obtained by applying a predefined function (e.g., median and average calculation, etc.) to motion vector prediction candidates. In this case, the video decoding device also knows the predefined function. In addition, since the neighboring blocks used to derive the motion vector prediction candidates are blocks that have already been encoded and decoded, the video decoding device may also already know the motion vectors of the neighboring blocks. Therefore, the video encoding device does not need to encode the information for identifying the motion vector prediction candidates. Accordingly, in this case, the information about the motion vector difference and the information about the reference image used to predict the current block are encoded.
[0074] On the other hand, motion vector prediction can also be determined by a scheme of selecting any one of the motion vector prediction candidates. In this case, the information for identifying the selected motion vector prediction candidate is additionally encoded together with the information about the motion vector difference and the information about the reference image used to predict the current block.
[0075] The subtractor 130 generates a residual block by subtracting the prediction block generated by the intra predictor 122 or the inter predictor 124 from the current block.
[0076] The transformer 140 transforms the residual signal in the residual block having pixel values in the spatial domain into transform coefficients in the frequency domain. The transformer 140 can transform the residual signal in the residual block by using the entire size of the residual block as the transform unit, or the residual block can also be divided into multiple sub-blocks, and the transform can be performed by using the sub-blocks as the transform units. Alternatively, the residual block is divided into two sub-blocks, i.e., a transform region and a non-transform region, to transform the residual signal by using only the transform region sub-block as the transform unit. Here, the transform region sub-block can be one of two rectangular blocks having a size ratio of 1:1 based on the horizontal axis (or the vertical axis). In this case, the entropy encoder 155 encodes a flag (cu_sbt_flag) indicating only the transformed sub-block, and the direction (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or the position information (cu_sbt_pos_flag), and signals it to the video decoding device. In addition, the size of the transform region sub-block can have a size ratio of 1:3 based on the horizontal axis (or the vertical axis). In this case, the entropy encoder 155 additionally encodes a flag (cu_sbt_quad_flag) for dividing the corresponding segmentation, and signals it to the video decoding device.
[0077] On the other hand, the transformer 140 can perform the transformation of the residual block separately in the horizontal direction and the vertical direction. For this transformation, various types of transformation functions or transformation matrices can be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation can be defined as a multiple transform set (MTS). The transformer 140 can select a pair of transformation functions in the MTS with the highest transformation efficiency and can transform the residual block on each of the horizontal direction and the vertical direction. The entropy encoder 155 encodes the information (mts_idx) regarding the pair of transformation functions in the MTS and signals it to the video decoding device.
[0078] The quantizer 145 quantizes the transform coefficients output from the transformer 140 using quantization parameters and outputs the quantized transform coefficients to the entropy encoder 155. The quantizer 145 can also quantize the relevant residual block immediately without performing transformation on any block or frame. The quantizer 145 can also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients in the transform block. The quantization matrix applied to the quantized transform coefficients arranged in two dimensions can be encoded and signaled to the video decoding device.
[0079] The rearrangement unit 150 can perform rearrangement of the coefficient values on the quantized residual values.
[0080] The rearrangement unit 150 can change the 2D coefficient array into a 1D coefficient sequence by using coefficient scanning. For example, the rearrangement unit 150 can use a zig-zag scan or a diagonal scan to scan the coefficients from the DC coefficient to the high-frequency region to output a 1D coefficient sequence. According to the size of the transform unit and the intra prediction mode, a vertical scan that scans the 2D coefficient array in the column direction and a horizontal scan that scans the 2D block type coefficients in the row direction can also be used instead of the zig-zag scan. In other words, according to the size of the transform unit and the intra prediction mode, the scan method to be used can be determined among the zig-zag scan, the diagonal scan, the vertical scan, and the horizontal scan.
[0081] The entropy encoder 155 encodes the sequence of 1D quantized transform coefficients output from the rearrangement unit 150 by using various coding schemes including context-based adaptive binary arithmetic coding (CABAC), exponential Golomb, etc. to generate a bitstream.
[0082] In addition, the entropy encoder 155 encodes information related to block partitioning (e.g., CTU size, CTU partitioning flag, QT partitioning flag, MTT partitioning type, and MTT partitioning direction, etc.) so that the video decoding device can partition blocks in the same way as the video encoding device. In addition, the entropy encoder 155 encodes information regarding the prediction type indicating whether the current block is encoded by intra prediction or inter prediction. The entropy encoder 155 encodes intra prediction information (i.e., information regarding the intra prediction mode) or inter prediction information (merge index in the case of the merge mode, and information regarding the reference image index and motion vector difference in the case of the AMVP mode) according to the prediction type. In addition, the entropy encoder 155 encodes information related to quantization (i.e., information regarding the quantization parameter and information regarding the quantization matrix).
[0083] The inverse quantizer 160 inverse quantizes the quantized transform coefficients output from the quantizer 145 to generate transform coefficients. The inverse transformer 165 transforms the transform coefficients output from the inverse quantizer 160 from the frequency domain to the spatial domain to reconstruct the residual block.
[0084] The adder 170 adds the reconstructed residual block and the prediction block generated by the predictor 120 to reconstruct the current block. When performing intra prediction on the next block, the pixels in the reconstructed current block are used as reference pixels.
[0085] The loop filter unit 180 performs filtering on the reconstructed pixels to reduce block artifacts, ringing artifacts, blurring artifacts, etc. that occur due to block-based prediction and transform / quantization. The loop filter unit 180, as an in-loop filter, may include all or some of a deblocking filter 182, a sample adaptive offset (SAO) filter 184, and an adaptive loop filter (ALF) 186.
[0086] The deblocking filter 182 filters the boundaries between the reconstructed blocks to remove the blocking artifacts that occur due to block-based coding / decoding, and the SAO filter 184 and the ALF 186 perform additional filtering on the deblocked video. The SAO filter 184 and the ALF 186 are filters for compensating for the difference between the reconstructed pixels and the original pixels that occurs due to lossy coding. The SAO filter 184 applies an offset in units of CTUs to enhance the subjective image quality and coding efficiency. On the other hand, the ALF 186 performs block-based filtering and applies different filters by dividing the degree of the boundary and the variation amount of the corresponding block to compensate for distortion. Information about the filter coefficients to be used for the ALF can be encoded and signaled to the video decoding apparatus.
[0087] The reconstructed blocks filtered by the deblocking filter 182, the SAO filter 184, and the ALF 186 are stored in the memory 190. When all the blocks in an image are reconstructed, the reconstructed image can be used as a reference image for inter prediction of the blocks within the image to be encoded subsequently.
[0088] The video encoding apparatus may store the bitstream of the encoded video data in a non-volatile storage medium or transmit the bitstream to the video decoding apparatus through a communication network.
[0089] Figure 5 is a functional block diagram of a video decoding apparatus that can implement the technology of the present invention. Hereinafter, with reference to Figure 5 ,the video decoding apparatus and the components of the apparatus are described.
[0090] The video decoding apparatus may include an entropy decoder 510, a rearrangement unit 515, an inverse quantizer 520, an inverse transformer 530, a predictor 540, an adder 550, a loop filter unit 560, and a memory 570.
[0091] Similar to Figure 1 the video encoding apparatus, each component of the video decoding apparatus may be implemented as hardware or software, or implemented as a combination of hardware and software. In addition, the functions of each component may be implemented as software, and the microprocessor may also be implemented to execute the functions of the software corresponding to each component.
[0092] The entropy decoder 510 extracts information related to block partitioning by decoding the bitstream generated by the video encoding apparatus to determine the current block to be decoded, and extracts the prediction information and the information about the residual signal required for reconstructing the current block.
[0093] The entropy decoder 510 determines the size of a coding tree unit (CTU) by extracting information about the CTU size from a sequence parameter set (SPS) or a picture parameter set (PPS), and divides an image into CTUs having the determined size. In addition, a CTU is determined as the highest layer (i.e., the root node) of a tree structure, and segmentation information of the CTU is extracted to divide the CTU by using the tree structure.
[0094] For example, when dividing a CTU by using a QTBTTT structure, first, a first flag (QT_split_flag) related to the segmentation of a quantization tree (QT) is extracted to divide each node into four lower-layer nodes. In addition, a second flag (mtt_split_flag) related to the segmentation of a multi-tree transform (MTT), a segmentation direction (vertical / horizontal), and / or a segmentation type (binary / trinary) are extracted with respect to a node corresponding to a leaf node of the QT to divide the corresponding leaf node into an MTT structure. As a result, each node below the leaf node of the QT is recursively divided into a binary tree (BT) or a ternary tree (TT) structure.
[0095] As another example, when dividing a CTU by using a QTBTTT structure, a CU segmentation flag (split_cu_flag) indicating whether to divide a coding unit (CU) is extracted. When dividing the corresponding block, the first flag (QT_split_flag) may also be extracted. During the division process, for each node, zero or more recursive MTT divisions may occur after zero or more recursive QT divisions. For example, for a CTU, the MTT division may occur immediately, or conversely, only multiple QT divisions may occur.
[0096] As another example, when dividing a CTU by using a QTBT structure, a first flag (QT_split_flag) related to the segmentation of a QT is extracted to divide each node into four lower-layer nodes. In addition, a segmentation flag (split_flag) indicating whether to further divide the node corresponding to the leaf node of the QT into a BT and segmentation direction information are extracted.
[0097] On the other hand, when the entropy decoder 510 determines a current block to be decoded by using the division of a tree structure, the entropy decoder 510 extracts information about a prediction type indicating whether the current block is intra-frame predicted or inter-frame predicted. When the prediction type information indicates intra-frame prediction, the entropy decoder 510 extracts a syntax element for intra-frame prediction information (intra-frame prediction mode) of the current block. When the prediction type information indicates inter-frame prediction, the entropy decoder 510 extracts information about syntax elements representing inter-frame prediction information, that is, a motion vector and a reference image to which the motion vector refers.
[0098] In addition, the entropy decoder 510 extracts quantization-related information and extracts information about the quantized transform coefficients of the current block as information about the residual signal.
[0099] The rearrangement unit 515 can change the sequence of 1D quantized transform coefficients entropy decoded by the entropy decoder 510 back into a 2D coefficient array (i.e., a block) in the reverse order of the coefficient scan order performed by the video coding device.
[0100] The inverse quantizer 520 inverse quantizes the quantized transform coefficients and inverse quantizes the quantized transform coefficients by using the quantization parameter. The inverse quantizer 520 can also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in 2D. The inverse quantizer 520 can perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video coding device to the 2D array of quantized transform coefficients.
[0101] The inverse transformer 530 reconstructs the residual signal by inverse-transforming the inverse quantized transform coefficients from the frequency domain to the spatial domain to generate a residual block of the current block.
[0102] In addition, when the inverse transformer 530 inverse-transforms a partial region (sub-block) of the transform block, the inverse transformer 530 extracts a flag (cu_sbt_flag) for inverse-transforming only the sub-block of the transform block, direction (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block. The inverse transformer 530 also inverse-transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to reconstruct the residual signal, and fills the un-inverse-transformed region with the value "0" as the residual signal to generate the final residual block of the current block.
[0103] In addition, when applying MTS, the inverse transformer 530 determines the transform function or transform matrix to be applied in each of the horizontal and vertical directions by using the MTS information (mts_idx) signaled from the video coding device. The inverse transformer 530 also performs inverse transformation on the transform coefficients in the transform block in the horizontal and vertical directions by using the determined transform function.
[0104] The predictor 540 can include an intra predictor 542 and an inter predictor 544. When the prediction type of the current block is intra prediction, the intra predictor 542 is activated, and when the prediction type of the current block is inter prediction, the inter predictor 544 is activated.
[0105] The intra predictor 542 determines the intra prediction mode of the current block among multiple intra prediction modes according to the syntax element of the intra prediction mode extracted from the entropy decoder 510. The intra predictor 542 also predicts the current block according to the intra prediction mode by using the adjacent reference pixels of the current block.
[0106] The inter-frame predictor 544 determines the motion vector of the current block and the reference image for motion vector reference by using the syntax element of the inter-frame prediction mode extracted from the entropy decoder 510.
[0107] The adder 550 reconstructs the current block by adding the residual block output from the inverse transformer 530 to the prediction block output from the inter-frame predictor 544 or the intra-frame predictor 542. When performing intra-frame prediction on a block to be decoded subsequently, the pixels in the reconstructed current block are used as reference pixels.
[0108] The loop filter unit 560 as an in-loop filter may include a deblocking filter 562, a SAO filter 564, and an ALF 566. The deblocking filter 562 performs deblocking filtering on the boundaries between the reconstructed blocks to remove block artifacts that occur due to block-based unit decoding. The SAO filter 564 and the ALF 566 perform additional filtering on the reconstructed blocks after deblocking filtering to compensate for the difference between the reconstructed pixels and the original pixels that occurs due to lossy coding. The filter coefficients of the ALF are determined by using the information about the filter coefficients decoded from the bitstream.
[0109] The reconstructed blocks filtered by the deblocking filter 562, the SAO filter 564, and the ALF 566 are stored in the memory 570. When all the blocks in an image are reconstructed, the reconstructed image can be used as a reference image for inter-frame prediction of the blocks within the image to be encoded subsequently.
[0110] In some embodiments, the present invention relates to encoding and decoding video images as described above. More specifically, the present invention provides a video encoding and decoding method and apparatus that, when generating the most probable mode list (MPM list) of the current block and the base block is in an irregular intra-prediction mode or an irregular intra-mode, changes the base block to a block having a regular intra-prediction mode, and then adaptively generates the MPM list by using the changed base block.
[0111] The following embodiments may be performed by the intra-frame predictor 122 in a video encoding apparatus. The following embodiments may also be performed by the intra-frame predictor 542 in a video decoding apparatus.
[0112] The video encoding apparatus may generate signaling information associated with the present embodiment from the perspective of optimizing rate distortion when encoding the current block. The video encoding apparatus may encode the signaling information by using the entropy encoder 155 and send the encoded signaling information to the video decoding apparatus. The video decoding apparatus may decode the signaling information associated with the decoding of the current block from the bitstream by using the entropy decoder 510.
[0113] In the following description, the term "target block" may be used interchangeably with the current block or coding unit (CU), or may refer to some regions of the coding unit.
[0114] In addition, a value of true for a flag indicates the case where the flag is set to 1. Further, a value of false for a flag indicates the case where the flag is set to 0.
[0115] I. Intra Prediction Mode and Most Probable Mode (MPM) Techniques
[0116] As described above, in intra prediction, a predictor for a luminance channel can be generated based on 67 intra-prediction modes (IPMs). The 67 IPMs mean 67 intra-prediction modes that can be signaled in prediction modes -14 to 80, including non-angular prediction modes such as the planar mode and the DC mode. On the other hand, prediction modes -14 to -1 and 67 to 80 are referred to as Wide Angular Intra Prediction (WAIP) modes, which are larger angle modes signaled according to the aspect ratio of the block.
[0117] Hereinafter, a regular intra-prediction mode or a regular intra mode refers to 67 intra-prediction modes that can be signaled in prediction modes -14 to prediction mode 80 according to the block aspect ratio, and the 67 intra-prediction modes include non-angular intra-prediction modes such as the planar mode and the DC mode. The set of regular intra-prediction modes refers to the set of the above 67 prediction modes. In addition, an irregular intra-prediction mode or an irregular intra mode represents a prediction mode not included in the set of regular intra-prediction modes. The set of irregular intra modes may include intra block copy (IBC), palette mode, Matrix-weighted Intra Prediction (MIP), Block-based Delta Pulse Code Modulation (BDPCM), and inter prediction.
[0118] When generating a predictor by using one of the 67 prediction modes, a video coding device signals the prediction mode by using the most probable mode (MPM) to efficiently transmit the prediction mode information.
[0119] MPM makes use of the characteristic that when a block is coded in an intra-prediction mode, the prediction modes of adjacent blocks may be similar to each other. As Figure 6As shown, the base block is defined by a block including a pixel L on the left side of the lower left pixel of the current block and a pixel A on the upper side of the upper right pixel of the current block. Additionally, mode L and mode A are defined by the prediction modes of the base blocks including pixel L and pixel A respectively. Six MPM candidates can be selected based on mode L and mode A to generate an MPM list. If mode L or mode A is not a regular intra prediction mode, that is, if mode L or mode A is an irregular intra mode, or if pixel L or pixel A is unavailable because the current block is located at the boundary of a CTU (Coding Tree Unit), tile, slice, sub-picture, picture, etc., the base block including the unavailable pixel L or pixel A replaces its prediction mode with one of the set of regular intra prediction modes. For example, if mode L or mode A is one of IBC, palette, MIP, inter prediction, or if the current block is located on the boundary of a CTU, tile, slice, sub-picture, picture, etc., and pixel L or pixel A is unavailable, the prediction mode of the base block including the unavailable pixel L or pixel A is considered to be planar. If mode L or mode A is BDPCM, then mode L or mode A is replaced by one of the set of regular intra prediction modes, which is the horizontal mode of mode 18 or the vertical mode of mode 50 corresponding to the mode direction of BDPCM. Then, an MPM list can be generated according to the following MPM list generation algorithm (hereinafter referred to as the "generation algorithm").
[0120] First, if mode L or mode A is the same, and when mode L is greater than INTRA_DC, the ones selected as MPM candidates are {planar, mode L, 2 + ((mode L + 61) % 64), 2 + ((mode L - 1) % 64), 2 + ((mode L + 60) % 64), 2 + (mode L % 64)}.
[0121] Next, if mode L and mode A are not the same, and when mode L or mode A is greater than INTRA_DC, the MPM candidates are organized as follows. Here, minAB = Min(mode L, mode A), and maxAB = Max(mode L, mode A).
[0122] If both mode L and mode A are greater than INTRA_DC, and when maxAB - minAB = 1, the ones selected as MPM candidates are {planar, mode L, mode A, 2 + ((minAB + 61) % 64), 2 + ((maxAB - 1) % 64), 2 + ((minAB + 60) % 64)}.
[0123] If both mode L and mode A are greater than INTRA_DC, and when maxAB - minAB ≥ 62, the ones selected as MPM candidates are {Plane, mode L, mode A, 2 + ((minAB - 1) % 64), 2 + ((maxAB + 61) % 64), 2 + (minAB % 64)}.
[0124] If both mode L and mode A are greater than INTRA_DC, and when maxAB - minAB = 2, the ones selected as MPM candidates are {Plane, mode L, mode A, 2 + ((minAB - 1) % 64), 2 + ((minAB + 61) % 64), 2 + ((maxAB - 1) % 64)}.
[0125] If both mode L and mode A are greater than INTRA_DC, and when 2 < maxAB - minAB < 62, the ones selected as MPM candidates are {Plane, mode L, mode A, 2 + ((minAB + 61) % 64), 2 + ((minAB - 1) % 64), 2 + ((maxAB + 61) % 64)}.
[0126] If mode L and mode A are not the same, and when one of mode L and mode A is greater than INTRA_DC, the ones selected as MPM candidates are {Plane, maxAB, 2 + ((maxAB + 61) % 64), 2 + ((maxAB - 1) % 64), 2 + ((maxAB + 60) % 64), 2 + (maxAB % 64)}.
[0127] In addition, if both mode L and mode A are equal to or less than INTRA_DC, the ones selected as MPM candidates are {Plane, INTRA_DC, INTRA_ANGULAR50, INTRA_ANGULAR18, INTRA_ANGULAR46, INTRA_ANGULAR54}.
[0128] On the other hand, if MPM is used, the video decoding device analyzes the intra prediction mode of the current block as shown in Table 1.
[0129] [Table 1]
[0130]
[0131] The video decoding device decodes the intra_luma_mpm_flag to determine whether to utilize the MPM list. If the intra_luma_mpm_flag is true and the intra_luma_ref_idx is 0, then the flag intra_luma_not_planar_flag indicating whether to utilize the planar mode can be signaled from the video encoding device to the video decoding device. The first element in the MPM list, MPM index 0, is in the planar mode and is thus signaled with the intra_luma_not_planar_flag. If the intra_luma_not_planar_flag is false, the intra prediction mode is set to the planar mode. On the other hand, if the intra_luma_not_planar_flag is true, the intra_luma_mpm_idx can be additionally signaled to indicate MPM indices 1 - 5. If the intra_luma_not_planar_flag does not exist, it can be inferred as true.
[0132] On the other hand, if the intra_luma_ref_idx is non - zero, the planar mode is not utilized. Thus, the intra_luma_not_planar_flag is not sent and is assumed to be true. Additionally, since the intra_luma_not_planar_flag is true, the intra_luma_mpm_idx can be signaled.
[0133] The following embodiments are described with reference to a video decoding device, but they can also be implemented by a video encoding device in the same or similar manner.
[0134] II. Embodiments of the present invention
[0135] As Figure 6 In the example of, the problem of the fixed positions at L and A of the pixels for base block determination is that prediction information from blocks other than the block including pixel L and pixel A cannot be utilized during the MPM list generation process. Thus, the characteristic that the current block has a high probability of utilizing the modes in the MPM list cannot be utilized, resulting in a reduction in video codec efficiency or a deterioration in video quality. When the mode of the base block is a non - regular intra mode, the present invention can solve the prior art problem by adaptively converting the non - regular intra mode into a regular intra prediction mode to effectively generate the MPM list.
[0136] Figure 7 is a schematic diagram showing the generation of the MPM list when the mode of the block including pixel L is a non - regular intra mode.
[0137] As Figure 7As shown, blocks N3 and N1 including pixels L and A can be determined as base blocks for generating the MPM list. Since the mode of block N3 is IBC (Intra Block Copy) included in the set of non-regular intra-frame modes, the video decoding device replaces IBC as described above to convert the mode of block N3 to the planar mode. Since all the prediction modes of the base block are not angular modes, the generation algorithm generates an MPM list of {0 (planar), 1 (DC), 50 (vertical), 18 (horizontal), 46, 54}. However, if the current block selects mode 37 which is not in the MPM list, a large number of bits are required to transmit this mode 37.
[0138] Figure 8 is a schematic diagram showing the change of the base block when the mode of the block including pixel L is a non-regular intra-frame mode according to at least one embodiment of the present invention.
[0139] In this case, as in the example of Figure 8 changing the pixel used to determine the base block from pixel L to pixel Up-L to change the base block to block N2, or directly changing the base block to block N2 can enable the video decoding device to utilize mode 36 when generating the MPM list. Mode 36 is the prediction mode of block N2 and is included in the set of regular intra-frame prediction modes. Since one of the modes of the base block is an angular mode, the generation algorithm generates an MPM list of {0 (planar), 36, 35, 37, 34, 38}. Therefore, the mode 37 selected by the current block exists in the MPM list, which can reduce the bits for transmitting this mode 37. The change of the base block satisfies the characteristic that the modes in the MPM list are highly likely to be utilized by the current block, which can improve the coding and decoding efficiency and enhance the video quality.
[0140] Figure 9 is a schematic diagram showing the construction of the MPM list when the mode of the block including pixel A is a non-regular intra-frame mode.
[0141] Figure 9 The example of Figure 7 shows an example different from the example of Figure 7 where the prediction mode of block N2 including pixel A is a non-regular intra-frame mode. Since the prediction mode of block N2 is IBC, the video decoding device replaces IBC as described above to convert the mode of block N2 to the planar mode. Since all the prediction modes of the base block are not angular modes, the generation algorithm generates an MPM list of {0 (planar), 1 (DC), 50 (vertical), 18 (horizontal), 46, 54}. However, similar to the example in
[0142] Figure 10It is a schematic diagram showing the change of the base block when the mode of the block including pixel A is a non-regular intra mode according to at least one embodiment of the present invention.
[0143] In this case, as in the example of Figure 10 , the base block is changed to the N1 block by changing the pixel for determining the base block from pixel A to pixel Lf-A, or directly changing the base block to the N1 block can enable the video decoding device to utilize the prediction mode of the N1 block included in the set of regular intra prediction modes, i.e., mode 36, when generating the MPM list. Since one of the modes of the base block is an angular mode, the generation algorithm generates an MPM list of {0 (plane), 36, 35, 37, 34, 38}. Therefore, the mode 37 selected by the current block exists in the MPM list, which can reduce the bits for transmitting mode 37. Figure 10 The example of Figure 8 can produce the same effect as the example of
[0144] To solve the prior art problems in the process of generating the MPM list, the video decoding device of the present invention checks whether the mode of the base block is a non-regular intra mode, and if so, it adaptively changes the non-regular intra mode to a regular intra prediction mode to effectively generate the MPM list, as in the above example. For example, if there is a base block with a non-regular intra prediction mode during the MPM list generation process, the video decoding device exchanges the base block by using the information of the current block and its adjacent reconstructed blocks, high-level information, and other signal information related to the operation of the present invention, so as to convert the mode of the changed block to a regular intra prediction mode. The video decoding device can generate the MPM list by using the prediction mode of the changed block.
[0145] The applicability of the present invention can be signaled or inferred. If there are multiple base block change methods when applying the present invention, the video decoding device can parse one of the block change methods and can continue to exchange the base block according to the parsed method. Alternatively, the video decoding device can continue with the base block change according to a predetermined method. The method of determining whether to apply the method of the present invention and the method of determining the method for base block change can be combined in various ways.
[0146] For example, the applicability of the present invention can be signaled to the bitstream, and a predetermined method can be used for the base block change method. Alternatively, the applicability of the present invention can be inferred, and the method for base block change can be signaled to the bitstream. Example embodiments of the present invention are as follows.
[0147] As used herein, the term "proximity" refers to the spatial proximity of two objects, and the terms "adjacent" (including "proximity") and the concept of an object mean that another object exists within a specific spatial distance of an object.
[0148] <Embodiment 1> Changing the base block
[0149] In this embodiment, if the mode of the base block is a non-regular intra mode, the video decoding device exchanges the base block by considering the information of the current block and its adjacent reconstructed blocks, so as to convert the mode of the changed block into a regular intra prediction mode. The existing MPM technology is based on the feature that the current block and its adjacent blocks have similar prediction modes. Therefore, the method for changing the base block also needs to consider the information of the current block and its adjacent reconstructed blocks, including the prediction mode of the adjacent blocks. The information of the current block and its adjacent reconstructed blocks may include the width / height / area / aspect ratio / position of the current block, the width / height / area / prediction mode / aspect ratio / position / distance to the current block of the adjacent blocks of the current block, the position / number / value / distance to the current block of the adjacent pixels of the current block, etc.
[0150] There are two specific methods for converting to the base block.
[0151] Figures 11a to 11d is a schematic diagram showing a predetermined adjacent pixel search order according to at least one embodiment of the present invention.
[0152] In the first method, if the prediction mode of the block including pixels is determined to be a regular intra prediction mode by performing a search according to a predetermined adjacent pixel search order, as Figures 11a to 11d shown in the example of, then the block including pixels is converted into a base block.
[0153] If the base block is on the left side of the current block, the base block can be searched in the following order: ① (lower side → upper side), ② (upper side → lower side), ⑤ (lower side → upper side → right side), ⑦ (zigzag order starting from the upper left pixel and away from pixel L), ⑧ (zigzag order starting from the upper left pixel and away from pixel A). If the base block is on the upper side of the current block, the base block can be searched in the following order: ③ (right side → left side), ④ (left side → right side), ⑥ (right side → left side → lower side), ⑦ (zigzag order starting from the upper left pixel and away from pixel L), ⑧ (zigzag order starting from the upper left pixel and away from pixel A). Except for Figures 11a to 11d the example of, various preset adjacent pixel search orders can be utilized.
[0154] In the second method, one of the current block and the adjacent reconstructed information items is used as a predetermined base block selection metric (e.g., proximity length, area, aspect ratio, etc.), and if the preset condition (e.g., maximum value) of each metric is satisfied and the prediction mode of the block is a regular intra prediction mode, the block is converted into a base block.
[0155] Based on the above method, if the mode of the base block is a non-regular intra mode, the method of changing the base block to convert the mode of the block into a regular intra prediction mode can be one of the following methods: (Embodiment 1-1) Search the adjacent pixels of the current block according to a predetermined search order, (Embodiment 1-2) Change the base block to a block having a regular intra prediction mode and the longest proximity length to the current block among the adjacent blocks, (Embodiment 1-3) Change the base block to a block having a regular intra prediction mode and the largest area among the adjacent blocks, and (Embodiment 1-4) Change the base block to a block having a regular intra prediction mode and an aspect ratio equal to (or most similar to) the aspect ratio of the current block among the adjacent blocks.
[0156] In the following example method, the adjacent blocks or pixels considered are all blocks adjacent to the current block. However, the range of adjacent blocks or pixels that can be considered can vary according to the embodiment, including those blocks and pixels slightly farther away.
[0157] The following describes each embodiment in detail.
[0158] <Embodiment 1-1> Method of Searching Adjacent Pixels of a Current Block According to a Predetermined Search Order
[0159] In this embodiment, the video decoding device searches the adjacent pixels of the current block according to a predetermined search order, excluding the initial pixels, and changes the base block to a block including the first pixel predicted in the regular intra prediction mode. At this time, pixels included in the base block including pixel L or pixel A can be excluded from the search.
[0160] Figure 12 is a schematic diagram showing the change of the base block by using the adjacent pixel search order according to at least one embodiment of the present invention.
[0161] In an example case, as Figure 12 shown, the prediction mode of the base block (the block including pixel L) adjacent to the left side of the current block is a non-regular intra mode. The video decoding device can search the pixels adjacent to the left side of the current block in the order of ① (from the lower side to the upper side) and can change the base block to the searched block, thereby converting the mode of the block into a regular intra prediction mode. In Figure 12In the example, since the mode of the N3 block including pixel L is IBC, the video decoding device searches for adjacent pixels of the current block in the order of ① (from the lower side to the upper side). At this time, since the prediction mode of the N2 block including pixel Up-L is mode 50, and the N2 block is the first block in the search order predicted according to the regular intra prediction mode, the video decoding device changes the base block to the N2 block. After changing the base block to the N2 block, an MPM list can be generated by using mode 50, which is the prediction mode of the N2 block.
[0162] In another example case, as Figure 13 shown, the prediction mode of the base block (the block including pixel A) adjacent to the upper side of the current block is an irregular intra mode. The video decoding device can search for pixels adjacent to the upper side of the current block in the order of ③ (right side → left side) and can change the base block to the searched block, thereby converting the mode of the block to the regular intra prediction mode. In Figure 13 the example, since the mode of the N2 block including pixel A is IBC, the video decoding device searches for adjacent pixels of the current block in the order of ③ (right side → left side). At this time, since the prediction mode of the N1 block including pixel Lf-A is mode 66, and the N1 block is the first block in the search order predicted according to the regular intra prediction mode, the video decoding device changes the base block to the N1 block. After changing the base block to the N1 block, an MPM list can be generated by using mode 66, which is the prediction mode of the N1 block.
[0163] <Embodiment 1-2> Method of changing the base block to a block having a regular intra prediction mode and the longest adjacent length to the current block among adjacent blocks
[0164] In this embodiment, the video decoding device can change the base block to a block whose prediction mode is the regular intra prediction mode and whose adjacent length to the current block is the longest among the adjacent blocks of the current block. When swapping the base block adjacent to the left side of the current block (the block including pixel L), search for blocks around the left side. The video decoding device can measure the adjacent length between the right side of the left adjacent block and the left side of the current block, and can change the base block to the block having the longest adjacent length. When changing the base block adjacent to the upper side of the current block (the block including pixel A), search for blocks around the upper side. The video decoding device can measure the adjacent length between the lower side of the upper adjacent block and the upper side of the current block, and can change the base block to the block having the longest adjacent length.
[0165] Figure 14 is a schematic diagram showing the change of the base block by using the adjacent length according to at least one embodiment of the present invention.
[0166] In yet another example case, as Figure 14As shown, the prediction mode of the base block (including the block of pixel L) adjacent to the left side of the current block is a non-regular intra mode. The video decoding apparatus may search for the block having the longest adjacent length to the current block among the adjacent blocks adjacent to the left side of the current block, and may change the base block to the searched block, thereby converting the mode of the block to a regular intra prediction mode. In Figure 14 In the example of, since the mode of the N3 block including pixel L is IBC, the video decoding apparatus searches for the block having the longest adjacent length to the current block among the adjacent blocks adjacent to the left side of the current block. At this time, the video decoding apparatus changes the base block to the N2 block because the N2 block has the longest adjacent length of 8 among the adjacent blocks adjacent to the left side of the current block, and the prediction mode of the N2 block is mode 50 included in the set of regular intra prediction modes. After changing the base block to the N2 block, an MPM list may be generated by using mode 50, which is the prediction mode of the N2 block.
[0167] In yet another example case, as Figure 15 shown, the prediction mode of the base block (including the block of pixel A) adjacent to the upper side of the current block is a non-regular intra mode. The video decoding apparatus may search for the block having the longest adjacent length to the current block among the adjacent blocks adjacent to the upper side of the current block, and may change the base block to the searched block, thereby converting the mode of the block to a regular intra prediction mode. In Figure 15 In the example of, since the mode of the N2 block including pixel A is IBC, the video decoding apparatus searches for the block having the longest adjacent length to the current block among the adjacent blocks adjacent to the upper side of the current block. At this time, the video decoding apparatus changes the base block to the N1 block because the N1 block has the longest adjacent length of 8 among the adjacent blocks adjacent to the upper side of the current block, and the prediction mode of the N1 block is mode 66 included in the set of regular intra prediction modes. After changing the base block to the N1 block, an MPM list may be generated by using mode 66, which is the prediction mode of the N1 block.
[0168] On the other hand, if two or more blocks have the same longest adjacent length, the base block may be selected as one of the two or more blocks based on a pre-allocated priority. In this case, the priority may be in an order such as {plane, DC, horizontal mode, vertical mode,...}, ascending order of mode numbers, descending order of mode numbers, etc. based on the prediction mode of each block. Alternatively, the higher the priority may be assigned to the block closer to the upper left side. In addition to the above examples, one of the blocks having the longest adjacent length may be selected based on a predefined rule.
[0169] <Embodiment 1-3> Method of changing a base block to a block having a regular intra prediction mode and the largest area among adjacent blocks
[0170] In this embodiment, the video decoding device may change the base block to a block having a prediction mode with a regular intra prediction mode and the largest area among adjacent blocks of the current block. When changing the base block adjacent to the left side of the current block (the block including pixel L), the video decoding device may exchange the base block by searching for a block adjacent to the left side, and when changing the base block adjacent to the upper side of the current block (the block including pixel A), the video decoding device may exchange the base block by searching for a block adjacent to the upper side.
[0171] Figure 16 is a schematic diagram showing the change of the base block according to at least one embodiment of the present invention by using the area.
[0172] In one example case, as Figure 16 shown, the prediction mode of the base block (the block including pixel L) adjacent to the left side of the current block is an irregular intra mode. The video decoding device may search for a block having the largest area among adjacent blocks adjacent to the left side of the current block, and may change the base block to the searched block, thereby converting the mode of the block to a regular intra prediction mode. In Figure 16 the example, since the mode of the N3 block including pixel L is IBC, the video decoding device searches for a block having the largest area among adjacent blocks adjacent to the left side of the current block. At this time, the video decoding device changes the base block to the N2 block because the area of the N2 block is the largest 64 among adjacent blocks adjacent to the left side of the current block, and the prediction mode of the N2 block is mode 50 included in the set of regular intra prediction modes. After changing the base block to the N2 block, an MPM list may be generated by using mode 50, which is the prediction mode of the N2 block.
[0173] In another example case, as Figure 17 shown, the prediction mode of the base block (the block including pixel A) adjacent to the upper side of the current block is an irregular intra mode. The video decoding device may search for a block having the largest area among adjacent blocks adjacent to the upper side of the current block, and may change the base block to the searched block, thereby converting the mode of the block to a regular intra prediction mode. In Figure 17 the example, since the mode of the N2 block including pixel A is IBC, the video decoding device searches for a block having the largest area among adjacent blocks adjacent to the upper side of the current block. At this time, the video decoding device changes the base block to the N1 block because the area of the N1 block is the largest 64 among adjacent blocks adjacent to the upper side of the current block, and the prediction mode of the N1 block is 66 included in the set of regular intra prediction modes. After changing the base block to the N1 block, an MPM list may be generated by using mode 66, which is the prediction mode of the N1 block.
[0174] On the other hand, if two or more blocks have the same maximum area, the base block can be selected as one of the two or more blocks based on a pre-allocated priority, as in Embodiment 1-2.
[0175] <Embodiment 1-4> Method for changing a base block into a block having a regular intra prediction mode and an aspect ratio equal to (or most similar to) the aspect ratio of the current block among adjacent blocks
[0176] In this embodiment, the video decoding device may change the base block into a block having a regular intra prediction mode and an aspect ratio the same as (or most similar to) the aspect ratio of the current block among the adjacent blocks of the current block. When changing the base block adjacent to the left side of the current block (the block including pixel L), the video decoding device may exchange the base block by searching for a block adjacent to the left side, and when changing the base block adjacent to the upper side of the current block (the block including pixel A), the video decoding device may exchange the base block by searching for a block adjacent to the upper side.
[0177] Figure 18 is a schematic diagram showing the change of the base block using the aspect ratio according to at least one embodiment of the present invention.
[0178] In one example case, as Figure 18 shown, the prediction mode of the base block adjacent to the left side of the current block (the block including pixel L) is a non-regular intra mode. The video decoding device may search for a block having an aspect ratio the same as (or most similar to) the aspect ratio of the current block among the adjacent blocks adjacent to the left side of the current block, and may change the base block into the searched block, thereby converting the mode of the block into a regular intra prediction mode. In Figure 18 the example of, since the mode of the N3 block including pixel L is IBC, the video decoding device searches for a block having an aspect ratio the same as (or most similar to) the aspect ratio of the current block among the adjacent blocks adjacent to the left side of the current block. At this time, the video decoding device changes the base block into the N2 block because, among the adjacent blocks adjacent to the left side of the current block, the aspect ratio of the N2 block is the same as the aspect ratio of the current block, and the prediction mode of the N2 block is mode 50 included in the regular intra prediction mode set. After changing the base block into the N2 block, an MPM list may be generated by using mode 50, which is the prediction mode of the N2 block.
[0179] In another example case, as Figure 19 shown, the prediction mode of the base block adjacent to the upper side of the current block (the block including pixel A) is a non-regular intra mode. The video decoding device may search for a block having an aspect ratio the same as (or most similar to) the aspect ratio of the current block among the adjacent blocks adjacent to the upper side of the current block, and may change the base block into the searched block, thereby converting the mode of the block into a regular intra prediction mode. In Figure 19In the example, since the pattern of the N2 block including pixel A is IBC, the video decoding device searches for a block having the same (or most similar) aspect ratio as the current block among the adjacent blocks adjacent to the upper side of the current block. At this time, the video decoding device changes the base block to the N1 block because, among the adjacent blocks adjacent to the upper side of the current block, the aspect ratio of the block N1 is the same as the aspect ratio of the current block, and the prediction pattern of the block N1 is a pattern 66 included in the set of regular intra prediction patterns. After changing the base block to the N1 block, an MPM list can be generated using the pattern 66, which is the prediction pattern of the N1 block.
[0180] On the other hand, if two or more blocks have the same (or most similar) aspect ratio as the current block, the base block can be selected as one of the two or more blocks according to a pre-allocated priority, as in Embodiment 1-2.
[0181] In addition to the above-described embodiments of exchanging the base block by using the information of the current block and its adjacent reconstructed blocks, the base block can be exchanged by using various other methods. For example, if the base block still maintains an irregular intra mode even after the base block has been exchanged by using the above method, the mode of the base block can be replaced with one of the existing regular intra prediction patterns. In addition, the foregoing example describes a case where the prediction pattern of one of the two blocks including pixel L and pixel A is an irregular intra mode. However, when the two blocks including pixel L and pixel A have an irregular intra mode, the base block replacement can still be performed by extending the operation of the above example.
[0182] <Embodiment 2> Determining the applicability of the present invention and the method for changing the base block
[0183] In this embodiment, the video decoding device determines whether to apply the present invention, and if so, determines one of a plurality of base block change methods. The determination of the applicability of the present invention can be based on the information signaled from the video encoding device. Alternatively, the video decoding device can infer the applicability of the present invention. When applying the present invention, an index indicating one of the plurality of base block change methods according to Embodiment 1 can be signaled from the video encoding device. Alternatively, in the absence of separate signaling, the video decoding device can determine the base block change method as a predetermined method.
[0184] For example, the flag of adaptive_mpm_list_construction_enabled_flag can be used to signal the applicability of the present invention. When applying the present invention, one of the plurality of base block change methods can be signaled by using the index adaptive_mpm_list_construction_idx.
[0185] The adaptive_mpm_list_construction_enabled_flag (“adaptive MPM list generation flag”) indicates the applicability of the present invention. If the adaptive MPM list generation flag is true, the video decoding device applies the present invention. On the other hand, if the adaptive MPM list generation flag is false, the video decoding device chooses not to apply the present invention. If the present invention is not applied, a normal method is used. If the prediction mode of the base block is a non-regular intra mode, the video decoding device will use a normal method to replace the prediction mode of the base block with a mode included in the set of regular intra prediction modes.
[0186] The adaptive_mpm_list_construction_idx (“adaptive MPM list generation index”) indicates one of multiple change methods for the base block. For example, based on the adaptive MPM list generation index, the base block change methods can be classified as shown in Table 2.
[0187] [Table 2]
[0188] adaptive_mpm_list_construction_idx Changing method 0 Embodiment 1-1 1 Embodiment 1-2 2 Embodiment 1-3 3 Embodiment 1-4
[0189] As described above, the applicability of the present invention can be determined based on a flag or inferred signaling, and the change method for the base block can be determined based on an index or signaling of a preset method. The syntax associated with determining the applicability of the present invention and determining the base block change method can be combined, for example, as shown in Table 3.
[0190] [Table 3]
[0191]
[0192] Hereinafter, detailed individual descriptions of the embodiments of the combinations shown in Table 3 are given.
[0193] <Embodiment 2-1> Signaling both the applicability of the present invention and the change method for the base block
[0194] In this embodiment, the video decoding device parses both the adaptive_mpm_list_construction_enabled_flag that determines the applicability of the present invention and the adaptive_mpm_list_construction_idx that determines the change method for the base block. The above information can be signaled at various levels such as SPS / VPS / PPS / SH (slice header) / CTU (coding tree unit) / CU (coding unit). As an example, when the applicability of the present invention and the change method for the base block are signaled at the SPS level, the associated syntax can be represented as shown in Table 4.
[0195] [Table 4]
[0196] [SPS level]
[0197] sps_adaptive_mpm_list_construction_enabled_flag if(sps_adaptive_mpm_list_construction_enabled_flag) sps_adaptive_mpm_list_construction_idx
[0198] In Table 4, the video decoding device parses the sps_adaptive_mpm_list_construction_enabled_flag to determine the applicability of the present invention. If the flag is true, the video decoding device parses the sps_adaptive_mpm_list_construction_idx to determine the change method for the base block. According to Table 4, the coding and decoding efficiency can be improved because the applicability of the present invention and the change method for the base block can be determined at a higher level by transmitting a smaller amount of information at a lower level such as CTU and CU at one time.
[0199] As another example, as shown in Table 5, the applicability of the present invention and the change method for the base block can be signaled at the CU level.
[0200] [Table 5]
[0201] [CU level]
[0202]
[0203] In Table 5, the video decoding device parses the intra_luma_mpm_flag that indicates the utilization or non-utilization of the modes included in the MPM list on a per-CU basis. If the intra_luma_mpm_flag is true, the video decoding device parses the intra_luma_not_planar_flag to determine whether the current CU is utilizing the planar mode. If the intra_luma_not_planar_flag is true and the CU is utilizing a mode other than the planar mode in the MPM list, the video decoding device parses both the flag that determines the applicability of the present invention and the index that determines the change method for the base block. In this case, the applicability of the appropriate present invention and the base block change method can be signaled adaptively on a per-CU basis. Syntactically, if both the intra_luma_mpm_flag and the intra_luma_not_planar_flag are 1, additional information is signaled according to this embodiment, which can avoid signaling the additional information for all CUs at once.
[0204] As another example, as shown in Table 6, the applicability of the present invention can be signaled at the SPS level and the change method for the base block can be signaled at the CU level.
[0205] [Table 6]
[0206] [SPS level]
[0207] sps_adaptive_mpm_list_construction_enabled_flag
[0208] [CU level]
[0209]
[0210] In Table 6, the video decoding device parses the sps_adaptive_mpm_list_construction_enabled_flag at the SPS level to determine the applicability of the present invention. Then, at the CU level, the video decoding device parses the intra_luma_mpm_flag that indicates the utilization or non-utilization of the modes included in the MPM list on a per-CU basis. If the intra_luma_mpm_flag is true, the video decoding device parses the intra_luma_not_planar_flag to determine whether the current CU is utilizing the planar mode. If the intra_luma_not_planar_flag is true and a mode other than the planar mode in the MPM list is utilized, the video decoding device checks the applicability of the present invention determined at a higher level. When the present invention is applied, the video decoding device parses the index indicating the change method for the base block.
[0211] According to Table 6, the coding and decoding efficiency can be improved because the applicability of the present invention can be determined at a lower level such as CTU and CU once without additional information transfer at a higher level. In addition, an appropriate base block change method can be signaled adaptively on a per-CU basis. Table 6 shows an example of signaling the applicability of the present invention at the SPS level and signaling the base block change method at the CU level, but in various other combinations, the applicability of the present invention can be signaled at a higher level and the base block change method can be signaled at a level lower than the higher level.
[0212] <Embodiment 2-2> Signaling the applicability of the present invention and determining the base block change method as a preset method
[0213] In this embodiment, the video decoding device parses the adaptive_mpm_list_construction_enabled_flag that determines the applicability of the present invention and determines the base block change method as a preset method. The flag determining the applicability of the present invention can be signaled at various levels such as SPS / VPS / PPS / SH (slice header) / CTU (coding tree unit) / CU (coding unit). Accordingly, no additional information is transmitted to determine the change method for the base block, which can improve the coding and decoding efficiency.
[0214] In one example, the applicability of the present invention can be signaled at the SPS level as shown in Table 7.
[0215] [Table 7]
[0216] [SPS level]
[0217] sps_adaptive_mpm_list_construction_enabled_flag
[0218] In Table 7, the video decoding device parses the sps_adaptive_mpm_list_construction_enabled_flag to determine the applicability of the present invention. If this flag is true, the video decoding device determines the base block change method as a preset method. According to Table 7, the applicability of the present invention can be determined for lower levels such as CTU and CU at one time, while there are only a small number of bits at higher levels, which can improve the encoding and decoding efficiency.
[0219] As another example, the applicability of the present invention can be signaled at the CU level, as shown in Table 8.
[0220] [Table 8]
[0221] [CU level]
[0222]
[0223] In Table 8, the video decoding device parses the intra_luma_mpm_flag on a per-CU basis, which indicates whether the CU is using a mode included in the MPM list. If the intra_luma_mpm_flag is true, the video decoding device parses the intra_luma_not_planar_flag to determine whether the current CU is using the planar mode. If the intra_luma_not_planar_flag is true and the CU is using a mode other than the planar mode in the MPM list, the video decoding device parses the flag to determine the applicability of the present invention. In this case, the applicability of the appropriate present invention can be signaled adaptively on a per-CU basis. Syntactically, if both the intra_luma_mpm_flag and the intra_luma_not_planar_flag are 1, additional information is signaled according to this embodiment, which can avoid signaling additional information for all CUs at one time.
[0224] <Embodiment 2-3> Infer the applicability of the present invention and signal the change method for the base block
[0225] In this embodiment, the video decoding device uses high-level information to infer the applicability of the present invention and parses the adaptive_mpm_list_construction_idx that determines the change method for the base block. The index that determines the change method for the base block can be signaled at various levels such as SPS / VPS / PPS / SH (slice header) / CTU (coding tree unit) / CU (coding unit), etc.
[0226] To determine the applicability of the present invention, a video decoding device may consider the following high-level information. The video decoding device may consider information on the video type, such as natural content (NC), screen content (SC), light field content, XR content, point cloud content, etc. Additionally, the video decoding device may consider information on whether compression techniques are allowed, such as sps_ibc_enabled_flag, sps_palette_enabled_flag, sps_bdpcm_enabled_flag, sps_mip_enabled_flag, etc.
[0227] Depending on the video type, additional prediction techniques, such as IBC, palette, BDPCM, MIP, etc., are allowed to be included in the set of non-regular intra modes. Due to the frequent utilization of these modes, simply replacing a non-regular intra mode with one of the regular intra prediction modes may be inefficient, as done with conventional MPM techniques. Accordingly, based on high-level information such as the type of video, additional allowed prediction techniques, etc., the video decoding device can infer the applicability of the present invention. That is, based on the high-level information, it can be determined in advance whether the application of the present invention is advantageous. In this case, no additional information is transmitted, which can improve the coding and decoding efficiency.
[0228] Hereinafter, a method for considering the above high-level information is described in detail.
[0229] The first case is that there is a high-level syntax (sps_video_type) for determining the video type. If the video type is NC (sps_video_type = 0), the video decoding device infers that adaptive_mpm_list_construction_enabled_flag is false. Additionally, if the video type is SC (sps_video_type = 1), the video decoding device may infer that adaptive_mpm_list_construction_enabled_flag is true.
[0230] In one example, as shown in Table 9, the applicability of the present invention can be inferred based on the video type at the SPS level, and the base block change method can be signaled accordingly.
[0231] [Table 9]
[0232] [SPS level]
[0233] sps_video_type if(sps_video_type == 0) sps_adaptive_mpm_list_construction_enabled_flag = 0 else sps_adaptive_mpm_list_construction_enabled_flag = 1 if(sps_adaptive_mpm_list_construction_enabled_flag) sps_adaptive_mpm_list_construction_idx
[0234] In Table 9, the video decoding device parses sps_video_type to determine the video type, and then infers the applicability of the present invention based on the video type. When applying the present invention, the video decoding device further parses sps_adaptive_mpm_list_construction_idx to determine the change method for the base block. According to Table 9, the coding and decoding efficiency can be improved because the applicability of the present invention can be determined for lower levels such as CTU, CU, etc. at once without additional information transmission in higher levels, and the change method for the base block can be determined for lower levels such as CTU, CU, etc. at once with only a small number of bits.
[0235] As another example shown in Table 10, the applicability of the present invention can be inferred based on the video type at the SPS level, and the change method for the base block can be signaled at the CU level.
[0236] [Table 10]
[0237] [SPS level]
[0238] sps_video_type if(sps_video_type == 0) sps_adaptive_mpm_list_construction_enabled_flag = 0 else sps_adaptive_mpm_list_construction_enabled_flag = 1
[0239] [CU level]
[0240]
[0241] In Table 10, the video decoding device parses sps_video_type at the SPS level to determine the video type, and infers the applicability of the present invention based on the video type. When applying the present invention, the video decoding device parses intra_luma_mpm_flag at the CU level to indicate the utilization or non-utilization of the modes included in the MPM list on a per-CU basis. If intra_luma_mpm_flag is true, the video decoding device parses intra_luma_not_planar_flag to determine whether the current CU is using the planar mode. If intra_luma_not_planar_flag is true and a mode other than the planar mode in the MPM list is used, the video decoding device checks the applicability of the present invention determined at a higher level. When applying the present invention, the video decoding device parses the index adaptive_mpm_list_construction_idx indicating the change method for the base block.
[0242] According to Table 10, the encoding and decoding efficiency can be improved because the applicability of the present invention is inferred without the transmission of additional information at a higher level. In addition, when the present invention is applied, an appropriate base block change method can be signaled adaptively on a per-CU basis. Syntactically, if both intra_luma_mpm_flag and intra_luma_not_planar_flag are 1, additional information is signaled according to this embodiment, which avoids signaling additional information for all CUs at once. Table 10 shows an example of inferring the applicability of the present invention at the SPS level and signaling the base block change method at the CU level. However, in various other combinations, the applicability of the present invention can be inferred by using higher-level information and the base block change method can be signaled at a level lower than the higher level.
[0243] In the second case, if, according to the compression technology admissibility information, at least one of the prediction techniques corresponding to the non-regular intra mode is admissible, the video decoding apparatus may infer that adaptive_mpm_list_construction_enabled_flag is true. However, if none of the prediction techniques corresponding to the non-regular intra mode is permitted, the video decoding apparatus may infer that adaptive_mpm_list_construction_enabled_flag is false.
[0244] In one example, as shown in Table 11, the applicability of the present invention can be inferred based on the compression technology admissibility information at the SPS level, and the change method for the base block can be signaled accordingly.
[0245] [Table 11]
[0246] [SPS level]
[0247]
[0248] In Table 11, the video decoding apparatus analyzes the admissibility of the compression technology corresponding to the non-regular intra mode to determine whether it is admissible, and then infers the applicability of the present invention based on whether the compression technology is admissible. When the present invention is applied, the video decoding apparatus further analyzes sps_adaptive_mpm_list_construction_idx to determine the change method for the base block. According to Table 11, the encoding and decoding efficiency can be improved because the applicability of the present invention can be determined once for lower levels such as CTUs and CUs without the transmission of additional information at a higher level, and the change method for the base block can be determined once for lower levels such as CTUs and CUs with only a small number of bits.
[0249] As another example shown in Table 12, the applicability of the present invention can be inferred at the SPS level based on the admissibility of the compression technique, and the base block change method can be signaled at the CU level.
[0250] [Table 12]
[0251] [SPS level]
[0252]
[0253] [CU level]
[0254]
[0255] In Table 12, the video decoding device parses the admissibility of the compression technique corresponding to the non-regular intra mode at the SPS level to determine whether it is admissible, and then infers the applicability of the present invention based on the admissibility of the compression technique. When applying the present invention, the video decoding device parses the intra_luma_mpm_flag that indicates the utilization or non-utilization of the modes included in the MPM list on a per-CU basis. If the intra_luma_mpm_flag is true, the video decoding device parses the intra_luma_not_planar_flag to determine whether the current CU is using the planar mode. If the intra_luma_not_planar_flag is true and the current CU is using a mode other than the planar mode in the MPM list, the video decoding device can confirm the applicability of the present invention determined at a higher level. When applying the present invention, the video decoding device parses the index adaptive_mpm_list_construction_idx that indicates the change method for the base block.
[0256] According to Table 12, the coding and decoding efficiency can be improved because the applicability of the present invention is inferred without the transmission of additional information at a higher level. In addition, when applying the present invention, an appropriate base block change method can be signaled adaptively on a per-CU basis. Syntactically, if both the intra_luma_mpm_flag and the intra_luma_not_planar_flag are 1, additional information is signaled according to this embodiment, thus avoiding signaling additional information for all CUs at once. Table 12 shows an example of inferring the applicability of the present invention at the SPS level and signaling the base block change method at the CU level, but in various other combinations, the applicability of the present invention can be inferred by using higher-level information and the base block change method can be signaled at a level lower than a relatively high level.
[0257] <Embodiment 2-4> Infers the applicability of the present invention and determines the basic block change method as a preset method
[0258] In this embodiment, the video decoding device uses high-level information to infer the applicability of the present invention and determines the basic block change method as a preset method. When determining the applicability of the present invention, the same high-level information exemplified in Embodiment 2-3 can be considered. Accordingly, it can be pre-determined whether it is beneficial to apply the present invention based on the same high-level information in Embodiment 2-3. Therefore, the coding and decoding efficiency can be improved because no additional information is transmitted to determine the applicability of the present invention.
[0259] A method for considering high-level information, such as the method exemplified in Embodiment 2-3, is described in detail.
[0260] In the first case, there is a high-level syntax (sps_video_type) for determining the type of video. If the video type is NC (sps_video_type = 0), the video decoding device infers that adaptive_mpm_list_construction_enabled_flag is false. Additionally, if the video type is SC (sps_video_type = 1), the video decoding device can infer that adaptive_mpm_list_construction_enabled_flag is true.
[0261] In an example shown in Table 13, the applicability of the present invention can be inferred based on the video type at the SPS level, and accordingly the basic block change method can be determined as a preset method.
[0262] [Table 13]
[0263] [SPS level]
[0264] sps_video_type if(sps_video_type == 0) sps_adaptive_mpm_list_construction_enabled_flag = 0 else sps_adaptive_mpm_list_construction_enabled_flag = 1
[0265] In Table 13, the video decoding device parses sps_video_type at the SPS level to determine the video type and infers the applicability of the present invention based on the video type. When applying the present invention, the video decoding device determines the basic block change method as a preset method. According to Table 13, the applicability of the present invention and the basic block change can be determined for lower levels such as CTU and CU at once without additional information transmission at higher levels, thereby improving the coding and decoding efficiency.
[0266] In the second case, if, according to the compression technology admissibility information, at least one of the prediction techniques corresponding to the non-regular intra mode is admissible, the video decoding device may infer that the adaptive_mpm_list_construction_enabled_flag is true. However, if none of the prediction techniques corresponding to the non-regular intra mode is admissible, the video decoding device may infer that the adaptive_mpm_list_construction_enabled_flag is 0.
[0267] In one example shown in Table 14, the applicability of the present invention may be inferred based on the compression technology admissibility information at the SPS level, and the base block change method may be determined as a preset method.
[0268] [Table 14]
[0269] [SPS level]
[0270]
[0271] In Table 14, the video decoding device analyzes the admissibility of the compression technology corresponding to the non-regular intra mode to determine whether it is admissible, and then infers the applicability of the present invention based on the admissibility of the compression technology. When applying the present invention, the video decoding device determines the base block change method as a preset method. According to Table 14, the applicability of the present invention and the base block change can be determined for lower levels such as CTU and CU at one time without additional information transmission from higher levels, thereby improving the coding and decoding efficiency.
[0272] Now refer to Figure 20 and Figure 21 , and a method for generating an MPM list based on a base block that has been changed according to settings at a higher level will be described below. Figure 20 and Figure 21 may be equivalent to Table 4 / Table 6 of Embodiment 2-1, Table 7 of Embodiment 2-2, Embodiment 2-3, and Embodiment 2-4.
[0273] Figure 20 is a flowchart of a method for encoding a current block by a video encoding device according to at least one embodiment of the present invention.
[0274] The video encoding device obtains an adaptive MPM list generation flag from a higher level (S2000). Here, the adaptive MPM list generation flag indicates whether to generate an MPM list based on the change of a base block. In terms of rate distortion optimization, the video encoding device can determine the adaptive MPM list generation flag at a higher level. Alternatively, the video encoding device can use high-level information to derive the adaptive MPM list generation flag at a higher level. Here, the high-level information can be the type of the video or information about the admissibility of a prediction technique using an irregular intra mode.
[0275] The base block can be a block including a pixel to the left of the pixel at the lower left of the current block or a block including a pixel above the pixel at the upper right of the current block.
[0276] The video encoding device encodes the adaptive MPM list generation flag (S2002).
[0277] The video encoding device checks the adaptive MPM list generation flag (S2004).
[0278] If the adaptive MPM list generation flag is true (S2004 is yes), the video encoding device performs the following steps.
[0279] The video encoding device obtains the prediction mode of the base block of the current block (S2006).
[0280] The video encoding device checks the prediction mode of the base block (S2008).
[0281] When the prediction mode of the base block is an irregular intra mode (S2008 is yes), the video encoding device performs the following steps.
[0282] The video encoding device determines a change method for the base block (S2010).
[0283] The video encoding device can determine the change method for the base block from the perspective of optimizing rate distortion. Alternatively, the video encoding device can set the change method for the base block to a predetermined method. The change method for the base block can be one of the methods exemplified in Embodiment 1.
[0284] The video encoding device changes the base block to a block with a regular intra prediction mode according to the change method (S2012).
[0285] The video encoding device generates an MPM list for the current block by using the intra prediction mode of the changed base block (S2014).
[0286] If the change method for the base block has been determined in terms of rate distortion optimization, the video encoding device encodes an index indicating the change method (S2016).
[0287] If the adaptive MPM list generation flag is false (S2004 is NO), the video encoding device generates an MPM list by using the prediction mode of the base block (S2020).
[0288] In addition, if the prediction mode of the base block is not the irregular intra mode (S2008 is NO), the video encoding device generates an MPM list by using the prediction mode of the base block (S2020).
[0289] Then, the video encoding device can generate a predicted block of the current block by using the intra prediction mode determined from the MPM list, and can encode the MPM index indicating the determined intra prediction mode.
[0290] Figure 21 is a flowchart of a method for a video decoding device to reconstruct a current block according to at least one embodiment of the present invention.
[0291] The video decoding device obtains an adaptive MPM list generation flag from a higher level (S2100). Here, the adaptive MPM list generation flag indicates whether to generate an MPM list in response to a change in the base block. The video decoding device can decode the adaptive MPM list generation flag from a bitstream at a higher level. Alternatively, the video decoding device can use high-level information to derive the adaptive MPM list generation flag from a higher level. Here, the high-level information can be the type of video or information about the admissibility of a prediction technique using the irregular intra mode.
[0292] The base block can be a block including pixels to the left of the pixel at the lower left of the current block or a block including pixels above the pixel at the upper right of the current block.
[0293] The video decoding device checks the adaptive MPM list generation flag (S2102).
[0294] If the adaptive MPM list generation flag is true (S2012 is YES), the video decoding device performs the following steps.
[0295] The video decoding device obtains the prediction mode of the base block of the current block (S2104).
[0296] The video decoding device checks the prediction mode of the base block (S2106).
[0297] If the prediction mode of the base block is the irregular intra mode (S2106 is YES), the video decoding device performs the following steps.
[0298] The video decoding device obtains the change method for the base block (S2108).
[0299] The video decoding apparatus may decode, from a bitstream, a change method for a base block. Alternatively, the video decoding apparatus may set a change method for a base block to a predetermined method. The change method for the base block may be one of the methods exemplified in Embodiment 1.
[0300] The video decoding apparatus changes the base block into a block having a regular intra prediction mode according to the change method (S2110).
[0301] The video decoding apparatus generates an MPM list for a current block by using the intra prediction mode of the changed base block (S2112).
[0302] On the other hand, if the adaptive MPM list generation flag is false (No in S2102), the video decoding apparatus generates an MPM list by using the prediction mode of the base block (S2120).
[0303] In addition, if the prediction mode of the base block is not an irregular intra mode (No in S2106), the video decoding apparatus generates an MPM list by using the prediction mode of the base block (S2120).
[0304] Thereafter, the video decoding apparatus may decode an MPM index from the bitstream and may generate a prediction block for the current block by using the intra prediction mode determined from the MPM list according to the MPM index.
[0305] Reference Figure 22 and Figure 23 , a method of generating an MPM list based on a base block whose change is set according to each CU will be described below. Figure 22 and Figure 23 The diagrams of may be equivalent to Table 5 of Embodiment 2-1 and Table 8 of Embodiment 2-2.
[0306] Figure 22 is a flowchart of a method of encoding a current block by a video encoding apparatus according to another embodiment of the present invention.
[0307] The video encoding apparatus obtains a base block of a current block (S2200). Here, the base block may be a block including pixels on the left side of the lower left pixel of the current block or a block including pixels on the upper side of the upper right pixel of the current block.
[0308] The video encoding apparatus generates a first MPM list for the current block by using the prediction mode of the base block (S2202).
[0309] The video encoding apparatus checks the prediction mode of the base block (S2204).
[0310] If the prediction mode of the base block is an irregular intra mode (Yes in S2204), the video encoding apparatus performs the following steps.
[0311] The video encoding device determines a change method for the base block (S2206).
[0312] In terms of rate - distortion optimization, the video encoding device can determine a change method for the base block. Alternatively, the video encoding device can set the change method for the base block to a predetermined method. The change method for the base block can be one of the methods exemplified in Embodiment 1.
[0313] The video encoding device changes the base block to a block with a regular intra - prediction mode according to the change method (S2208).
[0314] The video encoding device generates a second MPM list for the current block by using the intra - prediction mode of the changed base block (S2210).
[0315] On the other hand, if the prediction mode of the base block is not an irregular intra - mode (S2204 is no), the above steps of generating the second MPM list can be omitted.
[0316] The video encoding device determines an adaptive MPM list generation flag based on the first MPM list and the second MPM list (S2212). Here, the adaptive MPM list generation flag indicates whether to generate the second MPM list in response to the change of the base block. In terms of rate - distortion optimization, the video encoding device can determine the adaptive MPM list generation flag on a per - CU basis. For example, if the first MPM list is the best, the adaptive MPM list generation flag can be determined to be false. On the other hand, if the second MPM list is the best, the adaptive MPM list generation flag can be determined to be true. In addition, if the second MPM list is not generated, the adaptive MPM list generation flag can be determined to be false.
[0317] The video encoding device encodes the adaptive MPM list generation flag (S2214).
[0318] The video encoding device checks the adaptive MPM list generation flag (S2216).
[0319] If the adaptive MPM list generation flag is true (S2216 is yes) and a change method for the base block is determined in terms of rate - distortion optimization, the video encoding device encodes an index indicating the change method for the base block (S2218).
[0320] Then, the video encoding device can generate a prediction block for the current block by using the intra - prediction mode determined from the MPM list, and can encode an MPM index indicating the determined intra - prediction mode.
[0321] Figure 23A flowchart of a method for reconstructing a current block by a video decoding device according to another embodiment of the present invention.
[0322] The video decoding device obtains an adaptive MPM list generation flag for the current block (S2300). Here, the adaptive MPM list generation flag indicates whether to generate an MPM list in response to a change in the base block. The base block may be a block including pixels to the left of the pixel at the lower left of the current block or a block including pixels above the pixel at the upper right of the current block.
[0323] The video decoding device checks the adaptive MPM list generation flag (S2302).
[0324] If the adaptive MPM list generation flag is true (S2302 is yes), the video decoding device performs the following steps.
[0325] The video decoding device obtains a base block with an irregular intra mode (S2304). Since the per-CU adaptive MPM list generation flag is true, a base block with an irregular intra mode can be obtained for the current block.
[0326] The video decoding device obtains a change method for the base block (S2306).
[0327] The video decoding device may decode the change method for the base block from the bitstream. Alternatively, the video decoding device may set the change method for the base block to a predetermined method. The change method for the base block may be one of the methods exemplified in Embodiment 1.
[0328] The video decoding device changes the base block to a block with a regular intra prediction mode according to the change method (S2308).
[0329] The video decoding device generates an MPM list for the current block by using the intra prediction mode of the changed base block (S2310).
[0330] On the other hand, if the adaptive MPM list generation flag is false (S2302 is no), the video decoding device generates an MPM list by using the prediction mode of the base block (S2320).
[0331] Thereafter, the video decoding device may decode an MPM index from the bitstream and may generate a prediction block for the current block by using the intra prediction mode determined from the MPM list based on the MPM index.
[0332] Although the steps in the various flowcharts described are presented as sequential, these steps merely illustrate the technical ideas of some embodiments of the present invention. Accordingly, those of ordinary skill in the art to which the present invention pertains can perform the steps by changing the order described in the respective figures or by performing two or more steps in parallel. Thus, the steps in the respective flowcharts are not limited to the order shown as occurring in time.
[0333] It should be understood that the above description presents illustrative embodiments that can be implemented in various other ways. The functions described in some embodiments can be implemented by hardware, software, firmware, and / or combinations thereof. It should also be understood that the functional components described in the present invention are labeled as "…… unit" to highlight the possibility of their independent implementation.
[0334] On the other hand, the various methods or functions described in some embodiments can be implemented as instructions stored in a non - volatile recording medium, which can be read and executed by one or more processors. The non - volatile recording medium can include various types of recording devices that store data in a form readable by a computer system. For example, the non - volatile recording medium can include storage media such as erasable programmable read - only memory (EPROM), flash drives, optical disk drives, magnetic hard disk drives, and solid - state drives (SSD), etc.
[0335] Although the exemplary embodiments of the present invention have been described for illustrative purposes, those of ordinary skill in the art to which the present invention pertains should understand that various modifications, additions, and substitutions can be made without departing from the spirit and scope of the present invention. Accordingly, the embodiments of the present invention have been described for the sake of brevity and clarity. The scope of the technical ideas of the embodiments of the present invention is not limited by the examples. Correspondingly, those of ordinary skill in the art to which the present invention pertains should understand that the scope of the present invention should not be limited by the embodiments explicitly described above, but rather by the claims and their equivalents.
[0336] Reference Numerals
[0337] 122: Intra Predictor
[0338] 155: Entropy Encoder
[0339] 510: Entropy Decoder
[0340] 542: Intra Predictor.
[0341] Cross - Reference to Related Applications
[0342] This application claims priority and the benefit of Korean Patent Application No. 10-2022-0168825, filed on December 6, 2022, and Korean Patent Application No. 10-2023-0143180, filed on October 24, 2023, the entire contents of each of which are incorporated herein by reference.
Claims
1. A method for reconstructing a current block by a video decoding device, the method comprising: Obtaining a prediction mode of a base block of the current block, wherein the base block is a block including pixels to the left of the pixel at the lower left of the current block or a block including pixels above the pixel at the upper right of the current block; and Checking the prediction mode of the base block, wherein, when the prediction mode of the base block is a non-regular intra mode, the method further comprises: Obtaining a modification method for the base block; Modifying the base block into a block having a regular intra prediction mode according to the modification method; and Generating a most probable mode list (MPM list) of the current block by using the regular intra prediction mode of the modified base block.
2. The method according to claim 1, wherein, The modification method comprises: Searching adjacent pixels of the current block according to a predetermined search order, thereby excluding the pixels included in the base block; and Modifying the base block into a block including a first pixel predicted according to the regular intra prediction mode.
3. The method according to claim 1, wherein, The modification method comprises: Searching, on one side of the base block, for a block having a regular intra prediction mode; and Modifying the base block into the block having the longest length adjacent to the current block among the searched blocks.
4. The method according to claim 1, wherein, The modification method comprises: Searching, on one side of the base block, for a block having a regular intra prediction mode; and Modifying the base block into the block having the largest area among the searched blocks.
5. The method according to claim 1, wherein The modification method comprises: Searching, on one side of the base block, for a block having a regular intra prediction mode; and Modifying the base block into the block having the same or most similar aspect ratio to that of the current block among the searched blocks.
6. The method according to claim 1, wherein, Obtaining the modification method comprises: Decoding, from a bitstream, an index indicating the modification method; and Setting the modification method as the method indicated by the index.
7. The method according to claim 1, wherein Obtaining the modification method comprises: Setting the modification method as a preset modification method.
8. The method according to claim 1, further comprising: Obtaining an adaptive MPM list generation flag, the adaptive MPM list generation flag indicating whether to generate an MPM list in response to the modification of the base block; and Checking the adaptive MPM list generation flag, wherein, when the adaptive MPM list generation flag is true, the method further comprises: Obtaining the base block and checking the prediction mode of the base block.
9. The method according to claim 8, further comprising, when the adaptive MPM list generation flag is false: Generating an MPM list by using the prediction mode of the base block.
10. The method according to claim 8, wherein, Obtaining the adaptive MPM list generation flag comprises: Decoding the adaptive MPM list generation flag from a bitstream.
11. The method according to claim 8, wherein, Obtaining the adaptive MPM list generation flag comprises: Deriving the adaptive MPM list generation flag by using high-level information, the high-level information being the type of the video or information on the admissibility of a prediction technique using a non-regular intra mode.
12. The method according to claim 1, further comprising: When the prediction mode of the base block is not a non-regular intra mode: Generating an MPM list by using the prediction mode of the base block.
13. The method according to claim 1, wherein The non-regular intra mode is a prediction mode other than the regular intra prediction mode, which includes a non-angle prediction mode and an angle prediction mode, and the non-angle prediction mode includes a planar mode and a DC mode.
14. A method for encoding a current block by a video encoding device, the method comprising: Obtain the prediction mode of the base block of the current block, where the base block is a block including pixels to the left of the pixel at the lower left of the current block or a block including pixels above the pixel at the upper right of the current block; and Check the prediction mode of the base block, where, when the prediction mode of the base block is a non-regular intra mode, the method further includes: Determine a change method for the base block; Change the base block to a block with a regular intra prediction mode according to the change method; and Generate the most probable mode list (MPM list) of the current block by using the regular intra prediction mode of the changed base block.
15. The method according to claim 14, further including: Obtain an adaptive MPM list generation flag from a higher level, the adaptive MPM list generation flag indicating whether to generate the MPM list in response to the change of the base block; and Encode the adaptive MPM list generation flag.
16. The method according to claim 15, further including: Check the adaptive MPM list generation flag, where, when the adaptive MPM list generation flag is true, the method further includes: Obtain the prediction mode of the base block and check the prediction mode of the base block.
17. The method according to claim 14, further including: When the prediction mode of the base block is a non-regular intra mode and the change method is determined in terms of rate-distortion optimization, encode the index indicating the change method.
18. A computer-readable recording medium storing a bitstream generated by a video coding method, the video coding method including: Obtain the prediction mode of the base block of the current block, where the base block is a block including pixels to the left of the pixel at the lower left of the current block or a block including pixels above the pixel at the upper right of the current block; and Check the prediction mode of the base block, where, when the prediction mode of the base block is a non-regular intra mode, the video coding method further includes: Determine a change method for the base block; Change the base block to a block with a regular intra prediction mode according to the change method; and Generate the most probable mode list (MPM list) of the current block by using the regular intra prediction mode of the changed base block.
Citation Information
Patent Citations
Switching Converter for Accelerated Dynamic Voltage Scaling and Method for Controlling the same
KR1020220168825A
Integrated power module
KR1020230143180A