Video coding method, video coding apparatus, bit stream storage method, video decoding method, video decoding apparatus, and storage medium

By generating independent intra-predictions for sub-partitions and applying weighted averages, the method addresses inefficiencies in VVC's intra prediction, enhancing coding efficiency and hardware compatibility for video encoding.

JP2025109945AActive Publication Date: 2025-07-25BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025085128
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-02-05
Filing Date
2025-05-21
Publication Date
2025-07-25
Estimated Expiration
2040-02-05

AI Technical Summary

Technical Problem

Existing video encoding standards like VVC struggle with inefficient intra prediction methods, particularly in handling high-resolution video content, leading to suboptimal coding efficiency and hardware implementation challenges.

Method used

The proposed method involves generating independent intra-predictions for each sub-partition of an encoding block, using only a subset of possible intra-prediction modes, and applying weighted averages for combined predictions, while simplifying the intra-subdivision encoding mode to improve coding efficiency and hardware compatibility.

Benefits of technology

This approach enhances coding efficiency and visual quality by optimizing intra-prediction processes, reducing computational complexity, and improving hardware implementation of video encoding systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025109945000001_ABST
    Figure 2025109945000001_ABST
Patent Text Reader

Abstract

To provide a video coding method.SOLUTION: When a determination is made that multiple sub partitions satisfy a first condition of an intra prediction, the intra prediction is executed for the multiple sub partitions of a coded block in an intra sub partition (ISP) mode. Executing the intra prediction for the multiple sub partitions of the coded block includes merging at least two sub partitions of the multiple sub partitions into one prediction region of the intra prediction, and generating prediction samples in that one prediction region of the intra prediction on the basis of multiple reference samples of the coded block adjacent to that one prediction region. The first condition includes that the multiple sub-partitions comprise multiple vertical sub-partitions and each of the sub partitions has a width less than or equal to 2.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to video encoding and compression. Specifically, the present disclosure relates to systems and methods for performing video encoding using an intra-subdivision encoding mode. More specifically, it relates to a video decoding method, a video decoding apparatus, and a storage medium.

Background Art

[0002] This section provides background information related to the present disclosure. The information included in this section should not necessarily be construed as prior art.

[0003] To compress video data, any of various video encoding techniques can be used. Video encoding can be performed according to one or more video encoding standards. Some exemplary video encoding standards include VVC (versatile video coding), JEM (joint exploration test model) encoding, H.265 / HEVC (high-efficiency video coding), H.264 / AVC (advanced video coding), and MPEG (moving picture experts group) encoding.

[0004] Video encoding generally utilizes prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit redundancies inherent in a video image or sequence. One goal of video encoding techniques is to compress video data into a form that uses a lower bitrate while avoiding or minimizing degradation of video quality.

[0005] The first version of the HEVC standard, completed in October 2013, provides about 50% bitrate savings or equivalent perceptual quality compared to the previous-generation video coding standard (H.264 / MPEG AVC). The HEVC standard provides significant coding improvements compared to previous technologies, but there is evidence that even better coding efficiency than HEVC can be achieved by using additional coding tools. Based on this evidence, both the Video Coding Experts Group (VCEG) and the Moving Picture Experts Group (MPEG) initiated exploratory research to develop new coding technologies for future video coding standardization. The Joint Video Exploration Team (JVET) was established in October 2015 by ITU-T VECG and ISO / IEC MPEG and initiated important research on advanced technologies that enable significant improvements in coding efficiency. One reference software model called the Joint Exploration Model (JEM) was maintained by JVET by integrating some additional coding tools on top of the HEVC Test Model (HM).

[0006] In October 2017, ITU-T and ISO / IEC issued a Call for Proposals (CfP) on video compression with capabilities beyond HEVC. In April 2018, 23 CfP responses were received and evaluated at the 10th JVET. These responses showed about 40% compression efficiency gain over the HEVC standard. Based on such evaluation results, JVET launched a new project towards the development of the next-generation video coding standard, "Versatile Video Coding (VVC)". Also in April 2018, one reference software codebase called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard.

[0007] This section provides a general overview of the present disclosure and is not an all-inclusive disclosure of its full scope or all of its features.

Summary of the Invention

Means for Solving the Problem

[0008] According to a first aspect of the present disclosure, a video encoding method is executed in a computing device having one or more processors and a memory storing a plurality of programs executed by the one or more processors. The method includes independently generating respective intra-predictions for each of a plurality of corresponding sub-partitions, each intra-prediction being generated using a plurality of reference samples from a current encoding block.

[0009] According to a second aspect of the present disclosure, a video encoding method is executed in a computing device having one or more processors and a memory storing a plurality of programs executed by the one or more processors. The method includes generating respective intra-predictions for each of a plurality of corresponding sub-partitions for the luminance component of an intra-sub-partition (ISP) encoding block using only N of M possible intra-prediction modes, where M and N are positive integers and N is less than M.

[0010] According to a third aspect of the present disclosure, a video encoding method is executed in a computing device having one or more processors and a memory storing a plurality of programs executed by the one or more processors. The method includes generating an intra-prediction for the chrominance component of an intra-sub-partition (ISP) encoding block using only N of M possible intra-prediction modes, where M and N are positive integers and N is less than M.

[0011] According to a fourth aspect of the present disclosure, a video encoding method is executed in a computing device having one or more processors and a memory storing a plurality of programs executed by the one or more processors. The method generates respective luminance intra-predictions for each of a plurality of corresponding sub-divisions of an entire intra-subdivision (ISP) encoding block for a luminance component and for a chrominance component, and generates a chrominance intra-prediction for the entire intra-subdivision (ISP) encoding block.

[0012] According to a fifth aspect of the present disclosure, a video encoding method is executed in a computing device having one or more processors and a memory storing a plurality of programs executed by the one or more processors. The method generates a first prediction using an intra-subdivision mode, generates a second prediction using an inter-prediction mode, and combines the first prediction and the second prediction to generate a final prediction by applying a weighted average to the first prediction and the second prediction.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 6C

Figure 7

Figure 8A

Figure 8B

Figure 8C

Figure 9A

Figure 9B

Figure 10

Figure 11

Figure 12

DETAILED DESCRIPTION OF THE INVENTION

[0014] Hereinafter, a set of exemplary and non-limiting embodiments of the present disclosure will be described in conjunction with the accompanying drawings. Modifications to the structure, method, or function may be implemented by those skilled in the relevant art based on the examples presented herein, and all such modifications are included within the scope of the present disclosure. If there are no contradictions, the teachings of different embodiments can be combined with each other, but this is not necessarily the case.

[0015] The terms used in the present disclosure are not intended to limit the present disclosure but to describe specific examples. The singular forms such as "one" used in the present disclosure and the appended claims also refer to the plural form unless the context clearly dictates otherwise. It should be understood that the term "and / or" used herein refers to any or all possible combinations of one or more of the associated listed items.

[0016] Here, various information can be described using terms such as "first", "second", "third", etc., but the information should not be limited by these terms. These terms are only used to distinguish one category of information from another. For example, without departing from the scope of the present disclosure, the first information may be referred to as the second information, and similarly, the second information may be referred to as the first information. As used herein, the term "if" means "when", "upon", or "in response to" depending on the context.

[0017] Throughout this specification, references to one or more "embodiments", "an embodiment", "another embodiment", etc. mean that one or more specific features, structures, or characteristics described in connection with an embodiment are included in at least one embodiment of the present disclosure. Thus, the appearances of the phrases one or more "in an embodiment", "in an embodiment", "in another embodiment", etc. at various positions throughout this specification do not necessarily all refer to the same embodiment. Furthermore, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable manner.

[0018] Conceptually, many video coding standards are similar as described above. For example, substantially all video coding specifications use block-based processing and share similar video coding block diagrams to achieve video compression. Similar to HEVC, the VVC standard is built on a block-based hybrid video coding framework.

[0019] FIG. 1 shows a block diagram of an exemplary encoder 100 that can be used in conjunction with many video coding standards. In encoder 100, a video frame is divided into a plurality of video blocks for processing. For each given video block, a prediction is formed based on either an inter-prediction approach or an intra-prediction approach. In inter-prediction, one or more predictors are formed by motion estimation and motion compensation based on pixels from previously reconstructed frames. In intra-prediction, the predictor is formed based on reconstructed pixels within the current frame. By mode decision, the best predictor can be selected to predict the current block.

[0020] A prediction residual representing the difference between the current video block and its predictor is sent to the transform circuit 102. Then, the transform coefficients are sent from the transform circuit 102 to the quantization circuit 104 for entropy reduction. Then, the quantized coefficients are supplied to the entropy coding circuit 106 to generate a compressed video bitstream. As shown in FIG. 1, prediction-related information 110 from the inter-prediction circuit and / or the intra-prediction circuit 112, such as video block partition information, motion vectors, reference picture indices, and intra-prediction modes, is also supplied through the entropy coding circuit 106 and stored in the compressed video bitstream 114.

[0021] In the encoder 100, in order to reconstruct pixels for prediction purposes, decoder-related circuits are also required. First, the prediction residual is reconstructed through the inverse quantization circuit 116 and the inverse transform circuit 118. This reconstructed prediction residual is combined with the block predictor 120 to generate non-filtered reconstructed pixels for the current video block.

[0022] To improve the coding efficiency and visual quality, an in-loop filter 115 is generally used. For example, the deblocking filter is available in the current versions of AVC, HEVC, and VVC. In HEVC, an additional in-loop filter called SAO (Sample Adaptive Offset) is defined to further improve the coding efficiency. In the current VVC standard, yet another in-loop filter called ALF (Adaptive Loop Filter) is actively being considered and is likely to be included in the final standard.

[0023] These in-loop filter operations are optional. Performing these operations helps to improve the coding efficiency and visual quality. They can also be turned off as a decision by the encoder 100 to save computational complexity.

[0024] Note that intra prediction is usually based on non-filtered reconstructed pixels, while inter prediction is based on filtered reconstructed pixels when these filter options are turned on by the encoder 100.

[0025] FIG. 2 is a block diagram showing an exemplary decoder 200 that can be used with many video coding standards. This decoder 200 is similar to the reconstruction-related part present in the encoder 100 of FIG. 1. In decoder 200 (FIG. 2), the input video bitstream 201 is first decoded through an entropy decoding circuit 202 to derive quantization coefficient levels and prediction-related information. The quantized coefficient levels are processed through an inverse quantization circuit 204 and an inverse transform circuit 206 to obtain a reconstructed prediction residual. The block prediction mechanism implemented by the intra / inter mode selector 212 is configured to execute either an intra prediction procedure 208 or a motion compensation procedure 210 based on the decoded prediction information. By using an adder 214 to sum the reconstructed prediction residual from the inverse transform circuit 206 and the prediction output generated by the block predictor mechanism, a set of unfiltered reconstructed pixels is obtained. The reconstructed blocks can further pass through an in-loop filter 209 before being stored in a picture buffer 213 that functions as a reference picture store. The reconstructed video in the picture buffer 213 is then sent to drive a display device and can be used to predict future video blocks. In situations where the in-loop filter 209 is turned on, filtering operations are performed on these reconstructed pixels to derive the final reconstructed video output 222.

[0026] Returning to FIG. 1 here, the input video signal to the encoder 100 is processed block by block. Each block is called a coding unit (CU). In VTM-1.0, the CU can be up to 128×128 pixels. In HEVC (High Efficiency Video Encoding), JEM (Joint Exploration Test Model), and VVC (Versatile Video Encoding), the basic unit of compression is called a coding tree unit (CTU). However, in contrast to the HEVC standard that divides blocks based only on a quadtree, in the VVC standard, one CTU is divided into CUs to adapt to local characteristics that vary based on a quadtree / binary tree / trinary tree structure. Furthermore, the concept of multiple partition unit types in the HEVC standard does not exist in the VVC standard. That is, the separation of CU, PU (Prediction Unit), and TU (Transform Unit) does not exist in the VVC standard. Instead, each CU is always used as the basic unit for both prediction and transformation without further partitioning. The maximum CTU size of HEVC and JEM is defined as a maximum of 64×64 luma pixels and, in the case of 4:2:0 chroma format, two blocks of 32×32 chroma pixels. The maximum allowable size of the luma block within the CTU is specified as 128×128 (although the maximum size of the luma transform block is 64×64).

[0027] FIG. 3 shows five exemplary block partitions for a multi-type tree structure. The five exemplary block partitions include a four-way split 301, a horizontal two-way split 302, a vertical two-way split 303, a horizontal three-way split 304, and a vertical three-way split 305. In a situation where a multi-type tree structure is utilized, one CTU is first divided by a quadtree structure. Next, each quadtree leaf node can be further divided by binary tree and trinary tree structures.

[0028] One or more of the exemplary block partitions 301, 302, 303, 304, or 305 of FIG. 3 can be used to perform spatial prediction and / or temporal prediction using the configuration shown in FIG. 1. Spatial prediction (or "intra prediction") uses pixels from samples of already encoded adjacent blocks (referred to as reference samples) within the same video picture / slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal.

[0029] Temporal prediction (also referred to as "inter prediction" or "motion compensated prediction") uses reconstructed pixels from already encoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Also, if multiple reference pictures are supported, one reference picture index is further transmitted, and the reference picture index is used to identify from which reference picture in the reference picture store the temporal prediction signal comes.

[0030] After spatial and / or temporal prediction is performed, the intra-mode / inter-mode decision circuit 121 within the encoder 100 selects the best prediction mode, for example, based on the rate-distortion optimization method. Next, the block predictor 120 subtracts from the current video block, and the resulting prediction residual is decorrelated using the transform circuit 102 and quantization circuit 104. The obtained quantized residual coefficients are inverse quantized by the inverse quantization circuit 116, inverse transformed by the inverse transform circuit 118 to form a reconstructed residual, and then returned to the prediction block to form the reconstructed signal of the CU. Further in-loop filters 115, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive loop filter (ALF), can be applied to the reconstructed CU before the reconstructed CU is placed in the reference picture store of the picture buffer 117 and used to encode future video blocks. For forming the output video bitstream 114, all of the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are sent to the entropy coding unit 106, further compressed, and packed to form a bitstream.

[0031] The basic intra prediction method applied in the VVC standard remains generally the same as the HEVC standard, except that several modules are further extended and / or improved in the VVC standard, such as the intra sub-partition (ISP) coding mode, extended intra prediction by the intra wide-angle direction, position-dependent intra prediction combination (PDPC), and 4-tap intra interpolation. One broad aspect of the present disclosure aims to improve the existing ISP design in the VVC standard. Further, other coding tools included in the VVC standard and closely related to the techniques proposed in the present disclosure (e.g., tools in the intra prediction and transform coding processes) are described in more detail below.

[0032] [Intra Prediction Mode with Intra Wide-Angle Direction] Similar to the case of the HEVC standard, the VVC standard uses a set of previously decoded samples adjacent to (i.e., above or to the left of) one current CU to predict the samples of the CU. However, to capture the finer edge directions present in natural videos (especially for video content with high resolution, e.g., 4K), the number of intra angle modes is extended from 33 modes in the HEVC standard to 93 modes in the VVC standard. In addition to the angular directions, both the HEVC standard and the VVC standard provide for a planar mode (assuming a smoothly varying surface with horizontal and vertical slopes derived from the boundaries) and a DC mode (assuming a flat surface).

[0033] FIG. 4 illustrates an exemplary set of intra modes 400 for use with the VVC standard, and FIG. 5 illustrates a plurality of reference lines for performing intra prediction. Referring to FIG. 4, the exemplary set of intra modes 400 includes modes 0, 1, -14, -12, -10, -8, -6, -4, -2, 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, and 80. Mode 0 corresponds to the planar mode and mode 1 corresponds to the DC mode. Similar to the intra prediction process of the HEVC standard, all of the defined intra modes (i.e., planar, DC, and angular directions) in the VVC standard utilize a set of adjacent reconstructed samples above and to the left of the predicted block as references for intra prediction. However, unlike the HEVC standard where only the closest row / column of the reconstructed samples (row 0, 501 in FIG. 5) is used as a reference, a multi-reference line (MRL) is introduced in the VVC where two additional rows / columns (i.e., row 1, 503 and row 3, 505 in FIG. 5) are used for the intra prediction process. The index of the selected reference row / column is signaled from the encoder 100 (FIG. 1) to the decoder 200 (FIG. 2). If a row / column that is not the closest in FIG. 5, such as row 1, 503 or row 3, 505, is selected, the planar mode and the DC mode in FIG. 4 are excluded from the set of intra modes that can be used to predict the current block.

[0034] Figure 6A illustrates a reference sample and a first set of angular directions 602 and 604 used for intra prediction of a rectangular block (the value obtained by dividing the width W by the height H is equal to 2). The first set of reference samples includes a first sample 601 and a second sample 603. Figure 6B shows a second set of reference samples and angular directions 606 and 608 used for intra prediction of a vertically long rectangular block (W / H = 1 / 2). The second set of reference samples includes a third sample 605 and a fourth sample 607. Figure 6C shows a third set of reference samples and angular directions 610 and 612 used for intra prediction of a square block (W = H). The third set of reference samples includes a fifth sample 609 and a sixth sample 610. Assuming the use of the nearest neighbor, Figure 6C shows the positions of the third reference samples that can be used in the VVC standard to derive the predicted samples of one intra block. As shown in Figure 6C, due to the application of the quadtree / binary tree / trinary tree splitting structure, in addition to square coded blocks, rectangular coded blocks also exist in the intra prediction procedure in the context of the VVC standard.

[0035] Due to the unequal width and height of one given block, different sets of angular directions are selected for different block shapes, which is also called wide-angle intra prediction. Specifically, in addition to the planar and DC modes, for both square and rectangular coded blocks, as shown in Table 1, 65 out of 93 angular directions are also supported for each block shape. Such a design can not only efficiently capture the directional structures typically present in videos (by adaptively selecting the angular directions based on the block shape), but also ensure that a total of 67 intra modes (i.e., planar, DC, and 65 angular directions) are effectively available for each coded block. This can achieve good efficiency of the signal intra mode while providing a consistent design across different block sizes.

[0036]

Table 1

[0037] [Position-Dependent Intra Prediction Combining] As described above, the intra prediction samples are generated from either the set of unfiltered or filtered neighboring reference samples, which may introduce discontinuities along the block boundaries between the current coding block and its neighborhood. To solve such problems, boundary filtering is applied in the HEVC standard by combining the first row / column of the prediction samples in the DC, horizontal (i.e., mode 18 in FIG. 4) and vertical (i.e., mode 50) prediction modes with the unfiltered reference samples, by using a 2-tap filter (for the DC mode) or a gradient-based smoothing filter (for the horizontal and vertical prediction modes).

[0038] The position-dependent intra prediction combining (PDPC) tool in the VVC standard extends the aforementioned concept by using a weighted combination of the intra prediction samples and the unfiltered reference samples. In the current VVC working draft, PDPC is enabled for the following intra modes without signaling, i.e., planar, DC, horizontal (i.e., mode 18), vertical (i.e., mode 50), angular directions close to the bottom-left diagonal direction (i.e., modes 2, 3, 4, …, 10), and angular directions close to the top-right diagonal direction (i.e., modes 58, 59, 60, …, 66). Assuming that the prediction sample located at the coordinates (x, y) is pred(x, y), its corresponding value after PDPC is calculated as follows. pred(x,y)=(wL×R -1,y +wT×R x-1 -wTL×R -1,-1 +(64-wL-wT+wTL)×pred(x,y)+32)>>6 … (1) where, R x-1, R -1,y represents the reference samples above and to the left of the current sample (x, y), respectively, and R -1,-1 represents the reference sample at the upper left corner of the current block.

[0039] Figure 7 shows an exemplary set of positions for adjacent reconstructed samples used for position-dependent intra prediction combination (PDPC) of one coded block. The first reference sample 701 (R x-1 ) represents the reference sample located above the current predicted sample (x, y). The second reference sample 703 (R -1,y ) represents the reference sample located to the left of the current predicted sample (x, y). The third reference sample 705 (R -1,-1 ) represents the reference sample located at the upper left corner of the current predicted sample (x, y).

[0040] The reference samples including the first, second, and third reference samples 701, 703, and 705 are combined with the current predicted sample (x, y) during the PDPC process. Assuming that the current coded block has a size of W×H, as follows, the weights wL, wT, and wTL in Equation (1) are adaptively selected according to the prediction mode and sample position.

[0041] In the DC mode, wT = 32 >> ((y << 1) >> shift), wL = 32 >> ((x << 1) >> shift), wTL = (wL >> 4) + (wT >> 4)…(2)

[0042] In the planar mode, wT = 32 >> (y << 1) >> shift), wL = 32 >> ((x << 1) >> shift), wTL = 0…(3)

[0043] In the horizontal mode, wT = 32 >> ((y << 1) >> shift), wL = 32 >> ((x << 1) >> shift), wTL = wT …(4)

[0044] In the vertical mode, wT = 32 >> ((y << 1) >> shift), wL = 32 >> ((x << 1) >> shift), wTL = wL …(5)

[0045] In the lower left diagonal direction, wT = 16 >> ((y << 1) >> shift), wL = 16 >> ((x << 1) >> shift), wTL = 0 …(6)

[0046] In the upper right diagonal direction, wT = 16 >> ((y << 1) >> shift), wL = 16 >> ((x << 1) >> shift), wTL = 0 …(7)

[0047] Here, shift = (log2(W) - 2 + log2(H) - 2 + 2) >> 2

[0048] 〔Multiple Transform Selection and Shape-Adaptive Transform Selection〕

[0049] In addition to the DCT-II transform used in the HEVC standard, by introducing additional core transforms of DCT-VIII and DST-VII, the Multiple Transform Selection (MTS) tool becomes available in the VVC standard. In the VVC standard, the adaptive selection of transforms is made possible at the coded block level by signaling one MTS flag in the bitstream. Specifically, when the MTS flag is equal to 0 for a block, a pair of fixed transforms (e.g., DCT-II) are applied in the horizontal and vertical directions. Otherwise (when the MTS flag is equal to 1), two additional flags are further signaled for the block to indicate the transform type (either DCT-VIII or DST-VII) for each direction.

[0050] On one hand, due to the introduction of the quadtree / binary tree / tritree-based block partitioning structure in the VVC standard, the distribution of the intra prediction residuals is strongly correlated with the block shape. Therefore, when MTS is disabled (i.e., the MTS flag is equal to 0 for one coded block), one shape-adaptive transform selection method is applied to all intra-coded blocks for which the DCT-II and DST-VII transforms are implicitly enabled based on the width and height of the current block. More specifically, for each rectangular block, this method uses the DST-VII transform in the direction related to the short side of one block and the DCT-II transform in the direction related to the long side of the block. For each square block, DST-VII is applied in both directions. Furthermore, to avoid introducing new transforms for different block sizes, the DST-VII transform is only enabled when the short side of one intra-coded block is 16 or less. Otherwise, the DCT-II transform is always applied.

[0051] Table 2 shows the enabled horizontal and vertical transforms for intra-coded blocks based on the shape-adaptive transform selection method in VVC.

[0052]

Table 2

[0053] 〔Intra Sub-Partition Coding Mode〕 The conventional intra mode generates the intra prediction samples of a block by using only the reconstructed samples adjacent to one coded block. The spatial correlation between the prediction samples and the reference samples based on such a method is approximately proportional to the distance between the prediction samples and the reference samples. Therefore, the samples in the inner part (especially the samples located at the lower right corner of the block) usually have a lower prediction quality than the samples near the block boundary. To further improve the intra prediction efficiency, short-distance intra prediction (SDIP) has been proposed and well studied. This method divides one intra-coded block into a plurality of sub-blocks in the horizontal or vertical direction for prediction. Usually, a square block is divided into four sub-blocks. For example, an 8×8 block can be divided into four 2×8 or four 8×2 sub-blocks. One extreme case of such sub-block-based intra prediction is the so-called line-based prediction, where the block is divided into one-dimensional rows / columns for prediction. For example, one W×H (width × height) block can be divided into H sub-blocks of size W×1 or W sub-blocks of size 1×H for intra prediction. Each of the obtained rows / columns is encoded in the same way as a normal two-dimensional (2D) block (as shown in FIGS. 6A, 6B, 6C, and 7), i.e., predicted by one of the available intra modes, the prediction error is decorrelated based on transformation and quantization, and is sent to the decoder 200 (FIG. 2) for reconstruction. As a result, the reconstructed samples within one sub-block (e.g., a row / column) can be used as a reference for predicting the samples within the next sub-block. The above process is repeated until all sub-blocks within the current block are predicted and encoded. Further, to reduce the signaling overhead, all sub-blocks within one coded block share the same intra mode.

[0054] In SDIP, different sub-block partitions may provide different encoding efficiencies. Generally, line-based prediction provides the "shortest prediction distance" among different partitions, thus providing the best encoding efficiency. On the other hand, codec hardware implementations have the worst encoding / decoding throughput. For example, considering a block with 4×4 sub-blocks and the same block with 4×1 or 1×4 sub-blocks, the throughput in the latter case is only one-fourth of that in the former case. In HEVC, the minimum intra prediction block size for luminance is 4×4.

[0055] FIG. 8A shows an exemplary set of short-distance intra prediction (SDIP) partitions for an 8×4 block 801, FIG. 8B shows an exemplary set of short-distance intra prediction (SDIP) partitions for a 4×8 block 803, and FIG. 8C shows an exemplary set of short-distance intra prediction (SDIP) partitions for a block 805 of any size. In recent years, a video encoding tool called sub-partition prediction (ISP) has been introduced into the VVC standard. Conceptually, ISP is very similar to SDIP. Specifically, depending on the block size, ISP divides the current encoding block into two or four sub-blocks either horizontally or vertically, and each sub-block contains at least 16 samples.

[0056] In summary, FIGS. 8A, 8B, and 8C show all possible partition cases considered for different encoding block sizes. Furthermore, to handle the interaction with other encoding tools in the VVC standard, the current ISP design also includes the following main aspects.

[0057] ● Interaction with wide-angle intra directions: ISP is combined with wide-angle intra directions. In the current design, the block size (i.e., width / height ratio) used to determine whether the normal intra direction or its corresponding wide-angle intra direction should be applied is one of the original encoding blocks, i.e., the block before sub-block partitioning.

[0058] ● Interaction with multiple reference lines: It is not possible to enable ISP together with multiple reference lines. Specifically, in the current VVC signaling design, the ISP enable / disable flag is signaled after the MRL index. If one intra-block has one non-zero MRL index (i.e., referring to non-nearest neighbor samples), the ISP enable / disable flag is not signaled and is assumed to be 0, that is, in this case, ISP is automatically disabled for the coded block.

[0059] ● Interaction with the most probable mode: Similar to the normal intra mode, the intra mode used for one ISP block is signaled via the most probable mode (MPM) mechanism. However, compared with the normal intra mode, the following changes have been made to the MPM method of ISP. 1) The split ISP block enables only the intra modes included in the MPM list and disables all other intra modes not in the MPM list. 2) For the split ISP block, its MPM list excludes the DC mode, prioritizes the horizontal intra mode for ISP horizontal split, and prioritizes the vertical mode for ISP vertical split.

[0060] ● At least one non-zero coefficient block flag (CBF): In the current VVC, the CBF flag is signaled for each transform unit (TU) to identify that the transform block contains one or more transform coefficient levels not equal to 0. Given a specific block using ISP, the decoder assumes that there is a non-zero CBF in at least one of the sub-divisions. Therefore, if n is the number of sub-divisions and the first n - 1 sub-divisions generate a zero CBF, the CBF of the nth sub-division is assumed to be 1. Thus, there is no need to transmit and decode it.

[0061] ● Interaction with Multiple Transform Selection: ISP is exclusively applicable with MTS. That is, when one coding block uses ISP, its MTS flag is not notified and is always presumed to be 0 (disabled). However, instead of always using DCT-II transform, a fixed set of core transforms (including DST-VII and DCT-II) is implicitly applied to the ISP coding block based on the block size. Specifically, assuming that W and H are the width and height of one ISP sub-division, its horizontal transform and vertical transform are selected according to the following rules as described in Table 3.

[0062]

Table 3

[0063] 〔Cross-Component Linear Model Prediction〕 Figure 9A is a plot of chroma values as a function of luminance values, where the plot is used to derive a set of linear model parameters. More specifically, as follows, a straight line 901 representing the relationship between chroma values and luminance values is used to derive a set of linear model parameters. To reduce cross-component redundancy, the cross-component linear model (CCLM) prediction mode is used in VVC, and for this purpose, chroma samples are predicted based on the reconstructed luminance samples of the same CU by using a linear model as follows. pred C (i,j)=α·rec L ’(i,j)+β …(8) Here, pred C (i,j) represents the predicted chroma sample within the CU, and rec L ’(i,j) represents the downsampled and reconstructed luminance sample of the same CU. The linear model parameters α and β are derived from a straight line 901 representing the relationship between luminance values and chroma values from two samples as illustrated in Figure 9A. These are the minimum luminance sample A(X A, Y A ) and the maximum luminance sample B(X B , Y B ) within the set of adjacent luminance samples. Here, X A , Y A are the values of the x coordinate (luminance value) and y coordinate (chroma value) of sample A, and X B , Y B are the values of the x coordinate and y coordinate of sample B. The linear model parameters are obtained according to the following equations. α = (y B - y A ) / (x B - x A ) β = y A - αx A Such a method is also called the min-Max method. The division in the above equations can be avoided and replaced with multiplication and shift.

[0064] Figure 9B shows the positions of the samples used for the derivation of the linear model parameters in Figure 9A. For a coded block having a square shape, the above two equations for the linear model parameters α and β are directly applicable. In the case of a non-square coded block, the adjacent samples on the longer boundary are first subsampled to have the same number of samples as the shorter boundary. Figure 9B shows the positions of the left and upper samples involved in the CCLM mode, including the N×N set of chroma samples 903 and the 2N×2N set of luminance samples 905, as well as the samples of the current block. In addition to using the top template and the left template to calculate the linear model coefficients together, these templates can also be alternatively used in two other LM modes called the LM_A and LM_L modes.

[0065] In the LM_A mode, only the pixel samples within the above template are used to calculate the linear model coefficients. To obtain more samples, the above template is extended to (W+W). In the LM_L mode, only the pixel samples within the left template are used to calculate the linear model coefficients. To obtain more samples, the left template is extended to (H+H). It should be noted that when the upper reference line is at the CTU boundary, only one luminance line (the general line buffer in intra prediction) is used to create the downsampled luminance samples.

[0066] In the case of chroma intra mode coding, a total of 8 intra modes are allowed for chroma intra mode coding. These modes include 5 conventional intra modes and 3 cross-component linear model modes (CCLM, LM_A, and LM_L). The chroma mode signaling and derivation process are shown in Table 4. Chroma mode coding directly depends on the intra prediction mode of the corresponding luminance block. In an I slice, since a separate block partitioning structure for the luminance component and the chroma component can be used, one chroma block can correspond to multiple luminance blocks. Therefore, in the case of the chroma DM mode, the intra prediction mode of the corresponding luminance block covering the center position of the current chroma block is directly inherited.

[0067] [Table 4] [Table 4: Derivation of Chroma Prediction Modes from Luminance Modes when CCLM_ is Enabled]

[0068] Although the ISP tools in VVC can improve the intra prediction efficiency, there is still room for further improving the performance of VVC. On the other hand, some parts of the existing ISP will benefit from further simplification to provide a more efficient codec hardware implementation and / or to provide improved coding efficiency. In the present disclosure, several methods are proposed to further improve the ISP coding efficiency, to simplify the existing ISP design, and / or to facilitate an improved hardware implementation.

[0069] [Independent (or parallel) sub - partition predictor generation for ISP] FIG. 10 shows the generation of reference samples for intra prediction 1007 for all sub - partitions using only the reference samples outside the current coding block 1000. The current coding block 1000 includes a first sub - partition 1, 1001, a second sub - partition 2, 1002, a third sub - partition 3, 1003, and a fourth sub - partition 4, 1004. In the present disclosure, it is proposed to generate intra predictions independently for each of the sub - partitions 1001, 1002, 1003, and 1004. In other words, all predictors for the sub - partitions 1001, 1002, 1003, and 1004 can be generated in parallel. In one embodiment, the predictors for all sub - partitions are generated using the same method as that used in the conventional non - sub - partition intra mode. Specifically, the reconstructed samples of one sub - partition are not used to generate intra prediction samples for any other sub - partition within the same coding unit, and all predictors for each of the sub - partitions 1001, 1002, 1003, 1004 are generated using the reference samples of the current coding block 1000 as shown in FIG. 10.

[0070] In the case of the ISP mode in the VVC standard, the width of each sub-division can be 2 or less. One detailed example is as follows. According to the ISP mode in the VVC standard, the dependency of 2×N (width × height) sub-block prediction on the reconstructed values of the previously decoded 2×N sub-blocks of the coding block is not allowed. As a result, the minimum width for sub-block prediction is 4 samples. For example, an 8×8 coding block encoded using ISP in vertical division is divided into 4 prediction regions each of size 2×8, and the left two 2×8 prediction regions are merged into the first 4×8 prediction region to perform intra prediction. The conversion circuit 102 (Figure 1) is applied to each 2×8 division.

[0071] According to this example, the right two 2×8 prediction regions are merged into the second 4×8 prediction region to perform intra prediction. The conversion circuit 102 is applied to each 2×8 division. It should be noted that the first 4×8 prediction region generates an intra predictor using the adjacent pixels of the current coding block, and the second 4×8 region uses the reconstructed pixels from the first 4×8 region (located to the left of the second 4×8 region), or the adjacent pixels from the current coding block (located above the second 4×8).

[0072] In another embodiment, it should be noted that only the horizontal (HOR) prediction mode (mode 18 as shown in Figure 4), and only prediction modes with mode indicators smaller than 18 (as shown in Figure 4) can be used to form the intra prediction of the horizontal sub-division, and only the vertical (VER) prediction mode (i.e., mode 50 shown in Figure 4) and prediction modes with mode indicators larger than 50 (as shown in Figure 4) can be used to form the intra prediction of the vertical sub-division. As a result, for each horizontal sub-division, the intra prediction can be performed independently and in parallel with the HOR prediction mode (shown as mode 18 in Figure 4) and all angle prediction modes with modes smaller than 18. Similarly, for each vertical sub-division, the intra prediction can be performed independently and in parallel with the VER prediction mode (mode 50 in Figure 4) and all angle prediction modes with modes greater than 50.

[0073] Intra Prediction Mode Coding for the Luminance Component for ISP In the present disclosure, it is proposed to permit only N modes out of all possible intra prediction modes for the luminance component for an ISP coding block (where N is a positive integer). In one embodiment, only one mode is permitted for the luminance component of the ISP coding block. For example, this single permitted mode may be the planar mode. In another example, this single permitted mode may be the DC mode. In a third example, this single permitted mode may be one of the HOR prediction mode, the VER prediction mode, or the diagonal (DIA) mode (mode 34 in FIG. 4) intra prediction modes.

[0074] In another embodiment, only one mode is permitted for the luminance component of the ISP coding block, and this mode may vary depending on the direction of sub-division, i.e., whether it is a horizontal sub-division or a vertical sub-division. For example, only the HOR prediction mode is permitted for horizontal sub-division, while only the VER prediction mode is permitted for vertical sub-division. In yet another example, only the VER prediction mode is permitted for horizontal sub-division, while only the HOR prediction mode is permitted for vertical sub-division.

[0075] In yet another embodiment, only two modes are permitted for the luminance component of the ISP coding block. Each of the respective modes may be selected in response to the corresponding direction of sub-division, i.e., whether the direction of sub-division is horizontal or vertical. For example, only the planar and HOR prediction modes are permitted for horizontal sub-division, while only the planar and VER prediction modes are permitted for vertical sub-division.

[0076] To signal the N modes allowed for the luminance component of the ISP symbolization block, the conventional most probable mode (MPM) mechanism is not utilized. Instead, the inventors propose to signal the intra mode for the ISP symbolization block using the binary codeword determined for each intra mode. The codeword can be generated using any of a variety of different processes, including a truncated binary (TB) binarization process, a fixed-length binarization process, a truncated Rice (TR) binarization process, a k-th Exp-Golomb binarization process, a limited EGk binarization process, etc. These binary codeword generation processes are clearly defined in the HEVC specification. Truncated Rice with the Rice parameter equal to zero is also known as truncated unary binarization. An example of a set of codewords using different binarization methods is shown in Table 5.

[0077]

Table 5

[0078] [Independent (or parallel) sub-partition predictor generation for ISP] In yet another embodiment, the MPM derivation process for the normal intra mode is directly reused for the ISP mode, and the signaling method for the MPM flag and MPM index is kept the same as the existing ISP design.

[0079] [Intra prediction mode coding for the chrominance component for ISP] In the present disclosure, for the chrominance component for an ISP symbolization block (N c is a positive integer), N from all possible chrominance intra prediction modes cIt is proposed to permit only one mode. In one embodiment, only one mode is permitted for the chroma component of the ISP encoding block. For example, this single permitted mode may be the direct mode (DM). In another example, this single permitted mode may be the LM mode. In a third example, this single permitted mode can be one of the HOR prediction mode or the VER prediction mode. The direct mode (DM) is configured to apply the same intra prediction mode used by the corresponding luminance block to the chroma block.

[0080] In another embodiment, only two modes are allowed for the chroma component of the ISP encoding block. In one example, only DM and LM are permitted for the chroma component of the ISP encoding block.

[0081] In yet another embodiment, only four modes are permitted for the chroma component of the ISP encoding block. In one example, only DM, LM, LM_L, and LM_A are permitted for the chroma component of the ISP encoding block.

[0082] For the chroma component of the ISP encoding block, N c To signal the modes for the chroma component of the ISP encoding block, the conventional MPM mechanism is not used. Instead, a fixed binary codeword is used to indicate the chroma mode selected in the bitstream. For example, the chroma intra prediction mode of the ISP encoding block can be signaled using the determined binary codeword. The codeword can be generated using different processes including the truncated binary (TB) binarization process, the fixed-length binarization process, the truncated Rice (TR) binarization process, the k-th Exp-Golomb binarization process, the limited EGk binarization process, etc.

[0083] 〔Encoding Block Size of Chroma Component of ISP Encoding Block〕 In the present disclosure, it is proposed not to permit sub - division coding for the chrominance component for the ISP coding block. Instead, for the chrominance component of the ISP coding block, normal full - block - based intra - prediction is used. In other words, in the case of the ISP coding block, sub - division is performed only on its luminance component which does not have a sub - division for its chrominance component.

[0084] 〔Combination of ISP and Inter - prediction〕 FIG. 11 shows the combination of an inter - predictor sample and an intra - predictor sample for the first sub - division 1001 of FIG. 10. Similarly, FIG. 12 shows the combination of an inter - predictor sample and an intra - prediction sample for the second sub - division 1002 of FIG. 10. More specifically, in order to further improve the coding efficiency, a new prediction mode is provided in which the prediction is generated as a weighted combination (e.g., weighted averaging) of the ISP mode and the inter - prediction mode. The generation of the intra - predictor for the ISP mode is the same as that described above in relation to FIG. 10. The inter - predictor may be generated by a process of the merge mode or the inter - mode.

[0085] In the exemplary example of FIG. 11, the inter - predictor sample 1101 of the current block (including all sub - divisions) is generated by performing motion compensation using a merge candidate indicated by a merge index. The intra - predictor sample 1103 is generated by performing intra - prediction using the signalized intra - mode. It should be noted that this process can generate the intra - predictor sample of a non - first sub - division (i.e., the second sub - division 1002) as shown in FIG. 12 using the reconstructed samples of the previous sub - division. After the inter - and intra - predictor samples are generated, they are weighted - averaged to generate the final predictor sample for the sub - division. The combined mode can be treated as an intra - mode. Alternatively, the combined mode may be treated as an inter - mode or a merge mode instead of the intra - mode.

[0086] [CBF Signaling for ISP Symbolization Block] To simplify the design of the ISP, in this disclosure, it is proposed to always signal the CBF for the last sub - division.

[0087] In another embodiment of this disclosure, it is proposed to infer the value of the CBF of the last sub - division at the decoder side without signaling it. For example, the value of the CBF of the last sub - division is always inferred as 1. In another example, the value of the CBF for the last sub - division is always inferred as zero.

[0088] According to one embodiment of this disclosure, a video encoding method includes independently generating respective intra - predictions for each of a plurality of corresponding sub - divisions, and each intra - prediction is generated using a plurality of reference samples from the current encoding block.

[0089] In some examples, the reconstructed samples from the first sub - division of the plurality of corresponding sub - divisions are not used to generate respective intra - predictions for any other sub - division of the plurality of corresponding sub - divisions.

[0090] In some examples, each of the plurality of corresponding sub - divisions has a width of 2 or less.

[0091] In some examples, the plurality of corresponding sub - divisions include a plurality of vertical sub - divisions and a plurality of horizontal sub - divisions, and the method further includes generating a first set of intra - predictions for the plurality of horizontal sub - divisions using only the horizontal prediction mode, and generating a second set of intra - predictions for the plurality of vertical sub - divisions using only the vertical prediction mode.

[0092] In some examples, the horizontal prediction mode is executed using a mode index less than 18.

[0093] In some examples, the vertical prediction mode is executed using a mode index greater than 50.

[0094] In some examples, the horizontal prediction mode is executed independently and in parallel for each of a plurality of horizontal sub - partitions.

[0095] In some examples, the vertical prediction mode is executed independently and in parallel for each of a plurality of vertical sub - partitions.

[0096] In some examples, the plurality of corresponding sub - partitions includes the last sub - partition, and the method further comprises signaling a coefficient block flag (CBF) value of the last sub - partition.

[0097] In some examples, the plurality of corresponding sub - partitions includes the last sub - partition, and the method further comprises inferring, at a decoder, a coefficient block flag (CBF) value of the last sub - partition.

[0098] In some examples, the coefficient block flag (CBF) value is inferred as always 1.

[0099] In some examples, the coefficient block flag (CBF) value is inferred as always zero.

[0100] According to another embodiment of the present disclosure, a video encoding method includes generating a respective intra - prediction for each of a plurality of corresponding sub - partitions for a luminance component of an intra - sub - partition (ISP) encoding block using only N out of M possible intra - prediction modes, where M and N are positive integers and N is less than M.

[0101] In some examples, each of the plurality of corresponding sub - partitions has a width of 2 or less.

[0102] In some examples, N is equal to 1, such that only a single mode is allowed for the luminance component.

[0103] In some examples, the single mode is the planar mode.

[0104] In some examples, the single mode is the DC mode.

[0105] In some examples, the single mode is any one of a horizontal (HOR) prediction mode, a vertical (VER) prediction mode, or a diagonal (DIA) prediction mode.

[0106] In some examples, the single mode is selected in response to the direction of sub-division, the horizontal (HOR) prediction mode is selected in response to the direction of sub-division being horizontal, and the vertical (VER) prediction mode is selected in response to the direction of sub-division being vertical.

[0107] In some examples, the single mode is selected in response to the direction of sub-division, the horizontal (HOR) prediction mode is selected in response to the direction of sub-division being vertical, and the vertical (VER) prediction mode is selected in response to the direction of sub-division being horizontal.

[0108] In some examples, N is equal to 2 and two modes are allowed for the luminance component.

[0109] In some examples, the first set of two modes is selected in response to the direction of the first sub-division, and the second set of two modes is selected in response to the direction of the second sub-division.

[0110] In some examples, the first set of two modes includes a planar mode and a horizontal (HOR) prediction mode, the first sub-division direction includes a horizontal sub-division, the second set of modes includes a planar mode and a vertical (VER) prediction mode, and the second sub-division direction includes a vertical sub-division.

[0111] In some examples, each mode of the N modes is signaled using a corresponding binary codeword from a set of predetermined binary codewords.

[0112] In some examples, a set of predetermined binary codewords is generated using at least one of a truncating binary (TB) binarization process, a fixed-length binarization process, a truncating Rice (TR) binarization process, a truncating unary binarization process, a k-th Exp-Golumb binarization process, or a limited EGk binarization process.

[0113] According to another embodiment of the present disclosure, a method of video encoding includes generating an intra prediction for a chrominance component of an intra-subdivision (ISP) encoding block using only N of M possible intra prediction modes, where M and N are positive integers and N is less than M.

[0114] In some examples, N is equal to 1, and as a result, only a single mode is allowed for the chrominance component.

[0115] In some examples, the single mode is a direct mode (DM), a linear model (LM) mode, a horizontal (HOR) prediction mode, or a vertical (VER) prediction mode.

[0116] In some examples, N is equal to 2, and as a result, two modes are allowed for the chrominance component.

[0117] In some examples, N is equal to 4, and as a result, four modes are allowed for the chrominance component.

[0118] In some examples, each of the N modes is signaled using a corresponding binary codeword from a set of predetermined binary codewords.

[0119] In some examples, a set of predetermined binary codewords is generated using at least one of a truncating binary (TB) binarization process, a fixed-length binarization process, a truncating Rice (TR) binarization process, a truncating unary binarization process, a k-th Exp-Golumb binarization process, or a limited EGk binarization process.

[0120] According to another embodiment of the present disclosure, a video encoding method includes generating a respective intra prediction for each of a plurality of corresponding sub - partitions of an entire intra - sub - partition (ISP) encoding block for a luminance component, and for a chrominance component, and generating a chrominance intra prediction for the entire intra - sub - partition (ISP) encoding block.

[0121] In some examples, each of the plurality of corresponding sub - partitions has a width of 2 or less.

[0122] According to another embodiment of the present disclosure, a video encoding method generates a first prediction using an intra - sub - partition mode, generates a second prediction using an inter - prediction mode, and combines the first and second predictions to generate a final prediction by applying a weighted average to the first and second predictions.

[0123] In some examples, the second prediction is generated using at least one of a merge mode or an inter mode.

[0124] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. In this way, the computer-readable medium generally can correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to search for instructions, code, and / or data structures for the implementations described in this application. A computer program product may include a computer-readable medium.

[0125] Furthermore, the above method may be implemented using one or more circuits including devices such as application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. The device may use the circuits in combination with other hardware or software components to perform the methods described above. Each module, sub-module, unit, or sub-unit disclosed above may be at least partially implemented using one or more circuits.

[0126] Other embodiments of the present invention will be apparent to those skilled in the art in view of the present specification and by practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that fall within the scope of the known or customary practice in the art and that depart from the present disclosure in accordance with the general principles of the present invention. The specification and each embodiment are intended to be illustrative only, and the true scope and spirit of the present invention are set forth in the claims.

[0127] It will be understood that the present invention is not limited precisely to the examples described above and shown in the accompanying drawings, and that various modifications and changes can be made without departing from the scope of the present invention. The scope of the present invention is intended to be limited only by the appended claims.

[0128] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 801,214, filed Feb. 5, 2019. The entire disclosure of the foregoing application is hereby incorporated by reference in its entirety.

Claims

1. When it is determined that a plurality of sub - partitions satisfy the first condition of intra - prediction, perform the intra - prediction for the plurality of sub - partitions of an encoded block in an intra - sub - partition (ISP) mode, Performing the intra - prediction for the plurality of sub - partitions of the encoded block, merge at least two of the plurality of sub - partitions into one prediction region of the intra - prediction, generate prediction samples of the one prediction region of the intra - prediction based on a plurality of reference samples of the encoded block adjacent to the one prediction region, including, The first condition includes that the plurality of sub - partitions are a plurality of vertical sub - partitions, and each of the sub - partitions has a width of 2 or less. Video encoding method.

2. Samples reconstructed from one of the at least two sub - partitions are not used to perform the intra - prediction of any other of the at least two sub - partitions. The video encoding method according to claim 1.

3. Further including generating a set of prediction samples for the plurality of vertical sub - partitions using a vertical prediction mode with a mode index greater than 50. The video encoding method according to claim 1.

4. Further including independently and in parallel performing the intra - prediction for the samples of the one prediction region. The video encoding method according to claim 1.

5. The plurality of sub - partitions includes a last sub - partition, further including signaling a coefficient block flag (CBF) value for the last sub - partition. The video encoding method according to claim 1.

6. The plurality of sub - partitions includes a last sub - partition, further including, in an encoder, inferring a coefficient block flag (CBF) value of the last sub - partition. The video encoding method according to claim 1.

7. The coefficient block flag (CBF) value is inferred as 1 or 0. The video encoding method according to claim 6.

8. A plurality of prediction samples of the intra - prediction for the plurality of sub - partitions are generated in parallel. The video encoding method according to claim 1.

9. One or more processors, a non - transitory computer - readable memory storing instructions executable by the one or more processors, including, The one or more processors are configured to generate a bitstream and store the bitstream, and when it is determined that a plurality of sub - partitions satisfy a first condition of intra - prediction, perform intra - prediction on the plurality of sub - partitions of an encoding block in an intra - sub - partition (ISP) mode, and are configured as such, performing the intra - prediction on the plurality of sub - partitions of the encoding block includes merging at least two of the plurality of sub - partitions into one prediction region of the intra - prediction, and generating prediction samples of the one prediction region of the intra - prediction based on a plurality of reference samples of the encoding block adjacent to the one prediction region, and includes the first condition includes that the plurality of sub - partitions are a plurality of vertical sub - partitions and each of the sub - partitions has a width of 2 or less, a video encoding device.

10. Samples reconstructed from one of the at least two sub - partitions are not used to perform the intra - prediction of any other of the at least two sub - partitions, The video encoding device according to claim 9.

11. The one or more processors are further configured to generate a set of prediction samples for the plurality of vertical sub - partitions using a vertical prediction mode with a mode index greater than 50, and are configured as such, The video encoding device according to claim 9.

12. A non - transitory computer - readable storage medium storing a plurality of programs executed by a video encoding device having one or more processors, wherein when executed by the one or more processors, the plurality of programs cause the video encoding device to execute the video encoding method according to any one of claims 1 to 8, a non - transitory computer - readable storage medium.

13. A method for storing a bitstream, comprising executing the video encoding method according to any one of claims 1 to 8 to generate a bitstream, and storing the bitstream. A method.

14. when it is determined that a plurality of sub - partitions satisfy a first condition of intra - prediction, perform the intra - prediction on the plurality of sub - partitions of an encoding block in an intra - sub - partition (ISP) mode, wherein the encoding block is associated with a predetermined partitioning method, and the partitioning method includes 4 - way partitioning, horizontal 3 - way partitioning, vertical 3 - way partitioning, horizontal 2 - way partitioning, or vertical 2 - way partitioning, Performing the intra prediction for the plurality of sub - divisions of the encoding block includes merging at least two sub - divisions of the plurality of sub - divisions into one prediction region of the intra prediction, generating prediction samples for the one prediction region of the intra prediction based on a plurality of reference samples of the encoding block adjacent to the one prediction region, including The first condition includes that the plurality of sub - divisions are a plurality of vertical sub - divisions, and each of the sub - divisions has a width of 2 or less. Video decoding method.

15. Samples reconstructed from one of the at least two sub - divisions are not used to perform the intra prediction for any other of the at least two sub - divisions. The video decoding method according to claim 14.

16. Further including generating a set of prediction samples for the plurality of vertical sub - divisions using a vertical prediction mode having a mode index greater than 50. The video decoding method according to claim 14.

17. Further including independently and in parallel performing the intra prediction for samples of the one prediction region. The video decoding method according to claim 14.

18. The plurality of sub - divisions includes a last sub - division, further including signaling a coefficient block flag (CBF) value for the last sub - division. The video decoding method according to claim 14.

19. The plurality of sub - divisions includes a last sub - division, further including inferring a coefficient block flag (CBF) value for the last sub - division at a decoder. The video decoding method according to claim 14.

20. The coefficient block flag (CBF) value is inferred as 1 or 0. The video decoding method according to claim 19.

21. A plurality of prediction samples for the intra prediction for the plurality of sub - divisions are generated in parallel. The video decoding method according to claim 14.

22. One or more processors, a non - transitory computer - readable memory storing instructions executable by the one or more processors, including The one or more processors are configured to perform intra prediction for the plurality of sub - divisions of an encoding block in an intra - sub - division (ISP) mode when it is determined that the plurality of sub - divisions satisfy a first condition of the intra prediction. configured as such. The encoding block is associated with a predetermined splitting method, wherein the predetermined splitting method includes four-way splitting, three-way horizontal splitting, three-way vertical splitting, two-way horizontal splitting, or two-way vertical splitting, performing the intra prediction for the plurality of sub-splits of the encoding block merges at least two of the plurality of sub-splits into one prediction region of the intra prediction, generates prediction samples for the one prediction region of the intra prediction based on a plurality of reference samples of the encoding block adjacent to the prediction region, including wherein the first condition includes that the plurality of sub-splits are a plurality of vertical sub-splits, and each of the sub-splits has a width of 2 or less, Video decoding apparatus.

23. Samples reconstructed from one of the at least two sub-splits are not used to perform the intra prediction for any of the other sub-splits of the at least two sub-splits, The video decoding apparatus according to claim 22.

24. The one or more processors further generate a set of prediction samples for the plurality of vertical sub-splits using a vertical prediction mode having a mode index greater than 50, configured as The video decoding apparatus according to claim 22.

25. A non-transitory computer-readable storage medium storing a plurality of programs executed by a video decoding apparatus having one or more processors, wherein when executed by the one or more processors, the plurality of programs cause the video decoding apparatus to execute the video decoding method according to any one of claims 14 to 21, Non-transitory computer-readable storage medium.