Wide angle intra prediction with sub-partitions

By employing wide-angle intra-frame prediction and reference sample replacement techniques for sub-partitioning, the inefficiency of block shape adaptive prediction direction in video coding is addressed, thereby improving video compression efficiency and coding performance.

CN121967679APending Publication Date: 2026-05-01INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTERDIGITAL VC HOLDINGS INC
Filing Date
2020-04-09
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing video coding techniques suffer from inefficiencies in intra-frame prediction direction adaptation based on block shape, particularly in the availability of reference samples and prediction direction matching in sub-partition prediction.

Method used

Wide-angle intra-frame prediction of sub-partitions is adopted. The prediction direction of the sub-partitions is adjusted to adapt to the aspect ratio of the parent decoding unit. When the reference sample is unavailable, the missing sample is replaced from the parent CU. Improved reference sample mapping and filtering techniques are used.

Benefits of technology

It improves the compression efficiency of video encoding, especially in the luminance and chrominance components, achieving higher BD rate performance gains while maintaining stable encoding and decoding complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967679A_ABST
    Figure CN121967679A_ABST
Patent Text Reader

Abstract

A method and apparatus for performing prediction for encoding or decoding uses intra prediction employing sub-partitions. The sub-partitions are oriented horizontally or vertically and may use a wide angle mode that is different from the mode of the video block from which they originate. When a reference sample for a sub-partition is unavailable, for example due to direction, the reference sample for the sub-partition is a reference sample for the video block. In an embodiment, when a sub-partition is square, a conventional intra prediction direction is used. By using the mapping in the prediction direction vertically or horizontally, the reference samples are used from blocks above and right of the video block or blocks left and below the video block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of patent application No. 202080028075.4, filed on April 9, 2020, entitled "Wide-angle Intra-frame Prediction Using Sub-partitions". Technical Field

[0002] At least one embodiment of the present invention generally relates to a method or apparatus for video encoding or decoding. Background Technology

[0003] To achieve high compression efficiency, image and video decoding schemes typically employ prediction (including spatial and / or motion vector prediction) and transform to utilize spatial and temporal redundancy in the video content. Intra-frame or inter-frame prediction is usually used to leverage intra-frame or inter-frame correlations, followed by transform, quantization, and entropy decoding of the differences between the original and predicted images (typically represented as prediction error or prediction residual). To reconstruct the video, the compressed data is decoded through inverse processing corresponding to entropy decoding, quantization, transform, and prediction. Summary of the Invention

[0004] The shortcomings and disadvantages of the prior art can be addressed by the general aspects described in this paper, which involve intra-frame prediction directions that adapt block shape in encoding and decoding.

[0005] According to a first aspect, a method is provided. The method includes the steps of: determining an intra-prediction direction for a rectangular sub-partition of a video block as follows: determining the direction based on the size of the sub-partition when the ratio of the sub-partition size is within a specific range; and determining the direction based on the size of the video block when the ratio of the sub-partition size is outside the specific range; mapping reference samples from the upper right or lower left of the video block to upper right or lower left reference samples for prediction of the sub-partition of the video block; predicting samples of the rectangular sub-partition using reference samples from a row above the video block or from a column to the left of the video block, wherein the number of reference samples in the row above the video block or the column to the left of the video block is determined based on the size of the rectangular sub-partition; and using the prediction to encode the rectangular sub-partition of the video block in an intra-decoding mode.

[0006] According to a second aspect, a method is provided. The method includes the following steps: determining an intra-prediction direction for a rectangular sub-partition of a video block as follows: determining the direction based on the size of the sub-partition when the ratio of the sub-partition size is within a specific range; and determining the direction based on the size of the video block when the ratio of the sub-partition size is outside the specific range; mapping reference samples from the upper right or lower left of the video block to upper right or lower left reference samples for prediction of the sub-partition of the video block; predicting samples of the rectangular sub-partition using reference samples from a row above the video block or from a column to the left of the video block, wherein the number of reference samples in the row above the video block or the column to the left of the video block is determined based on the size of the rectangular sub-partition; and using the prediction in an intra-decoding mode to decode the rectangular sub-partition of the video block.

[0007] According to another aspect, an apparatus is provided. The apparatus includes a processor. The processor is configured to encode blocks of video or decode bitstreams by performing any of the methods described above.

[0008] According to another general aspect of at least one embodiment, an apparatus is provided, including an apparatus according to any of the decoding embodiments; and at least one of: (i) an antenna configured to receive a signal including a video block, (ii) a band limiter configured to limit the received signal to a band including the video block, or (iii) a display configured to display an output representing the video block.

[0009] According to another general aspect of at least one embodiment, a non-transitory computer-readable medium is provided, which contains data content generated according to any of the described encoding embodiments or variations.

[0010] According to another general aspect of at least one embodiment, a signal comprising video data generated according to any of the described encoding embodiments or variations is provided.

[0011] According to another general aspect of at least one embodiment, the bitstream is formatted to include data content generated according to any of the described encoding embodiments or variations.

[0012] According to another general aspect of at least one embodiment, a computer program product including instructions is provided that, when a computer executes the program, cause the computer to perform any of the described decoding embodiments or variations.

[0013] These and other aspects, features, and advantages of the general aspects will become apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Attached Figure Description

[0014] Figure 1 A reference sample for intra-frame prediction in VTM is shown.

[0015] Figure 2 Several reference lines are shown for intra-frame prediction in VTM.

[0016] Figure 3 The intra-prediction direction in the VTM is shown for a square target block.

[0017] Figure 4 Examples of 4×8 and 8×4 block partitioning are shown.

[0018] Figure 5 Examples of block partitioning are shown, except for 4×8, 8×4, and 4×4.

[0019] Figure 6 The wide-angle intra-frame prediction for non-square blocks is shown.

[0020] Figure 7 The ISP width angle is shown as calculated from the CU height and width.

[0021] Figure 8 A reference sample is shown in the case of wide-angle and horizontally split ISP.

[0022] Figure 9 An example of the proposed reference sample array used to predict sub-partitions is shown.

[0023] Figure 10 An example of reference array filling in a VTM for an exemplary case of two sub-partitions is shown.

[0024] Figure 11 The reference sample mapping from the top reference of the decoding unit is shown in the case of horizontal splitting.

[0025] Figure 12 The reference sample mapping from the top reference of the decoding unit is shown in the case of horizontal splitting.

[0026] Figure 13 This illustrates a general standard encoding scheme.

[0027] Figure 14 This illustrates a typical standard decoding scheme.

[0028] Figure 15 A typical processor arrangement in which the described embodiments can be implemented is shown.

[0029] Figure 16 This illustrates a general standard encoding scheme.

[0030] Figure 17 This illustrates a typical standard decoding scheme.

[0031] Figure 18 A typical processor arrangement in which the described embodiments can be implemented is shown. Detailed Implementation

[0032] The embodiments described herein belong to the field of video compression, and relate to video compression as well as video encoding and decoding.

[0033] The Universal Video Decoding (VVC) Test Model 4.0 (VTM) supports intra-frame prediction (ISP) with sub-partitions, where the target block is encoded using the option of horizontal or vertical sub-partitions. Prediction for each sub-partition uses the same prediction direction as the target decoding unit. Sub-partition prediction can be improved by employing wide-angle prediction suitable for the aspect ratio of each sub-partition, independent of the parent decoding unit. Secondly, sub-partition prediction uses decoded pixels from previous sub-partitions as reference samples. Therefore, reference samples in the upper right or lower left of a sub-partition may be unavailable. Reference samples used for the parent CU can be used to improve prediction along directions where those samples are needed.

[0034] In the Universal Video Decoding (VVC) Test Model (VTM), any target block in intra-frame prediction can have one of 67 prediction modes. Similar to HEVC, there are two modes: a planar mode, a DC mode, and the remaining 65 are directional modes. These 65 directional modes are selected from 95 directions, including 65 regular angles spanning from 45 degrees to -135 degrees if the target block is square, and possibly 28 wide-angle directions if the target block is rectangular. VTM encodes the prediction modes of the block using a set of most probable modes (MPMs), which consists of 6 prediction modes. If a prediction mode does not belong to the MPM set, it is truncated into binary encoding with 5 or 6 bits.

[0035] Intra-frame prediction in video compression refers to spatial prediction of pixel blocks using information from causally adjacent blocks (i.e., neighboring blocks that have already been decoded within the same frame). This is a powerful decoding tool because it allows for high compression efficiency both intra-frame and inter-frame, provided there is no better temporal prediction. Therefore, intra-frame prediction has been included as a core decoding tool in all video compression standards, including H.264 / AVC, HEVC, and others. In the following text, for illustrative purposes, we will refer to intra-frame prediction in the Universal Decoder-Codec (VVC) Software Test Model (VTM).

[0036] In VTM, the encoding of frames in a video sequence is based on a quadtree (QT) / polytree (MTT) block structure. Frames are divided into non-overlapping square decode tree units (CTUs), and each CTU undergoes a QT / MTT-based split into multiple decode units (CUs) based on a rate distortion standard. For ease of reference, we will use the terms "CU" and "block" interchangeably throughout this document.

[0037] In intra-frame prediction, CUs are spatially predicted from causally adjacent CUs (i.e., the top and left CUs). For this purpose, VTM uses a simple spatial model called prediction modes. Based on the decoded pixel values ​​in the top and left CUs, called reference pixels, the encoder constructs different predictions for the target block and selects the one that results in the best RD performance. Of the 95 defined modes, one is a planar mode (indexed as mode 0), one is a DC mode (indexed as mode 1), and the remaining 93 (indexed as mode 14…-1, 2…80) are corner modes. Of the 93 corner modes, only 65 neighboring modes are selected for any target CU based on their shape. Corner modes are designed to model the directional structure of objects within the frame. Therefore, the decoded pixel values ​​in the top and left CUs are repeated only along the defined directions to fill the target CU. Some prediction modes can lead to discontinuities along the top and left reference boundaries, so those prediction modes may include subsequent post-processing (called position-dependent intra-frame prediction combination (PDPC)) designed to smooth the pixel values ​​along those boundaries.

[0038] The intra-frame prediction process in VTM consists of three steps: (1) reference sample generation, (2) intra-sample prediction, and (3) post-processing of predicted samples. Figure 1 The diagram illustrates the reference sample generation process, showing reference samples used for intra-frame prediction in VTM. H and W represent the height and width of the current block, respectively. For a CU of size HxW, a top row of 2W decoded samples is formed from the top and upper right pixels of the previously reconstructed current CU. Similarly, a left column of 2H samples is formed from the reconstructed left and lower left pixels. Corner pixels at the upper left position are also used to fill the gap between the top row and the left column references. If some of the top or left samples are unavailable because the corresponding CU is not in the same slice, or the current CU is at a frame boundary, or for some other reason, then a method called reference sample substitution can be performed.

[0039] VTM also supports intra-frame prediction using multiple reference lines (MRLs). Its idea is based on... Figure 2 Several sets of reference lines are shown for prediction, and then the reference line that gives the best rate distortion performance is selected. The reference line used is signaled to the decoder in a variable-length code.

[0040] The next step (i.e., in-sample prediction) involves predicting the pixels of the target CU based on a reference sample. As mentioned earlier, to effectively predict different types of content, VTM supports a range of prediction models. Planar and DC prediction modes are used to predict smooth and gradually changing regions, while angular prediction modes are used to capture different directional structures. VTM supports 95 directional prediction modes, indexed from -14 to -1 and from 2 to 80, with only prediction modes 2-66 used for square CUs. These prediction modes correspond to different prediction directions from 45 degrees to -135 degrees in a clockwise direction, such as... Figure 3 The diagram illustrates the intra-prediction direction in a VTM used for square target blocks. Typically, non-square blocks can also be used with extended prediction directions, but this diagram shows square blocks. The numbers represent the prediction mode indices associated with the corresponding directions. Modes 2 through 33 indicate horizontal prediction, while modes 34 through 66 indicate vertical prediction.

[0041] Patterns with indices from -14 to -1 and from 67 to 80 are wide-angle patterns, used for rectangular blocks of different shapes. Patterns 14 to 1 are defined beyond pattern 2 (45 degrees) and are used for tall rectangular blocks (blocks whose height is greater than their width). Similarly, patterns 67 to 80 are defined beyond pattern 66 (-135 degrees) and are used for flat rectangular blocks (blocks whose width is greater than their height). The number of wide-angle patterns used for rectangular blocks depends on the block's aspect ratio. In any case, the total number of wide-angle patterns used for any block is 65, and the patterns are always continuous in direction.

[0042] The general aspects described in this paper aim to improve the prediction efficiency of intra-frame prediction using sub-partitions. First, it proposes using the wide-angle pattern of each sub-partition instead of the wide-angle pattern of the parent decoding unit to predict it. Second, it also proposes using the parent CU's reference samples to predict sub-partitions when those samples are unavailable. These two suggestions can be combined into current VTM codes to improve decoding performance. The increased complexity and memory requirements for incorporating these changes are minimal, with the potential for significant decoding gain.

[0043] In the Universal Video Decoding (VVC) Test Model (VTM), the encoding of frames in a video sequence is based on a Quadtree (QT) / Multi-Type Tree (MTT) block structure. Frames are divided into non-overlapping square decoder tree units (CTUs), each of which undergoes a QT / MTT-based split into multiple decoder units (CUs) based on a rate distortion criterion. In intra-frame prediction, CUs are spatially predicted from causally adjacent CUs (i.e., the top and left CUs). For this purpose, VTM uses a simple spatial model called the prediction mode. Based on the decoded pixel values ​​in the top and left CUs (called reference pixels), the encoder constructs different predictions for the target block and selects the one that results in the best RD performance.

[0044] In VTM 4.0, the target block has the option to choose between intra-prediction for the entire CU and intra-prediction using sub-partitions (ISP) for the CU. In the latter prediction, the target CU is divided into two or four equal-sized sub-partitions, which are then sequentially decoded using a prediction mode. That is, each sub-partition is decoded independently. Therefore, sub-partitions benefit from the availability of decoded samples from neighboring sub-partitions, which are the direct neighbors of the current sub-partition. This, in turn, leads to better prediction and thus higher compression efficiency.

[0045] VTM 4.0 also supports Wide Angle Intra Prediction (WAIP), which allows the use of wide-angle prediction directions for rectangular blocks. Depending on the aspect ratio of the target block, some regular prediction directions are replaced with corresponding wide-angle directions. The prediction directions have also been appropriately modified to accommodate different rectangular block shapes, ensuring that the defined prediction directions are aligned along the second diagonal of the block. Effectively, by using wide angles instead of regular angles, this method aims to provide better predictions that result in higher compression efficiency.

[0046] In VTM 4.0, the combination of ISP and WAIP is straightforward. For each sub-partition in the ISP, the prediction direction is the same as that of the parent CU. However, this is not mandatory. Since sub-partitions are decoded independently, the prediction direction of each sub-partition can be determined independently using WAIP. Secondly, except for the first sub-partition, the remaining sub-partitions use pixels decoded from previous sub-partitions as reference samples for prediction. This leads to an undesirable effect (i.e., depending on the prediction direction), some reference samples may be unavailable. This situation can be improved by using reference samples from the parent CU. Before describing the proposed pattern, ISP and WAIP are briefly presented in VTM 4.0 below. For easier reference, we will use the terms "CU" and "block" interchangeably in this paper.

[0047] The ISP tool in VTM 4.0 divides the luma intra-prediction block vertically or horizontally into 2 or 4 sub-partitions based on the block size. Each sub-partition must have at least 16 samples. Therefore, a 4×4 block is not divided into sub-partitions, while 4×8 and 8×4 blocks have only two sub-partitions. All other block sizes have only four sub-partitions. Sub-partitions can be horizontal or vertical. A 4x8 block can have only two 4x4 vertical sub-partitions, while an 8x4 block can have only two 4x4 horizontal sub-partitions. Similarly, as another example, a 4×16 block can each have four 4×4 vertical sub-partitions or four 1×16 horizontal sub-partitions. Figure 4 and Figure 5Examples of the two possibilities are shown.

[0048] Table 1: Number of sub-partitions depending on block size

[0049] For each of these sub-partitions, a prediction is constructed using the decoding prediction mode of the parent CU. This prediction signal is added to the decoded residual signal (generated by entropy decoding of the coefficients sent by the encoder, followed by inverse quantization and inverse transform) to reconstruct the pixels in the sub-partition. The reconstructed values ​​of each sub-partition, except the first, can then be used to generate the prediction for the next sub-partition.

[0050] Sub-partitions are processed in normal order, regardless of the intra-frame mode and splitting used. That is, the first sub-partition to be processed is the upper left sample containing the CU, and then the sub-partitions continue sequentially downward (horizontal splitting) or to the right (vertical splitting).

[0051] VTM 4.0 also supports intra-frame prediction using multiple reference lines (MRLs). The target block can choose to use the first, second, or fourth reference line that provides the best rate distortion performance. A one-bit (0) or two-bit (10 or 11) flag is used to signal the selected reference line to indicate either the first reference line or the second and fourth reference lines, respectively. In VTM 4.0, ISP is only applied to blocks with the first reference line. Therefore, if a block has an MRL index other than 0, the ISP decoding mode will be inferred as 0, and thus it will not be sent to the decoder.

[0052] The ISP method has been tested with intra-frame modes as part of the MPM list, which consists of six different modes out of 67 prediction modes. For any block tested with ISP, the MPM list was modified to exclude DC modes and prioritize horizontal intra-frame modes used for horizontal splits and vertical intra-frame modes used for vertical splits.

[0053] The basic idea behind wide-angle prediction is to adapt the prediction direction to the block shape while keeping the total number of prediction patterns the same. This is done by adding some prediction directions on the larger sides of the block and reducing the number of prediction directions on the shorter sides. The overall goal is to improve prediction accuracy, resulting in higher compression efficiency. Because the newly introduced directions exceed the usual 180-degree range of angles from 45 degrees to -135 degrees, they are called wide-angle directions.

[0054] When the target block is square, wide angles have no effect because the definition pattern used for the block remains unchanged. When the target block is flat (i.e., its width W is greater than its height H), some patterns close to 45 degrees are removed, and an equal number of wide angle patterns exceeding -135 degrees are added. The added directions are indexed as prediction patterns 67, 68, ..., etc. Similarly, when the target block is tall, some patterns close to -135 degrees are removed, and an equal number of wide angle patterns exceeding 45 degrees are added. The added directions are indexed as prediction patterns -1, -2, ..., etc., because prediction patterns 0 and 1 are retained for planar and DC predictions.

[0055] Since the direction of the secondary diagonal depends on the shape of the block, the total number of wide angles used depends on the shape of the block. Table 2 shows the number of regular patterns replaced by wide angle patterns for different block shapes.

[0056] Table 2: Number of wide-angle patterns depending on block shape

[0057] For any target block, the mapping from the regular replacement pattern to the wide-angle pattern is performed as follows: modeShift[5] = {0, 6, 10, 12, 14}; ratio = Abs(Log2(W / H)) if W>H and 1 <predMode<2 + modeShift[ratio] predMode = predMode + 65; else if H> W and (66 – modeShift[ratio]) <predMode<= 66 predMode = predMode - 67; In the case of ISP, the mapping from regular mode to wide-angle mode is done using the height and width of the parent CU, not the height and width of the sub-regions. Regardless of whether the new mode is wide-angle or regular, the same pattern is used for prediction in each sub-region. This is... Figure 7 The example shown is as follows.

[0058] As shown above, the current approach to ISPs with WAIPs is to determine the prediction angle from the aspect ratio of the parent CU and use it for predictions in each sub-partition. Since the sub-partitions are processed sequentially but independently, this constraint does not need to be bound. A second problem arising from processing individual sub-partitions is the unavailability of some reference samples, which can affect predictions in certain directions. Below, we provide two solutions to these problems.

[0059] For the sake of generality, we will assume in the following text that the rectangular block has a width W and a height H. The square target block is a special case where W = H.

[0060] WAIP for intra-frame sub-partitions For blocks using ISP, the current VTM code (VTM 4.0) processes each sub-partition individually in normal order. During prediction, the prediction tool first generates a reference array for the samples at the top of the partition and a reference array for the samples on the left. If the ISP split is horizontal, the top reference array is constructed using the decoded samples from the last row of the previous sub-partition. Similarly, if the ISP split is vertical, the left reference array is constructed using the decoded samples from the last column of the previous sub-partition. This is in Figure 8 As shown in the image.

[0061] The current VTM code determines the size of the reference array for each partition based on the size of the reference array required by the parent CU. We propose determining the prediction direction based on the size of the sub-partitions. As an advantage of the implementation, we propose also determining the length of the partition's reference array based on the partition size.

[0062] Assumption and This represents the width and height of each partition. Then, the lengths of the top and top-left reference arrays are determined as follows: and And exclude the top left pixel.

[0063] For any subpartition, such as Figure 9 As shown, the mapping from the normal mode to the wide-angle mode is performed as follows: modeShift[5] = {0, 6, 10, 12, 14}; whRatio = Abs(Log2(Wp / Hp)) if Wp>Hp and1 <predMode<2 + modeShift[whRatio] predMode = predMode + 65; else if Hp> Wp and (66 – modeShift[whRatio]) <predMode<= 66 predMode = predMode - 67; The above mapping is valid as long as the aspect ratio of the sub-partition is within the valid range given in Table 2. If Wp / Hp > 16 or Wp / Hp < 1 / 16, we recommend using the parent CU size for wide-angle derivation, as done in VTM.

[0064] In a simple variation of the above method, if the sub-partition has a square shape (i.e., if Wp = Hp), then the usual prediction pattern is applied; otherwise, the width angle is obtained using the dimensions of the CU. This simplification avoids applying the width angle prediction pattern (derived from the parent rectangle CU) to square sub-partitions. In other words, the width angle is derived from the sub-partition dimensions only when the sub-partition is square; otherwise, it is derived from the CU dimensions. Since there is no width angle for a square shape, this is equivalent to not having a width angle derivation for square sub-partitions while maintaining the integrity of the width angle derivation for rectangular sub-partitions. This is also presented in Example 1 below.

[0065] Since all sub-partitions in the ISP have the same size, the wide-angle pattern calculated for one sub-partition is the same for all sub-partitions.

[0066] Missing reference sample replacement The second improvement to the ISP is the use of reference samples from the top or left of the parent CU. Here, it is assumed that sub-partitions are processed in the normal order as in VTM 4.0. For horizontal splits, the top reference sample after sub-partition 2 uses decoded pixels from the previous sub-partition. Since the previous partition has the same width as the parent CU, the top-right reference sample is completely unavailable. In this case, the current VTM codec will use padding, where the last available reference pixel is repeated through the top-right portion of length W. This... Figure 10 The case of two sub-partitions is explained below. A similar case applies to vertical splits, where the lower left reference sample will be missing for sub-partition 2 and beyond. In this case, the current VTM codec will use padding, where the last available reference pixel is repeated through the lower left portion of length H.

[0067] These missing reference samples can be replaced from the top reference of the parent CU, such as... Figure 11 As explained, this is used for horizontal splitting into two sub-partitions. That is, after predicting the direction, the missing samples are mapped from the reference array of the corresponding parent CU. If some reference samples are still missing even after the replacement process, a normal filling process can be applied, in which the last reference samples are repeated to fill the reference array.

[0068] Missing samples mapped from the reference array of the corresponding parent CU after the predicted direction can be obtained using the same interpolation process as the prediction (linear, 4-tap...). To reduce complexity, the nearest neighbor can also be used (without interpolation). In the latter case, it corresponds to shifting the reference sample by the number of pixels corresponding to the predicted angle.

[0069] This reference sample replacement process has an undesirable effect: it can produce intensity discontinuities at the locations where pixels are copied from the top reference of the parent CU. To mitigate this effect, a simple low-pass filter can be used. For example, we can use a 3-tap filter [1 2 1] / 4.

[0070] exist Figure 8 In this approach, the mapped samples on the top reference are located along the prediction direction. Another method is to copy all the top-right reference samples of the CU onto the top-right reference samples of the second or subsequent sub-partitions, and then perform low-pass filtering at the top corner to mitigate discontinuities. This is in Figure 12 As shown in the image.

[0071] This replacement process is independent of whether the WAIP modifications from the previous section were applied. It can be applied to each subpartition without WAIP modifications, or it can be applied together with them.

[0072] In the following text, we assume the use of any general video codecs for ISP and WAIP. The VTM codec is an example of such a codec.

[0073] Example 1: In this example, if the target block is split in the ISP, then for any prediction mode, we use the size of the parent CU to determine the width angle, unless the sub-partition has a square shape. If the sub-partition has a square shape, then the regular prediction mode is applied to the sub-partition without mapping to the width angle. All other codec parameters remain unchanged as in VTM 4.0.

[0074] Example 2: In this example, if the target block is split in the ISP, for any prediction mode, we use the size of the parent CU to determine the width angle, as in VTM 4. However, starting from the second sub-partition, we replace the missing reference sample on the upper right (lower left) side from the top (left) reference array of the parent CU for horizontal (vertical) splitting. The reference sample on the top (left) reference array is positioned along the prediction direction, as shown below. Figure 11 As shown. All other codec parameters remain unchanged in VTM 4.0.

[0075] Example 3: In this example, if the target block is split in the ISP, for any prediction mode, we use the size of the parent CU to determine the width angle, as in VTM 4. However, starting from the second sub-partition, we replace the missing reference sample in the upper right (lower left) position with the top (left) reference array of the parent CU for horizontal (vertical) splitting. Figure 12As shown, after the second sub-partition, the reference samples on the upper (left) reference array are directly copied to the upper right (lower left) part of the reference array. All other codec parameters remain unchanged in VTM 4.0.

[0076] Example 4: In this example, if the target block is segmented in the ISP, then for any prediction mode, we use the size of the parent CU to determine the width angle, unless the sub-partition has a square shape. If the sub-partition has a square shape, then the regular prediction mode is applied to the sub-partition without mapping to the width angle. Additionally, starting from the second sub-partition, we replace the missing reference samples in the upper right (lower left) position from the reference array at the top (left) of the parent CU for horizontal (vertical) segmentation. The reference sample mapping is performed as in Example 2 or Example 3, while all other codec parameters in VTM 4.0 remain unchanged.

[0077] Example 5: In this example, if the target block is split in the ISP, then for any prediction mode, we use the size of the sub-partition instead of the size of the parent CU to determine the width angle. If the aspect ratio (Wp / Hp) of the sub-partition is greater than 16 or less than 1 / 16, the W / H ratio of the parent CU will be used to calculate the width angle and applied to all sub-partitions. Additionally, starting from the second sub-partition, we replace the missing reference sample in the upper right (lower left) position with the reference array at the top (left) of the parent CU for horizontal (vertical) splitting. The reference sample on the top (left) reference array is positioned along the prediction direction, such as... Figure 11 As shown, all other codec parameters in VTM 4.0 remain unchanged. To accelerate the prediction step, a reference array can be used for sub-partitions whose size is twice that of the parent CU. When using the CU size, the original reference array length can remain unchanged.

[0078] Example 6: In this example, if the target block is split in the ISP, then for any prediction mode, we use the size of the sub-partition instead of the size of the parent CU to determine the width angle. If the aspect ratio (Wp / Hp) of the sub-partition is greater than 16 or less than 1 / 16, the W / H ratio of the parent CU is used to calculate the width angle and applied to all sub-partitions. Additionally, starting from the second sub-partition, we replace the missing reference sample from the top (left) reference of the parent CU to the upper right (lower left) for horizontal (vertical) splitting. Figure 12 As shown, after the second sub-partition, the reference samples on the top (left) reference array are directly copied to the upper right (lower left) portion of the reference array, as all other codec parameters in VTM 4.0 remain unchanged. To accelerate the prediction step, the reference array can be used for sub-partitions whose size is twice that of the parent CU. When using the CU size, the original reference array length can remain unchanged.

[0079] Example 7: VTM 4.0 uses an ISP with 2 or 4 sub-partitions, provided that each sub-partition has at least 16 pixels. In this example, we relax these constraints and apply the ISP to all block shapes, provided that each sub-partition can be at least a single row or a single column. Depending on the CU size, the number of sub-partitions in the CU can also be higher, such as 8, 16, or 32. Regardless of whether the split is horizontal or vertical, the shape of the sub-partitions is restricted to be the same. If the aspect ratio (Wp / Hp) of a sub-partition is greater than 16 or less than 1 / 16, the W / H ratio of the parent CU is used to calculate the width angle and applied to all sub-partitions. Starting from the second sub-partition, we replace the missing reference sample in the upper right (lower left) position with the reference from the top (left) position of the parent CU for horizontal (vertical) splits. The mapping of the missing sample is performed as in Example 2 or Example 3.

[0080] Example 8: In this example, we relax the constraint that all sub-partitions must have the same shape. Sub-partitions can have unequal shapes, but are constrained to have height and width that are powers of 2. Therefore, the number of sub-partitions does not need to be a power of two as in Examples 1-7. For example, an 8×8 block can be divided into three partitions of sizes 4×8, 2×8, and 2×8. Any of the examples in Examples 1-7 can be changed accordingly.

[0081] Example 9: VTM 4.0 uses ISP on blocks without MRL. In this example, we relax this constraint. In general, ISP can be applied even if the CU uses multiple reference lines. In one variation, if the split is horizontal and the prediction direction is vertical, or if the split is vertical and the prediction direction is horizontal, then only the first sub-partition can use MRL. In another variation, all sub-partitions can use MRL regardless of the split type and prediction direction. In all these variations, we can use any of Examples 1 to 8 to calculate the wide angle for the sub-partition and / or map the missing reference samples from the second sub-partition onwards.

[0082] Example 10: VTM 4.0 uses ISP only on the luminance component. In this example, we relax this constraint and apply any of Examples 1-9 to both the luminance and chrominance components.

[0083] The proposed wide-angle mode derivation utilizing the ISP was implemented by incorporating changes in the VTM 4.0 reference software, as in Example 1. Testing was performed using a single frame from the JVET test sequence across all intra-frame (AI) configurations. The BD rate performance of the test method relative to the VTM 4.0 anchor results is shown in Table 3. It can be seen that the luminance BD rate is improved by 0.04% without increasing encoder and decoder complexity. The gains for Class C and Class E sequences are noteworthy, with a luminance gain of 0.11%.

[0084] Table 3: Example 1 compared to a VTM 4.0 anchor with a single frame from a JVET test sequence configured by AI. BD rate performance.

[0085]

[0086] The general aspect described aims to improve intra-frame prediction efficiency in the ISP by modifying the wide-angle derivation and replacing some missing reference samples in a sub-partition with reference samples from the parent CU. One advantage is its high compression efficiency without much additional complexity.

[0087] Figure 13 This document illustrates an embodiment of a method 1300 for encoding video data blocks using the general aspects described herein. The method begins at a start block 1301, and control proceeds to a function block 1310 for determining the intra-prediction direction of a rectangular sub-partition of the video block by: determining it based on the size of the sub-partition when the ratio of the sub-partition size is within a specific range; and determining it based on the size of the video block when the ratio of the sub-partition size is outside the specific range. Control then proceeds from block 1310 to block 1320 for mapping reference samples from the upper right or lower left of the video block to upper right or lower left reference samples for predicting the sub-partitions of the video block. Control proceeds from block 1320 to block 1330 for predicting samples of the rectangular sub-partition using reference samples from the row above the video block or from the column to the left of the video block, wherein the number of reference samples in the row above the video block or the column to the left of the video block is determined based on the size of the rectangular sub-partition. Control moves from box 1330 to box 1340 to encode video blocks using predictions in intra-frame decoding mode.

[0088] Figure 14This document illustrates an embodiment of a method 1400 for decoding a video data block using the general aspects described herein. The method begins at a start block 1401, and control proceeds to a function block 1410 for determining the intra-prediction direction of a rectangular sub-partition of the video block by: determining it based on the size of the sub-partition when the ratio of the sub-partition size is within a specific range; and determining it based on the size of the video block when the ratio of the sub-partition size is outside the specific range. Control then proceeds from block 1410 to block 1420 for mapping reference samples from the upper right or lower left of the video block to reference samples for predicting the upper right or lower left of the sub-partition of the video block. Control proceeds from block 1420 to block 1430 for predicting samples of the rectangular sub-partition using reference samples from a row above the video block or from a column to the left of the video block, wherein the number of reference samples in the row above the video block or the column to the left of the video block is determined based on the size of the rectangular sub-partition. Control moves from box 1430 to box 1440 to encode the video block using predictions in intra-frame decoding mode.

[0089] Figure 15 An embodiment of an apparatus 1500 for encoding or decoding video data blocks is shown. The apparatus includes a processor 1510 and can be interconnected to a memory 1520 via at least one port. Both the processor 1510 and the memory 1520 may also have one or more additional interconnects to external connections.

[0090] The processor 1510 is configured to encode or decode video data using an intra-frame prediction mode employing sub-partitions, or to encode or decode blocks of video data using an intra-frame prediction mode employing sub-partitions.

[0091] The general aspects described aim to improve intra-prediction efficiency by employing intra-prediction in sub-partitions and the necessary changes required for these modes. The advantage is higher compression efficiency without much additional complexity.

[0092] This application describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described as specific and are generally described in a manner that may sound restrictive, at least to illustrate individual characteristics. However, this is for the purpose of clarity and does not limit the application or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide other aspects. Furthermore, these aspects can also be combined and interchanged with those described in earlier documents.

[0093] The aspects described and anticipated in this application can be implemented in many different forms. Figure 16 , 17Some embodiments are provided in 18, but other embodiments are conceived, and... Figure 16 , 17 The discussion in section 18 does not limit the breadth of implementation. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having a bitstream generated according to any of the described methods stored thereon.

[0094] In this application, the terms "reconstruction" and "decoding" are used interchangeably, as are the terms "pixel" and "sample," and the terms "image," "picture," and "frame." Typically, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.

[0095] This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions may be modified or combined.

[0096] The various methods and other aspects described in this application can be used to modify the module, for example, such as Figure 16 , 17 The intra-frame prediction, entropy decoding, and / or decoding modules (160, 360, 145, 330) of the video encoder 100 and decoder 200 shown in Figure 18. Furthermore, aspects of the invention are not limited to VVC or HEVC and can be applied to, for example, other standards and recommendations, whether pre-existing or developed in the future, and any extensions to such standards and recommendations (including VVC and HEVC). Unless otherwise indicated or technically excluded, the aspects described in this application may be used alone or in combination.

[0097] Various numerical values ​​are used in this application. Specific values ​​are for illustrative purposes, and the aspects described are not limited to these specific values.

[0098] Figure 16 Encoder 100 is shown. Variations of encoder 100 can be envisioned, but for clarity, encoder 100 is described below without describing all anticipated variations.

[0099] Before being encoded, the video sequence may undergo pre-coding (101), for example, applying a color transformation to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input image components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization with one of the color components). Metadata may be associated with the pre-processing and appended to the bitstream.

[0100] In encoder 100, the image is encoded by encoder elements as described below. The image to be encoded is partitioned (102) in units such as CUs. Each unit is encoded using, for example, intra-frame or inter-frame modes. When a unit is encoded in intra-frame mode, intra-frame prediction (160) is performed. In inter-frame mode, motion estimation (175) and compensation (170) are performed. The encoder decides (105) whether to use intra-frame or inter-frame mode to encode the unit, and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (110) the prediction block from the original image block.

[0101] Then, the predicted residual is transformed (125) and quantized (130). The quantized transform coefficients, motion vectors, and other syntax elements are entropy decoded (145) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass the transform and quantization, i.e., decode the residual directly without applying transform or quantization.

[0102] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the prediction residual. The decoded prediction residual and prediction block are combined (155) to reconstruct an image block. An in-loop filter (165) is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering, thereby reducing coding artifacts. The filtered image is stored in a reference image buffer (180).

[0103] Figure 17 A block diagram of a video decoder 200 is shown. In decoder 200, the bitstream is decoded by decoder elements as described below. The video decoder 200 typically performs operations related to... Figure 16 The encoding rounds described herein are the inverse of the decoding rounds. Encoder 100 typically also performs video decoding as part of the encoded video data.

[0104] Specifically, the decoder's input includes a video bitstream, which can be generated by the video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other decoded information. Picture partitioning information indicates how the picture is partitioned. The decoder can therefore partition (235) the picture based on the decoded picture partitioning information. The transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residuals. The decoded prediction residuals are combined with the prediction blocks (255) to reconstruct the image blocks. The prediction blocks can be obtained from intra-frame prediction (260) or motion-compensated prediction (i.e., inter-frame prediction) (275) (270). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).

[0105] The decoded image can undergo further post-decoding processing (285), such as inverse color transformation (e.g., conversion from YCbCr4:2:0 to RGB 4:4:4) or inverse remapping of the remapping process performed in pre-encoding processing (101). Post-decoding processing can utilize metadata derived in pre-encoding processing and signaled in the bitstream.

[0106] Figure 18 A block diagram illustrating an example system in which various aspects and embodiments are implemented is shown. System 1000 can be implemented as a device including the various components described below and configured to perform one or more aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 1000 can be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more aspects described herein.

[0107] System 1000 includes at least one processor 1010 configured to execute instructions loaded thereon for implementing various aspects described herein, such as those described herein. Processor 1010 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). System 1000 includes a storage device 1040, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 1040 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0108] System 1000 includes an encoder / decoder module 1030 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents one or more modules that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both encoding and decoding modules. Alternatively, the encoder / decoder module 1030 may be implemented as a separate element of system 1000 or may be incorporated into processor 1010 as a combination of hardware and software as known to those skilled in the art.

[0109] Program code to be loaded onto processor 1010 or encoder / decoder 1030 to execute the various aspects described herein may be stored in storage device 1040 and subsequently loaded onto memory 1020 for execution by processor 1010. According to various embodiments, one or more of processor 1010, memory 1020, storage device 1040, and encoder / decoder module 1030 may store one or more of various items during the execution of the processes described herein. These stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing equations, formulas, operations, and operational logic.

[0110] In some embodiments, the memory within processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., the processing device may be processor 1010 or encoder / decoder module 1030) is used for one or more of these functions. External memory may be memory 1020 and / or storage device 1040, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory, such as RAM, is used as working memory for video decoding and decoding operations, such as for MPEG-2 (MPEG stands for Moving Picture Experts Group; MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Decoding; also known as H.265 and MPEG-H Part 2), or VVC (Various Video Decoding, a new standard developed by JVET (Joint Video Team Experts)).

[0111] As shown in box 1130, inputs to the components of system 1000 can be provided through various input devices. Such input devices include, but are not limited to: (i) an RF section that receives radio frequency (RF) signals transmitted over the air, for example by a broadcaster; (ii) a component (COMP) input terminal (or a set of COMP input terminals); (iii) a universal serial bus (USB) input terminal; and / or (iv) a high-resolution input. Clarity Multimedia Interface (HDMI) input terminal. Figure 18 Other examples not shown include composite videos.

[0112] In various embodiments, the input device of block 1130 has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal band to a band), (ii) down-converting the selected signal, (iii) further band-limiting to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements to perform these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the aforementioned (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0113] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 1000 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) can be implemented as needed, for example, within a separate input processing IC or processor 1010. Similarly, various aspects of USB or HDMI interface processing can be implemented as needed within a separate interface IC or within processor 1010. Demodulated, error-corrected, and demultiplexed streams are provided to various processing elements (including, for example, processor 1010) and encoder / decoder 1030 (which operates in conjunction with memory and storage elements) to process the data streams as needed for presentation on the output device.

[0114] Various components of the system 1000 can be housed within an integrated housing. Within the integrated housing, the various components can be interconnected and transmit data therebetween using suitable connection arrangements (e.g., internal buses (including inter-IC (I2C) buses), wiring, and printed circuit boards known in the art).

[0115] System 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.

[0116] In various embodiments, a wireless network (e.g., a Wi-Fi network, such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)) is used to stream or otherwise provide data to system 1000. In these embodiments, the Wi-Fi signal is received via a communication channel 1060 and a communication interface 1050 suitable for Wi-Fi communication. The communication channel 1060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box that transmits data via an HDMI connection in input block 1130 to provide streaming data to system 1000. Still other embodiments use an RF connection in input block 1130 to provide streaming data to system 1000. As described above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0117] System 1000 can provide output signals to various output devices, including a display 1100, a speaker 1110, and other peripheral devices 1120. The display 1100 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a flexible display, and / or a foldable display. The display 1100 can be used in a television, tablet computer, laptop computer, cellular phone (mobile phone), or other device. The display 1100 can also be integrated with other components (e.g., as in a smartphone) or stand alone (e.g., as an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 1120 include one or more of a standalone digital video disc (or digital multifunction disc) (for both, this can be simply referred to as a DVR), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120 that provide functionality based on the output of system 1000. For example, a disc player performs the function of playing the output of system 1000.

[0118] In various embodiments, signaling is used to transmit control signals between system 1000 and display 1100, speaker 1110, or other peripheral devices 1120 using communication protocols such as AV links, consumer electronics controls (CEC), or other communication protocols that enable device-to-device control with or without user intervention. Output devices may be communicatively coupled to system 1000 via dedicated connections through corresponding interfaces 1070, 1080, and 1090. Alternatively, output devices may be connected to system 1000 via communication interface 1050 using communication channel 1060. Display 1100 and speaker 1110 may be integrated into a single unit in an electronic device (e.g., a television set) along with other components of system 1000. In various embodiments, display interface 1070 includes a display driver, such as a timing controller ((T-control) chip).

[0119] For example, if the RF portion of input 1130 is part of a separate set-top box, then display 1100 and speaker 1110 may alternatively be separated from one or more other components. In various embodiments where display 1100 and speaker 1110 are external components, the output signal may be provided via a dedicated output connection, such as an HDMI port, USB port, or COMP output.

[0120] These embodiments can be implemented by processor 1010 or computer software implemented by a combination of hardware and software. As a non-limiting example, embodiments can be implemented by one or more integrated circuits. Memory 1020 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. Processor 1010 can be of any type suitable for the technical environment and can include one or more of microprocessors, general-purpose computers, special-purpose computers, and multi-core architecture-based processors, as non-limiting examples.

[0121] Various implementations involve decoding. As used herein, "decoding" can include, for example, all or part of a process performed on a received encoded sequence to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include processes performed by decoders of the various implementations described herein.

[0122] As a further example, in one embodiment, "decoding" refers only to entropy decoding; in another embodiment, "decoding" refers only to differential decoding; and in yet another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the specific context of the description and is believed to be fully understood by those skilled in the art.

[0123] Various implementations involve encoding. In a manner similar to the above discussion of “decoding,” “encoding,” as used herein, can include all or part of a process performed, for example, on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various embodiments, such processes also, or alternatively, include processes performed by encoders of the various implementations described herein.

[0124] As a further example, in one embodiment, “encoding” refers only to entropy encoding; in another embodiment, “encoding” refers only to differential encoding; and in yet another embodiment, “encoding” refers to a combination of differential and entropy encoding. Whether the phrase “encoding process” is intended to specifically refer to a subset of operations or generally to a broader encoding process will become clear from the context of the specific description and is believed to be fully understood by those skilled in the art.

[0125] Note that the syntax elements used here are descriptive terms. Therefore, the use of other syntax element names is not excluded.

[0126] When the accompanying drawings are presented as flowcharts, it should be understood that they also provide block diagrams of the corresponding devices. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that they also provide flowcharts of the corresponding methods / processes.

[0127] Various implementations may involve parametric modeling or rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, usually with constraints on computational complexity. This can be measured by rate-distortion optimization (RDO), or by least mean square (LMS), mean absolute error (MAE), or other such measures. Rate-distortion optimization is typically formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different approaches exist to address the rate-distortion optimization problem. For example, these approaches can be based on extensive testing of all coding options (including all considered modes or coding parameter values), providing a complete evaluation of their decoding costs and the associated distortion of the reconstructed signal after decoding and decoding. Faster methods can also be used to save coding complexity, particularly by calculating approximate distortion based on predicting or predicting the residual signal rather than the reconstructed signal. A hybrid of these two approaches can also be used, for example, by using approximate distortion only for some possible coding options and full distortion for others. Other methods evaluate only a subset of possible coding options. More generally, many methods employ any of a variety of techniques to perform optimization, but optimization is not necessarily a complete assessment of both coding cost and associated distortion.

[0128] The implementations and aspects described herein can be implemented, for example, in methods or processes, apparatuses, software programs, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the features in question can also be implemented in other forms (e.g., apparatuses or programs). For example, an apparatus can be implemented with appropriate hardware, software, and firmware. The method can be implemented, for example, in a processor, which generally refers to a processing device (including, for example, a computer, microprocessor, integrated circuit, or programmable logic device). Processors also include communication devices, such as computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0129] The references to "one embodiment," "one implementation," or "implementation," and other variations, mean that a particular feature, structure, characteristic, etc., described in connection with the embodiment is included in at least one embodiment. Therefore, the phrases "in one embodiment," "in an embodiment," "in an implementation," or "in an implementation," and any other variations appearing in various places throughout this application, do not necessarily refer to the same embodiment.

[0130] Additionally, this application may involve "determining" various types of information. Determining information may include, for example, one or more of the following: estimation information, calculation information, prediction information, or information retrieved from memory.

[0131] Furthermore, this application may relate to "accessing" various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or one or more of these.

[0132] Additionally, this application can refer to "receiving" various types of information. Like "accessing," receiving is intended to be a broad term. Receiving information can include, for example, accessing information or retrieving information (e.g., from memory). Furthermore, "receiving" is generally referred to in one or more ways during operations such as storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0133] It should be understood that, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one” is intended to cover the selection of only the first listed option (A), or only the selection of only the second listed option (B), or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” such wording is intended to include selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to multiple listed items.

[0134] Furthermore, as used herein, the term "signaling" specifically refers to instructing the corresponding decoder to do something. For example, in some embodiments, the encoder signals a particular one of multiple transforms, decoding modes, or tags. Thus, in one embodiment, the same transform, parameter, or mode is used on both the encoder and decoder sides. Therefore, for example, the encoder can send (explicit signaling) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling can be used without sending it (implicit signaling) to simply allow the decoder to know and select the specific parameter. Bit savings are achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. While the foregoing refers to the verb form of the term "signaling," the term "signaling" can also be used as a noun herein.

[0135] As will be apparent to those skilled in the art, implementations can generate various signals formatted to carry information, such as information that can be stored or transmitted. This information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave using the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via various wired or wireless links. The signal may be stored on a processor-readable medium.

[0136] We have described several embodiments across various claim classes and types. Features of these embodiments may be provided individually or in any combination. Furthermore, embodiments may include one or more of the following features, devices, or aspects, individually or in any combination, across various claim classes and types: • A process or apparatus for performing intra-frame encoding and decoding using intra-frame decoding sub-partitions.

[0137] • A process or apparatus for performing intra-frame encoding and decoding using multiple reference lines through intra-frame decoding sub-partitions.

[0138] • A process or apparatus for performing intra-frame encoding and decoding using intra-frame decoding sub-partitions and missing sample replacements.

[0139] • A process or apparatus for performing intra-frame encoding and decoding using intra-frame decoding sub-partitions and a mapping of reference pixels above a parent video block.

[0140] • A process or apparatus for performing intra-frame encoding and decoding using intra-frame decoding sub-partitions and unequal sub-partition sizes.

[0141] • A bit stream or signal that includes one or more of the described syntax elements or variations thereof.

[0142] • A bitstream or signal that includes grammatical information generated according to any of the described embodiments.

[0143] • Create and / or send and / or receive and / or decode according to any of the embodiments described.

[0144] • A method, process, apparatus, medium for storing instructions, medium for storing data, or signal according to any of the described embodiments.

[0145] • Inserting into the signaling enables the decoder to determine the syntax elements of the decoding mode in a manner corresponding to that used by the encoder.

[0146] • Creating and / or sending and / or receiving and / or decoding bitstreams or signals that include one or more of the described syntax elements or their variants.

[0147] • A TV, set-top box, cellular phone, tablet computer or other electronic device that performs one or more transformation methods according to any of the described embodiments.

[0148] • A TV, set-top box, cellular phone, tablet computer, or other electronic device that performs a transformation method (one or more) according to any of the described embodiments to determine and display (e.g., using a monitor, screen, or other type of display) the resulting image.

[0149] • A TV, set-top box, cellular phone, tablet computer, or other electronic device that selects, band-limits, or tunes (e.g., using a tuner) a channel to receive a signal including an encoded image, and performs a transformation method (one or more) according to any of the described embodiments.

[0150] A TV, set-top box, cellular phone, tablet computer or other electronic device that receives over the air (e.g., using an antenna) a signal including an encoded image and performs a transformation method (one or more).

Claims

1. A method comprising: The intra-prediction direction of a rectangular sub-partition of a video block is determined as follows: when the ratio of the sub-partition size is within a specific range, it is determined based on the size of the sub-partition; and when the ratio of the sub-partition size is outside the specific range, it is determined based on the size of the video block. Map the reference sample from the upper right or lower left of the video block to the upper right or lower left reference sample for the prediction of the sub-partition of the video block; The samples of the rectangular sub-partition are predicted using reference samples from the row above the video block or from the column to the left of the video block, wherein the number of reference samples in the row above the video block or the column to the left of the video block is determined based on the size of the rectangular sub-partition; and, The prediction is used in intra-frame decoding mode to encode the rectangular sub-partitions of the video block.

2. An apparatus comprising: The processor, which is configured to execute: The intra-prediction direction of a rectangular sub-partition of a video block is determined as follows: when the ratio of the sub-partition size is within a specific range, it is determined based on the size of the sub-partition; and when the ratio of the sub-partition size is outside the specific range, it is determined based on the size of the video block. Map the reference sample from the upper right or lower left of the video block to the upper right or lower left reference sample for the prediction of the sub-partition of the video block; The samples of the rectangular sub-partition are predicted using reference samples from the row above the video block or from the column to the left of the video block, wherein the number of reference samples in the row above the video block or the column to the left of the video block is determined based on the size of the rectangular sub-partition; and, The prediction is used in intra-frame decoding mode to encode the rectangular sub-partitions of the video block.

3. A method comprising: The intra-prediction direction of a rectangular sub-partition of a video block is determined as follows: when the ratio of the sub-partition size is within a specific range, it is determined based on the size of the sub-partition; and when the ratio of the sub-partition size is outside the specific range, it is determined based on the size of the video block. Map the reference sample from the upper right or lower left of the video block to the upper right or lower left reference sample for the prediction of the sub-partition of the video block; The samples of the rectangular sub-partition are predicted using reference samples from the row above the video block or from the column to the left of the video block, wherein the number of reference samples in the row above the video block or the column to the left of the video block is determined based on the size of the rectangular sub-partition; and, The prediction is used in intra-frame decoding mode to decode the rectangular sub-partition of the video block.

4. An apparatus comprising: The processor, which is configured to execute: The intra-prediction direction of a rectangular sub-partition of a video block is determined as follows: when the ratio of the sub-partition size is within a specific range, it is determined based on the size of the sub-partition; and when the ratio of the sub-partition size is outside the specific range, it is determined based on the size of the video block. Map the reference sample from the upper right or lower left of the video block to the upper right or lower left reference sample for the prediction of the sub-partition of the video block; The samples of the rectangular sub-partition are predicted using reference samples from the row above the video block or from the column to the left of the video block, wherein the number of reference samples in the row above the video block or the column to the left of the video block is determined based on the size of the rectangular sub-partition; and, The prediction is used in intra-frame decoding mode to decode the rectangular sub-partition of the video block.

5. The method according to claim 1 or 3, or the apparatus according to claim 2 or 4, wherein the non-wide-angle prediction mode is applied to the sub-partitions of the square.

6. The method of claim 1 or 3, or the apparatus of claim 2 or 4, wherein for sub-partitions beyond the sub-partition closest to the left reference array for the vertical sub-partition or beyond the sub-partition closest to the top edge for the horizontal sub-partition, missing reference samples are mapped from the video block.

7. The method or apparatus of claim 6, wherein the intra-frame prediction direction is based on the size of the video block for a sub-partition having a width / height ratio greater than 16 or less than 1 / 16.

8. The method of claim 1 or 3, or the apparatus of claim 2 or 4, wherein the intra-frame prediction direction is based on the size of the video block for a sub-partition having a width / height ratio greater than 16 or less than 1 / 16, and wherein for a sub-partition beyond the sub-partition closest to the left reference array for the vertical sub-partition or beyond the sub-partition closest to the top edge for the horizontal sub-partition, missing reference samples are directly mapped from the video block horizontally or vertically.

9. The method or apparatus of claim 6, wherein the sub-partition is at least a single row or a single column of the video block.

10. The method according to claim 1 or 3, or the apparatus according to claim 2 or 4, wherein the sub-partitions have unequal shapes but have a height and width that are powers of two.

11. The method of claim 1 or 3, or the apparatus of claim 2 or 4, wherein the sub-partition can use multiple reference lines.

12. An apparatus comprising: The apparatus according to any one of claims 4 to 11; as well as At least one of the following: (i) an antenna configured to receive a signal comprising the video block; (ii) a bandwidth limiter configured to limit the received signal to the bandwidth comprising the video block; and (iii) a display configured to display the output representing a video block.

13. A non-transitory computer-readable medium comprising data content generated by the method of any one of claims 1 to 11 or by the apparatus of any one of claims 2 to 11, for playback using a processor.

14. A signal comprising video data generated by the method of any one of claims 1 to 11 or by the apparatus of any one of claims 2 to 11, for playback using a processor.

15. A computer program product comprising instructions which, when executed by a computer, cause the computer to perform the method according to any one of claims 1, 3, and 5 to 11.