Prediction within extended reference image

The video codec utilizes extended reference samples and filtering techniques to enhance prediction efficiency and accuracy, addressing the inefficiencies in existing codecs by optimizing parallel processing in video encoding and decoding.

JP2026063090APending Publication Date: 2026-04-10FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2026-01-13
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing video codecs like H.265/HEVC do not efficiently utilize parallel processing capabilities for reference samples in video encoding and decoding, particularly in cases where nearest reference samples are unavailable.

Method used

Implementing a video codec that uses extended reference samples separated from prediction blocks by nearest reference samples, with the ability to replace unavailable nearest samples and apply filtering techniques to enhance prediction accuracy.

Benefits of technology

Enhances prediction efficiency by utilizing extended reference samples, improving processing speed and accuracy even when nearest samples are unavailable, thereby optimizing video encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063090000001_ABST
    Figure 2026063090000001_ABST
Patent Text Reader

Abstract

This provides a video encoder that encodes images from a video into encoded data using block-based predictive coding, including in-image prediction. [Solution] For in-image prediction, the video encoder uses multiple nearest reference samples 1062, 1064 of the image directly adjacent to the prediction block 1066 and multiple extended reference samples to encode the prediction block 1066 of the image, wherein each of the multiple extended reference samples is separated from the prediction block by the nearest reference sample having at least one index i=0 among the multiple reference samples. The video encoder further sequentially determines the availability or unavailability of each of the multiple nearest reference samples, replaces the nearest reference sample determined to be unavailable with a replacement sample, and uses the replacement sample for in-image prediction.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to video coding, particularly hybrid video coding including in-image prediction. The invention further relates to a video encoder, a video decoder, and a method for coding and decoding video, respectively. [Background technology]

[0002] H.265 / HEVC is a video codec that already provides tools to improve or enable parallel processing in encoders and / or decoders. For example, HEVC supports the subdivision of an image into an array of independently encoded tiles. Another concept supported by HEVC is related to WPP, which allows CTU rows or lines of an image to be processed in parallel from left to right (i.e., in stripes) as long as the smallest CTU offset is maintained in the processing of consecutive CTU lines. However, it is preferable to have a video codec at hand that more efficiently supports the parallel processing capabilities of video encoders and / or video decoders. [Overview of the Initiative] [Problems that the invention aims to solve]

[0003] Therefore, an object of the present invention is to provide a video codec that enables more efficient processing in encoders and / or decoders with respect to reference samples used to predict prediction blocks. [Means for solving the problem]

[0004] This objective is achieved by the subject matter of the independent claims of this application.

[0005] According to the embodiment, the video encoder is configured to encode images of a video into encoded data by block-based predictive coding, the block-based predictive coding includes in-image prediction. For in-image prediction, the video encoder is configured to use a plurality of extended reference samples of the image to encode prediction blocks of the image, each extended reference sample of the plurality of extended reference samples being separated from the prediction block by at least one nearest reference sample of the plurality of reference samples that is directly adjacent to the prediction block. The video encoder is further configured to sequentially determine the availability or unavailability of each of the plurality of nearest reference samples and to replace the nearest reference sample that is determined to be unavailable with a replacement sample. The video encoder is configured to use replacement samples for in-image prediction. This allows the use of the concept of prediction using the nearest reference sample even when such a sample is unavailable, which may occur, for example, when there is a line or row of samples in a buffer / memory, but there is not actually a column of samples in memory, and therefore it is unavailable.

[0006] According to the embodiment, the video encoder is configured to encode images of a video into encoded data by block-based predictive coding, the block-based predictive coding includes in-image prediction. In the in-image prediction, the video encoder is configured to use a plurality of extended reference samples of the image to encode prediction blocks of the image, each extended reference sample of the plurality of extended reference samples being separated from the prediction block by at least one nearest reference sample of the plurality of reference samples that is directly adjacent to the prediction block. The video encoder is further configured to use a plurality of filtered extended reference samples for in-image prediction, by filtering at least a subset of the plurality of extended reference samples using a bilateral filter to obtain a plurality of filtered extended reference samples.

[0007] According to the embodiment, the video encoder is configured to encode images of a video into encoded data by block-based predictive coding, the block-based predictive coding includes in-image prediction. In in-image prediction, the video encoder uses a plurality of extended reference samples of the image to encode prediction blocks of the image, each extended reference sample of the plurality of extended reference samples is separated from the prediction block by at least one nearest reference sample of the plurality of reference samples that is directly adjacent to the prediction block, the plurality of nearest reference samples are arranged along a first image direction of the prediction block and along a second image direction of the prediction block, and at least a portion of the nearest reference samples arranged along the second direction are mapped to the extended reference samples arranged along the first direction such that the mapped reference samples extend beyond the extension of the prediction block along the first image direction. The video encoder is configured to use the mapped extended reference samples for prediction.

[0008] According to the embodiment, the video encoder is configured to encode images of a video into encoded data by block-based predictive coding, the block-based predictive coding includes in-image prediction. For in-image prediction, the video encoder is configured to use a plurality of nearest reference samples of the image directly adjacent to the prediction block, and a plurality of extended reference samples, to encode the prediction block of the image, each extended reference sample of the plurality of reference samples being separated from the prediction block by at least one nearest reference sample of the plurality of reference samples, and the video encoder is configured to perform boundary filtering in a mode where extended samples are not used, and not to perform boundary filtering when extended samples are used, or the video encoder is configured to perform boundary filtering on at least a subset of the plurality of nearest reference samples and not to perform boundary filtering for extended samples.

[0009] According to the embodiment, the video encoder is configured to encode images of a video into encoded data by block-based predictive coding, the block-based predictive coding includes in-image prediction. In the in-image prediction, the video encoder is configured to determine a plurality of nearest reference samples of the image directly adjacent to the prediction block, and a plurality of extended reference samples, in order to encode the prediction block of the image, such that each extended reference sample of the plurality of reference samples is separated from the prediction block by at least one nearest reference sample of the plurality of extended reference samples. The video encoder is further configured to determine the prediction of the prediction block using the extended reference samples, filter the extended reference samples to obtain a plurality of filtered extended reference samples, and combine the prediction and the filtered extended reference samples to obtain a combined prediction of the prediction block.

[0010] According to the embodiment, the video encoder is configured to encode images of a video into encoded data by block-based predictive coding, the block-based predictive coding includes in-image prediction. For in-image prediction, the video encoder uses multiple nearest reference samples and / or multiple extended reference samples of the image directly adjacent to the prediction block to encode the prediction block of the image, each extended reference sample of the multiple extended reference samples being separated from the prediction block by at least one nearest reference sample of the multiple reference samples, and is configured to determine a first prediction of the prediction block using a first prediction mode of a set of prediction modes, the first set of prediction modes including a prediction mode that uses multiple nearest reference samples when there are no extended reference samples, and a second prediction mode of a second set of prediction modes including a subset of the prediction modes of the first set, the subset being associated with multiple extended reference samples, and is configured to determine a second prediction of the prediction block using a second prediction mode of a second set of prediction modes. i (x, y)) and weighted (w0;w i) are combined and configured to obtain a combined prediction (p(x, y)) as the prediction of the prediction block of the encoded data.

[0011] According to the embodiment, the video encoder encodes images of a video into encoded data by block-based predictive coding including in-image prediction, and for in-image prediction, uses multiple nearest reference samples and / or multiple extended reference samples of the image directly adjacent to the prediction block to encode the prediction block of the image, wherein each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample of the multiple reference samples, for example, if there are no extended reference samples, uses one of a first set of prediction modes for predicting the prediction block using the nearest reference sample, or uses the extended reference sample for predicting the prediction block. The decoder is configured to allow a conclusion in the decoder that a specific value of a parameter is selected or determined by using a prediction mode, which is one of a second set of modes, where the second set of prediction modes is a subset of the first set of prediction modes, notifying mode information (m) indicating the prediction mode to be used to predict a prediction block, then notifying parameter information (i) indicating a subset of extended reference samples to be used for the prediction mode if the prediction mode is included in the second set of prediction modes, and skipping notification of parameter information if the prediction mode to be used is not included in the second set of prediction modes, such that a certain property allows for the skipping of notifications, i.e., the absence of a signal is given a useful meaning. For example, absence may indicate that the nearest reference sample should be used.

[0012] According to the embodiment, the video encoder encodes images of a video into encoded data by block-based predictive coding including in-image prediction, and for in-image prediction, uses a plurality of reference samples and a plurality of extended reference samples to encode prediction blocks of an image, each extended reference sample of the plurality of extended reference samples is separated from the prediction block by at least one nearest reference sample of the plurality of reference samples, and is configured to use a prediction mode which is one of a first set of prediction modes for predicting a prediction block using the nearest reference sample, or a prediction mode which is one of a second set of prediction modes for predicting a prediction block using the extended reference sample, the second set of prediction modes being a subset of the first set of prediction modes. The video encoder can generate the first set and / or the second set using the available reference data, and / or determine the set using information derived from the image. The second set which is a subset of the first set includes the case where both sets are equal. The video encoder is configured to notify parameter information indicating a subset of multiple reference samples used for the prediction mode, where the subset of multiple reference samples includes only the nearest reference sample or extended reference samples, and then to notify mode information (m) indicating the prediction mode used to predict the prediction block, where the mode information indicates the prediction mode from a subset of modes, and the subset is limited to a set of prediction modes permitted according to parameter information (i). Identification of the limited set is possible because only the prediction modes associated with the indicated reference samples are applied based on the association of the reference samples used, i.e., the nearest or extended ones.

[0013] According to the embodiment, the video encoder encodes images of a video into encoded data by block-based predictive coding including in-image predictions, and for in-image predictions, uses multiple nearest reference samples and / or multiple extended reference samples of the image directly adjacent to a prediction block to encode a prediction block of the image, each extended reference sample of the multiple extended reference samples being separated from the prediction block by at least one nearest reference sample of the multiple reference samples, and determines a first prediction of a prediction block using a first prediction mode of a set of prediction modes, the first set of prediction modes including a prediction mode that uses multiple nearest reference samples when there are no extended reference samples, and determines a second prediction of a prediction block using a second prediction mode of a second set of prediction modes, the second set of prediction modes including a subset of the first set of prediction modes being associated with multiple extended reference samples. The video encoder is configured to combine the first and second predictions to obtain a combined prediction as a prediction of the prediction block in the encoded data.

[0014] According to the embodiment, the video encoder encodes images of a video into encoded data by block-based predictive coding including in-image prediction, and in in-image prediction, uses multiple extended reference samples of the image to encode prediction blocks of the image, each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample from a plurality of reference samples directly adjacent to the prediction block, and is configured to use multiple extended reference samples according to a predetermined set of multiple extended reference samples. Multiple extended reference samples may, for example, be included in a list of area indices identified by identifiers.

[0015] According to the embodiment, the video encoder is configured to encode a plurality of prediction blocks into encoded data by block-based predictive coding including in-image prediction, and to use a plurality of extended reference samples of the image to encode the prediction blocks of the plurality of prediction blocks for in-image prediction, wherein each extended reference sample of the plurality of extended reference samples is separated from the prediction block by at least one nearest reference sample of a plurality of reference samples directly adjacent to the prediction block. The video encoder is configured to determine that the extended reference samples are at least a partial part of the adjacent prediction blocks of the plurality of prediction blocks, to determine that the adjacent prediction blocks have not yet been predicted, and to notify information indicating that the extended prediction samples are associated with the prediction block and have been placed as unavailable samples for the adjacent prediction blocks.

[0016] According to the embodiment, the video decoder decodes an encoded image by encoding the data into video using block-based predictive decoding that includes in-image prediction, and in the in-image prediction, it uses multiple extended reference samples of the image to encode the prediction block of the image, each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample of multiple reference samples that is directly adjacent to the prediction block, the availability or unavailability of each of the multiple nearest reference samples is determined sequentially, the nearest reference sample determined to be unavailable is replaced with a replacement sample, and the replacement sample is used for in-image prediction.

[0017] According to an embodiment, a video decoder decodes an image encoded by encoding data into a video by block-based predictive decoding including intra prediction. In intra prediction, to decode a prediction block of an image, a plurality of extended reference samples of the image are used. Each extended reference sample of the plurality of extended reference samples is separated from the prediction block by one of the nearest reference samples among a plurality of reference samples that are at least directly adjacent to the prediction block. A bilateral filter is used to filter at least a subset of the plurality of extended reference samples to obtain a plurality of filtered extended reference samples, and is configured to use the plurality of filtered extended reference samples for intra prediction.

[0018] According to an embodiment, a video decoder decodes an image encoded by encoding data into a video by block-based predictive decoding including intra prediction. In intra prediction, to decode a prediction block of an image, a plurality of extended reference samples of the image are used. Each extended reference sample of the plurality of extended reference samples is separated from the prediction block by one of the nearest reference samples among a plurality of reference samples that are at least directly adjacent to the prediction block. The plurality of nearest reference samples are arranged along a first image direction of the prediction block and along a second image direction of the prediction block. At least a part of the nearest reference samples arranged along the second direction is mapped to extended reference samples arranged along the first direction so that the mapped reference samples exceed the extension of the prediction block along the first image direction, and is configured to use the mapped extended reference samples for prediction.

[0019] According to an embodiment, a video decoder decodes an image encoded by encoding data into a video by block-based predictive decoding including intra prediction. For intra prediction, in order to decode a prediction block of an image, the video decoder is configured to use a plurality of nearest reference samples and a plurality of extended reference samples of an image directly adjacent to the prediction block. Each extended reference sample of the plurality of extended reference samples is separated from the prediction block by at least one nearest reference sample of the plurality of reference samples. The video decoder performs boundary filtering in a mode where extended samples are not used. When extended samples are used, the video decoder does not use boundary filtering, or at least a subset of the plurality of nearest reference samples is boundary-filtered and boundary filtering is not used for the extended samples.

[0020] According to an embodiment, a video decoder decodes an image encoded by encoding data into a video by block-based predictive decoding including intra prediction. In intra prediction, in order to decode a prediction block of an image, the video decoder determines a plurality of nearest reference samples and a plurality of extended reference samples of an image directly adjacent to the prediction block. Each extended reference sample of the plurality of extended reference samples is separated from the prediction block by at least one nearest reference sample of the plurality of reference samples. The video decoder determines the prediction of the prediction block using the extended reference samples, filters the extended reference samples to obtain a plurality of filtered extended reference samples, and combines the prediction and the filtered extended reference samples to obtain a combined prediction of the prediction block.

[0021] According to the embodiment, the video decoder decodes an encoded image by encoding data into video using block-based predictive decoding that includes in-image prediction, and for in-image prediction, to decode the prediction block of the image, uses multiple nearest reference samples and / or multiple extended reference samples of the image directly adjacent to the prediction block, each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample of multiple reference samples, and determines a first prediction of the prediction block using a first prediction mode of a set of prediction modes, the first set of prediction modes includes a prediction mode that uses multiple nearest reference samples when there are no extended reference samples, and determines a second prediction of the prediction block using a second prediction mode of a second set of prediction modes, the second set of prediction modes includes a subset of the prediction modes of the first set, the subset being associated with multiple extended reference samples. The video decoder is configured to weight-combine the first and second predictions to obtain a combined prediction as the prediction of the prediction block in the encoded data.

[0022] According to the embodiment, the video decoder decodes an encoded image by encoding data into video by block-based predictive decoding including in-image prediction, and for in-image prediction, uses multiple nearest reference samples and / or multiple extended reference samples of the image directly adjacent to the prediction block to decode the prediction block of the image, each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample of the multiple reference samples, and uses a prediction mode which is one of a first set of prediction modes for predicting the prediction block using the nearest reference sample, or uses a prediction mode which is one of a second set of prediction modes for predicting the prediction block using the extended reference sample, the second set of prediction modes is a subset of the first set of prediction modes, receives mode information (m) indicating the prediction mode used to predict the prediction block, then receives parameter information (i) indicating a subset of extended reference samples used for the prediction mode, thereby indicating that the prediction mode is included in the second set of prediction modes, and if parameter information is not received, determines that the prediction mode used is not included in the second set of prediction modes, and determines the use of the nearest reference sample for prediction.

[0023] According to the embodiment, the video decoder decodes an encoded image by encoding data into video by block-based predictive decoding including in-image prediction, and for in-image prediction, uses a plurality of reference samples and a plurality of extended reference samples to decode a prediction block of an image, wherein each extended reference sample of the plurality of extended reference samples is separated from the prediction block by at least one nearest reference sample from the plurality of reference samples, and uses a prediction mode which is one of a first set of prediction modes for predicting the prediction block using the nearest reference sample, or uses a prediction mode which is one of a second set of prediction modes for predicting the prediction block using the extended reference sample, wherein the second set of prediction modes is a subset of the first set of prediction modes, and receives parameter information (i) indicating a subset of the plurality of reference samples used for the prediction mode, wherein the subset of the plurality of reference samples includes only the nearest reference sample or at least one extended reference sample, and then receives mode information (m) indicating a prediction mode used for predicting the prediction block, wherein the mode information indicates a prediction mode from a subset of modes, and the subset is limited to a set of prediction modes permitted according to the parameter information (i).

[0024] According to the embodiment, the video decoder decodes an encoded image by encoding data into video using block-based predictive decoding that includes in-image prediction, and for in-image prediction, uses multiple nearest reference samples and / or multiple extended reference samples of the image directly adjacent to the prediction block to decode the prediction block of the image, each extended reference sample of the multiple extended reference samples being separated from the prediction block by at least one nearest reference sample of multiple reference samples, and determines a first prediction of the prediction block using a first prediction mode of a set of prediction modes, the first set of prediction modes including a prediction mode that uses multiple nearest reference samples when there are no extended reference samples, and determines a second prediction of the prediction block using a second prediction mode of a second set of prediction modes, the second set of prediction modes including a subset of the first set of prediction modes being associated with multiple extended reference samples. The video decoder is configured to combine the first and second predictions to obtain a combined prediction as a prediction of the prediction block in the encoded data.

[0025] According to the embodiment, the video decoder decodes an encoded image by encoding data into video using block-based predictive decoding including in-image prediction, and in in-image prediction, uses multiple extended reference samples of the image to decode the prediction blocks of the image, each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample from a plurality of reference samples directly adjacent to the prediction block, and is configured to use multiple extended reference samples according to a predetermined set of multiple extended reference samples.

[0026] According to the embodiment, the video decoder is configured to decode an encoded image by encoding data into video by block-based predictive decoding including in-image prediction, wherein multiple predictive blocks are decoded for each image, and multiple extended reference samples of the image are used to decode the predictive blocks of the multiple predictive blocks for in-image prediction, wherein each extended reference sample of the multiple extended reference samples is separated from the predictive block by at least one nearest reference sample of multiple reference samples directly adjacent to the predictive block. The video decoder is configured to determine that the extended reference samples are at least a partial part of the adjacent predictive blocks of the multiple predictive blocks, to determine that the adjacent predictive blocks have not yet been predicted, and to receive information indicating that the extended predictive samples are associated with the predictive blocks and are placed as unavailable samples for the adjacent predictive blocks.

[0027] Further embodiments relate to methods for encoding and decoding video, as well as computer program products.

[0028] With regard to the embodiments described above in this patent application, please note that two or more of the embodiments described above, such as all embodiments, can be combined in such a way that they are simultaneously implemented in the video codec.

[0029] Furthermore, advantageous embodiments of this patent application are the subject of the dependent claims, and preferred embodiments of this patent application are described below with respect to the figures therein. [Brief explanation of the drawing]

[0030] [Figure 1] A schematic block diagram of a video encoder according to an embodiment, which includes a decoder according to the embodiment, is shown. [Figure 2] A schematic flowchart of the method for encoding a video stream according to the embodiment is shown. [Figure 3] Examples of directly adjacent (nearest) and extended reference samples used in the embodiment are shown. [Figure 4a] This embodiment shows an example of in-image predicted angles for five angles relative to a 4x2 block of prediction samples. [Figure 4b] This embodiment shows an example of in-image predicted angles for five angles relative to a 4x2 block of prediction samples. [Figure 4c] This embodiment shows an example of in-image predicted angles for five angles relative to a 4x2 block of prediction samples. [Figure 4d] This embodiment shows an example of in-image predicted angles for five angles relative to a 4x2 block of prediction samples. [Figure 4e] This embodiment shows an example of in-image predicted angles for five angles relative to a 4x2 block of prediction samples. [Figure 4f] A schematic diagram is shown to illustrate the direction of angle prediction used in the embodiment. [Figure 4g] A table illustrating an example of the dependency on the number of taps used in a filter is provided, which depends on the block size and prediction mode of the prediction block. [Figure 4h] A table illustrating an example of the dependency on the number of taps used in a filter is provided, which depends on the block size and prediction mode of the prediction block. [Figure 5a] This document illustrates an embodiment related to angle prediction using the definition of angle parameters. [Figure 5b] This document illustrates an embodiment related to angle prediction using the definition of angle parameters. [Figure 5c] This document illustrates an embodiment related to angle prediction using the definition of angle parameters. [Figure 6a] This shows the derivation of the vertical offset related to the mapping of the reference sample according to the embodiment. [Figure 6b] This shows the derivation of the vertical offset related to the mapping of the reference sample according to the embodiment. [Figure 6c]This shows the derivation of the vertical offset related to the mapping of the reference sample according to the embodiment. [Figure 7a] This shows the derivation of the horizontal offset according to the embodiment. [Figure 7b] This shows the derivation of the horizontal offset according to the embodiment. [Figure 7c] This shows the derivation of the horizontal offset according to the embodiment. [Figure 8] The upper left corner of the diagonal shows the embodiment and the use of the nearest reference sample relating to that embodiment. [Figure 9] This shows an example of projection of an extended left reference sample as a side reference adjacent to an extended top reference sample as the primary reference in the case of upper left diagonal prediction according to the embodiment. [Figure 10] This shows an example of projection of the nearest left reference sample as a side reference next to the extended upper reference sample according to the embodiment. [Figure 11] This shows an exemplary truncated single code for a specific set of reference regions relating to the embodiment. [Figure 12a] A schematic diagram of the usable block sizes according to the embodiment is shown. [Figure 12b] A schematic diagram of the usable block sizes according to the embodiment is shown. [Figure 13] This diagram shows a schematic representation of an embodiment of vertical angle prediction where the angle of the prediction block is 45 degrees. [Figure 14a] Examples of the nearest and extended reference samples required for diagonal vertical image prediction according to the embodiment are shown. [Figure 14b] Examples of the nearest and extended reference samples required for diagonal vertical image prediction according to the embodiment are shown. [Modes for carrying out the invention]

[0031] Equal or equivalent elements, or elements having equivalent or equivalent functions, are indicated in the following description by equivalent or equivalent reference numerals, even if they occur in different figures.

[0032] The following description includes several details to provide a more complete description of embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention can be carried out without these specific details. In other examples, well-known structures and apparatus are shown in block diagram form rather than in detail to avoid obscuring embodiments of the present invention. Furthermore, features of the different embodiments described below can be combined with each other unless otherwise specified.

[0033] In hybrid video coding, regions of an image sample are encoded by using in-image prediction to generate a prediction signal from available neighboring samples, i.e., reference samples. The prediction signal is subtracted from the original signal to obtain a residual signal. This residual signal, or prediction error, is further transformed, scaled, quantized, and entropy coded, as shown in Figure 1, which shows a schematic block diagram of a video encoder 1000 according to an embodiment of a hybrid video encoder that includes an in-image prediction block 1001. The video encoder 1000 is configured to receive an input video signal 1002 containing multiple images, the sequence of images forming a video. The video encoder 1000 includes a block 1004 configured to divide the signal 1002 into regions of samples, i.e., to form blocks from the input video signal 1002. A controller 1006 of the video encoder 1000 is configured to control block 1004 and a decoder 1008 which may be part of the encoder 1000. A decoder that receives and decodes the output bitstream 1012 and the generated output video signal 1014, i.e., the encoded data, can be implemented accordingly. In particular, the transform, scaling, and quantization block 1016, along with block 1018 for motion estimation of signal 1022, which is the input video signal 1002 divided into blocks by block 1004, can provide information on both quantized transform coefficients and motion information to enable entropy coding of the output bitstream 1012.

[0034] Next, the quantized and transformed coefficients are scaled and inversely transformed to generate a reconstructed residual signal before a potential in-loop filtering operation. This signal can then be added back to the prediction signal to obtain a reconstruction that is also available to the decoder. The reconstructed signal can then be used to predict subsequent samples in the same image in their coding order.

[0035] For details of in-image prediction, please refer to Figure 2. First, the reference samples used for prediction are generated in block 1042 based on the reconstructed samples. This stage also includes, for example, the replacement of adjacent samples that are unavailable at the boundaries of the image, slice, or tile. Secondly, in block 1044, the reference samples can be filtered to eliminate discontinuities in the reference signal. Thirdly, in block 1046, the prediction samples are calculated using the reference samples according to an intra-prediction mode. The prediction mode describes how the prediction signal is generated from the reference samples, for example, by averaging them in DC mode or by copying them along one prediction angle in angle prediction mode. The encoder needs to determine which intra-prediction mode to select, and the selected intra-prediction mode is communicated to the decoder via a bitstream by entropy coding. On the decoder side, the intra-prediction mode is extracted from the bitstream by entropy decoding. Fourthly, and perhaps lastly, in block 1048, the prediction samples can also be filtered to smooth the signal. In other words, Figure 2 shows a flowchart of the in-image prediction process or method. Generally, the correlation between samples in an image decreases as the distance increases. Therefore, directly adjacent samples are generally suitable as reference samples for predicting the area of ​​a sample. However, directly adjacent reference samples may represent edges or objects within a uniform region (occlusion). In these cases, the correlation between the sample to be predicted (uniform or textured region) and the directly adjacent reference sample (edge) will be low. Extended reference in-image prediction solves this problem by incorporating more distant reference samples that are not directly adjacent. While the concept of extending the nearest reference sample is known, several novel improvements to all parts of the in-image prediction process and notification are defined in embodiments of the present invention and are described below.

[0036] Extended reference image prediction allows for the generation of prediction signals for sample regions using extended reference samples. Extended reference samples are available reference samples that are not directly adjacent. The improved reference sample generation, filtering, prediction, and prediction filtering using extended reference samples according to embodiments are described in further detail below. Special cases combining prediction using extended reference samples with prediction using directly adjacent or unfiltered reference samples are described later. Then, various methods according to embodiments are described to improve the prediction mode and extended reference region notification of extended reference samples. Furthermore, embodiments for facilitating parallel coding with extended reference samples are described.

[0037] For generating reference samples, current video coding standards predict the current block using directly adjacent samples. The literature has proposed using multiple reference lines in addition to the nearest directly adjacent sample. These additional reference lines used in in-image prediction will be further referred to in detail below as extended reference samples. An improved method for replacing unavailable extended reference samples, according to embodiments, is then described.

[0038] An example showing the nearest reference sample line and three extended reference sample lines for a predicted 16x8 block is shown in Figure 3, which illustrates the directly adjacent (nearest) reference sample 1062 and extended reference samples 10641, 10642, and 10643.

[0039] The nearest reference sample 1062 and extended reference samples 10641, 10642, and 10643 are positioned adjacent to the predicted prediction block 1066, along two directions of the image, namely direction x and direction y, which is perpendicular to direction x. Along direction x, the prediction block includes extended W with samples ranging from 0 to W-1. Along direction y, the prediction block includes extended H with samples ranging from 0 to H-1.

[0040] A reference region with index i can indicate the distance between each reference sample, i.e., the nearest reference sample with index i=0, i.e., directly adjacent reference samples, where the extended reference sample is separated from the prediction block 1066 by at least the nearest reference sample 1062. For example, reference region index i can indicate the extension of the distance between the prediction block 1066 and each reference sample 1062 or 1064. As an example, increasing the parameter x along direction x can be called moving to the right, and conversely, decreasing x can be called moving to the left.

[0041] Alternatively, or furthermore, decreasing index i along the negative direction y can be referred to as moving upward or towards the top of the image, and increasing parameter y can be indicated as moving downward or towards the bottom of the image. Terms such as up, down, left, and right are used to simplify the understanding of the invention. According to other embodiments, such terms can be modified, altered, or replaced in any other direction without limiting the scope of this embodiment. As an example, reference samples 1062 and / or 1064 located to the left of the prediction block, i.e., having x < 0, can be referred to as left reference samples. Assuming the upper left end of prediction block 1066 has position 0,0, y Reference samples 1062 and / or 1064, positioned to have <0, can be called upper reference samples. The identified reference samples, left reference samples, and upper reference samples can be called corner reference samples. Thus, reference samples extending beyond extension W along the x-direction can be called right reference samples, and reference samples extending beyond extension H of prediction block 1066 can be called lower reference samples.

[0042] To indicate the reference samples used for prediction, each line of the reference samples is associated with a reference region index i. The nearest reference sample is given index i=0, the next line of the extended reference sample is i=1, and so on. Using the notation in Figure 3, the reference samples above are in the range x from 0 to M.

number

number

number

number

[0043] A video encoder according to an embodiment such as video encoder 1000 can be configured to encode images of a video into encoded data by block-based predictive coding, the block-based predictive coding includes in-image prediction. In in-image prediction, the video encoder can use multiple extended reference samples of the image to encode prediction blocks of the image, each extended reference sample of the multiple extended reference samples being separated from the prediction block by at least one nearest reference sample from the multiple reference samples that is directly adjacent to the prediction block. The video encoder can sequentially determine the availability or unavailability of each of the multiple nearest reference samples and can replace the nearest reference sample that is determined to be unavailable with a replacement sample. The video encoder can use replacement samples for in-image prediction.

[0044] To determine availability or unavailability, the video encoder sequentially checks the samples according to the sequence and can determine a replacement sample as a copy of the last extended reference sample determined to be available in the sequence, and / or the next extended reference sample determined to be available in the sequence.

[0045] The video encoder can further determine sequential availability or unavailability according to the sequence, and determine replacement samples based on combinations of extended reference samples that are placed in the sequence before a reference sample is determined to be available and an extended reference sample is determined to be unavailable, and after a reference sample is determined to be unavailable and is placed in sequence.

[0046] Alternatively, the video encoder may be configured to use multiple nearest reference samples and multiple extended reference samples of the image directly adjacent to the prediction block to encode the prediction block of the image for in-image prediction, wherein each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample of the multiple reference samples, and the availability or unavailability of each of the multiple extended reference samples can be determined. The video encoder may notify the use of multiple extended reference samples if some of the available extended reference samples of the multiple extended reference samples are above a predetermined threshold, and may skip the notification of use of multiple extended reference samples if some of the available extended reference samples of the multiple extended reference samples are below the predetermined threshold.

[0047] Therefore, each decoder, such as the video decoder 1008 or a video decoder for regenerating a video stream, can be configured to decode an image encoded with encoded data into video by block-based predictive decoding including in-image prediction, using multiple extended reference samples of the image to encode prediction blocks of the image, each extended reference sample of the multiple extended reference samples being separated from the prediction block by at least one nearest reference sample from a plurality of reference samples directly adjacent to the prediction block, and sequentially determining the availability or unavailability of each of the multiple extended reference samples. The video decoder can replace extended reference samples determined to be unavailable by replacement samples and use replacement samples for in-image prediction.

[0048] The video decoder may also be configured to determine availability or unavailability sequentially according to the sequence, determine a replacement sample as a copy of the last extended reference sample determined to be available in the sequence, and / or determine a replacement sample as a copy of the next extended reference sample determined to be available in the sequence.

[0049] Furthermore, the video decoder can be configured to determine availability or unavailability in order according to the sequence, and to determine replacement samples based on combinations of extended reference samples that are placed in the sequence before a reference sample is determined to be available and before an extended reference sample is determined to be available, and after a reference sample is determined to be unavailable and is then placed in order.

[0050] The video decoder may be configured to use, alternatively or further, multiple nearest reference samples and multiple extended reference samples of the image directly adjacent to the prediction block in order to decode the prediction block of the image for in-image prediction, wherein each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample of the multiple reference samples, and the availability or unavailability of each of the multiple extended reference samples is determined, and if a portion of the available extended reference samples of the multiple extended reference samples is above a predetermined threshold, it notifies received information indicating the use of the multiple extended reference samples, and if there is no information, i.e., if a portion of the available extended reference samples of the multiple extended reference samples is below the predetermined threshold, it may be configured to skip the use of the multiple extended reference samples.

[0051] If adjacent reference samples are unavailable, according to the embodiment, extended reference sample substitution can be performed. For example, if an unavailable sample is unavailable, the nearest available adjacent sample, the nearest two adjacent samples, or, if there are no available adjacent samples, for example, two bitdepth-1 These can be replaced by predetermined values, such as [values ​​omitted]. For example, reference samples are unavailable if they are outside the boundaries of an image, slice, or tile, or if constrained internal prediction is used which does not allow the use of samples from inter-image prediction regions as references for in-image prediction regions.

[0052] For example, if the predicted current block is on the left image boundary, the left and upper-left corner reference samples are unavailable. In this case, the left and upper-left corner reference samples are replaced by the first available upper reference sample. This first available upper reference sample is the first sample, i.e.

number

number

[0053] Using constrained intra-prediction, because it is encoded using inter-image prediction, one or more adjacent blocks may be unavailable. For example, the left H sample.

number

number

[0054] In one embodiment, the availability check process for each reference sample is performed sequentially, for example, from the bottom left to the top right, or vice versa, and the first unavailable sample along this direction is replaced with the last available sample. If there are no previously available samples, the unavailable sample is replaced with the next available sample. In an embodiment starting from the bottom right, the W bottom right sample...

number

number

number

[0055] If it is determined that most of the extended reference region samples are unavailable, using the extended reference samples offers no advantage compared to using the nearest reference samples. Therefore, a notification of the reference region index can be saved, and that notification can be limited to blocks where at least half of the extended reference samples are available.

[0056] A video encoder according to an embodiment such as the video encoder 1000 can be configured to encode images of a video into encoded data by block-based predictive coding including in-image prediction, to use multiple extended reference samples of an image to encode prediction blocks of an image in in-image prediction, to separate each extended reference sample of the multiple extended reference samples from the prediction block by at least one nearest reference sample from a plurality of reference samples directly adjacent to the prediction block, to filter at least a subset of the multiple extended reference samples using a bilateral filter to obtain multiple filtered extended reference samples, and to use multiple filtered extended reference samples for in-image prediction.

[0057] The video encoder according to this embodiment can be configured to obtain multiple combined reference values ​​by combining multiple filtered extended reference samples with multiple unfiltered extended reference samples, and the video encoder is configured to use multiple combined reference values ​​of in-image predictions.

[0058] Alternatively, the video encoder can be configured to filter multiple extended reference samples using one of three-tap, five-tap, or seven-tap filters.

[0059] The video encoder can further be configured to choose to predict a prediction block using an angle prediction mode, and the 3-tap, 5-tap, and 7-tap filters are configured as bilateral filters, and the video encoder is configured to choose to use one of the 3-tap, 5-tap, and 7-tap filters based on the angle used for angle prediction, where the angle is positioned between the horizontal and vertical directions of the angle prediction mode, and / or the video decoder is configured to choose to use one of the 3-tap, 5-tap, and 7-tap filters based on the block size of the prediction block. As shown in Figure 4f, the angle ε can be used to predict the prediction block 1066 with respect to the horizontal boundary 1072 and / or vertical boundary 1074 of the prediction block 1066, and can represent the angle of the direction of the angle prediction measured toward the diagonal 1076 between the horizontal and vertical directions. That is, the angle of the angle prediction is at most 45°. A larger angle ε increases the number of taps available in the filter. Alternatively, or further, the block size can define the basis or dependency for selecting the filter.

[0060] The corresponding video decoder can be configured to decode an image encoded with encoded data into video by block-based predictive decoding including in-image prediction, using multiple extended reference samples of the image to decode the prediction block of the image, each extended reference sample of the multiple extended reference samples being separated from the prediction block by at least one nearest reference sample from multiple reference samples directly adjacent to the prediction block, filtering at least a subset of the multiple extended reference samples using a bilateral filter to obtain multiple filtered extended reference samples, and using multiple filtered extended reference samples for in-image prediction.

[0061] The video decoder can also be configured to combine multiple filtered extended reference samples with multiple unfiltered extended reference samples to obtain multiple combined reference values, and the video decoder can be configured to use the multiple combined reference values ​​for in-image prediction.

[0062] Alternatively, the video decoder may be configured to filter multiple extended reference samples using one of three-tap, five-tap, and seven-tap filters. As described for the encoder, the three-tap, five-tap, and seven-tap filters are configured as bilateral filters, and the video decoder is configured to predict a prediction block using an angle prediction mode and to select to use one of three-tap, five-tap, and seven-tap filters based on the angle used for angle prediction, where the angle is located between the horizontal and vertical directions of the angle prediction mode, and / or the video decoder is configured to select to use one of three-tap, five-tap, and seven-tap filters based on the block size of the prediction block.

[0063] For example, instead of a bilateral filter, a 3-tap FIR filter can be used. This allows filtering only the nearest reference sample (even if a bilateral filter is not used) while leaving extended reference samples unfiltered.

[0064] When the sample area is large, discontinuities can occur in the reference sample, potentially distorting the prediction. A state-of-the-art solution to this is to apply a linear smoothing filter to the reference sample. For example, strong smoothing can be applied to discontinuities that can be detected by comparing them to a given threshold. This typically involves generating reference samples by interpolating between corner reference samples.

[0065] However, linear smoothing filters can also remove edge structures that need to be preserved. By applying a bilateral filter to an extended reference sample according to an embodiment of reference sample filtering, undesirable smoothing of sharp edges can be prevented. Since bilateral filtering is more efficient for large blocks and intra-predicted angles that deviate from the exact horizontal and exact vertical directions, the decision of whether to apply the filter and the length of the filter can depend on the block size and / or prediction mode. One example design can incorporate dependencies as shown in Figure 4g, which shows a dependency on block sizes smaller than 64×64 and larger with respect to W×H, and Figure 4h shows a different dependency on block sizes smaller than 64×64 and larger with respect to W×H. In the embodiment of Figure 4g, the in-image prediction mode can be, for example, one of several different angles in the planar mode, DC mode, nearly horizontal mode, nearly vertical mode, or angle mode. In the embodiment of Figure 4h, angles identified as farther horizontal and farther vertical can be further selected, for example, when the angle ε shown in Figure 4f is larger compared to nearly horizontal or nearly vertical. As can be seen, a larger block size can result in more taps to facilitate filtering of larger amounts of data, and furthermore, an increase in ε can also result in an increase in taps. According to Figure 4h, an embodiment can apply a small 3-tap filter to nearly horizontal and nearly vertical modes, and increase the filter length as the distance from the horizontal and vertical directions increases. Although shown as depending on both the prediction mode and block size, the selection of the filter or at least the number of taps can instead depend on either one of both and / or additional parameters only.

[0066] In another embodiment of reference sample filtering, in-image predictions using filtered reference samples can be combined with unfiltered reference samples using location-dependent weighting, as described in relation to the combination of location-dependent predictions. In this case, the reference samples for predictions using filtered reference samples can use a different reference sample filtering than in the case of unconnected predictions. For example, the filtering can be selected from sets of 3-tap, 5-tap, and 7-tap filters.

[0067] A video encoder according to an embodiment such as video encoder 1000 encodes images of a video into encoded data by block-based predictive coding including in-image prediction, and in in-image prediction, uses a plurality of extended reference samples of an image to encode prediction blocks of an image, each extended reference sample of the plurality of extended reference samples is separated from the prediction block by at least one nearest reference sample from a plurality of reference samples directly adjacent to the prediction block, the plurality of nearest reference samples are arranged along a first image direction of the prediction block and along a second image direction of the prediction block, and at least a portion of the nearest reference samples arranged along the second direction are mapped to extended reference samples arranged along the first direction such that the mapped reference samples extend beyond the extension of the prediction block along the first image direction, and the mapped extended reference samples can be used for prediction.

[0068] The video encoder can also be configured to map the nearest portion of the reference sample according to the prediction mode used to predict the predicted block. The video encoder can also be configured to map the nearest portion of the reference sample according to the direction used in the prediction mode for predicting the predicted block.

[0069] The corresponding video decoder decodes an image encoded with encoded data into video by block-based predictive decoding, including in-image prediction, and in in-image prediction, uses multiple extended reference samples of the image to decode the prediction block of the image, each extended reference sample of the multiple extended reference samples being separated from the prediction block by at least one nearest reference sample from multiple reference samples directly adjacent to the prediction block, and the multiple nearest reference samples are arranged along a first image direction of the prediction block and along a second image direction of the prediction block. The system can be configured to map at least a portion of the nearest reference samples located along a second direction to extended reference samples located along a first direction, so that the mapped reference samples extend beyond the extension of the prediction block along a first image direction, and to use the mapped extended reference samples for prediction.

[0070] The video decoder can be configured to map the nearest reference sample portion according to the prediction mode used to predict the predicted block. The video decoder can be configured to map the nearest reference sample portion according to the direction used in the prediction mode for predicting the predicted block.

[0071] In principle, all in-image predictions that use directly adjacent reference samples can be adapted to use extended reference samples. The following three prediction methods are employed in the literature.

[0072] ·plane ·DC ·angle In the following sections, each prediction will be described in detail for an extended reference sample in order to illustrate embodiments of the present invention.

[0073] Planar prediction is a bilinear interpolation of W×H prediction samples from the boundary shown in Figure 3. Since the right and bottom boundaries have not yet been reconstructed, the right boundary sample is the sample from the upper right corner.

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0074] DC prediction calculates the average of the reference sample value DC and all predicted samples within the W×H block.

number

number

number

number

number

[0075] According to the embodiment, a finer predictive angle granularity can be used so that the direction can point to a non-integer sample position. These can be generated by interpolation using the nearest integer sample position. Interpolation can be performed using a simple bilinear filter or a more advanced filter, such as a cubic or Gaussian 4-tap filter.

[0076] For predicted angles between horizontal and vertical lines, such as the top-left angle of a diagonal and all angles between it and the diagonal, predictions are generated using reference samples from the left and top. To facilitate such generation, reference sample projection can be used. For all vertex angles between the top-left diagonal and the vertical, the top reference sample can be displayed as the primary reference, and the left-side sample as the secondary reference. For all horizontal angles between the top-left horizontal line and the diagonal, the left-side reference sample can be displayed as the primary reference, and the top reference sample as the secondary reference. To simplify calculations, i.e., to avoid switching between primary and secondary reference sample calculations, the secondary reference sample is projected along the predicted angle to extend the line of the primary reference sample. (Top primary reference to the left)

number

number

[0077] Figures 5a to 5c illustrate embodiments relating to angle prediction using a definition of angle parameter A given with a 1 / 32 sample accuracy for the top-left diagonal predicted angle and two other vertical angles. The following embodiments assume a 1 / 32 sample accuracy and a vertical predicted angle between the vertical and the top-left diagonal, but any other arbitrary value can be implemented. The 33 predicted angle ranges are angle parameters ranging from 0 (vertical) to 32 (top-left diagonal), as shown in Figures 5a to 5c.

number

[0078] Angle parameters

number

number

number

number

number

number

number

[0079] Figures 7a to 7c show the height of the prediction block for the upper left horizontal prediction angle at A=32, 17, 1.

number

number

[0080] In this embodiment, the nearest integer reference sample is used instead of the subsample position by rounding the vertical offset. This simplifies the calculation. However, it is also possible to project the reference sample on the interpolated subsample side.

[0081] Figure 8 shows the diagonal upper left corner.

number

number

number

[0082] Figure 8 shows an example of projection according to an embodiment, where the nearest upper reference sample is used as the primary reference and the nearest left reference sample is used as the side reference, in the case of the upper left horizontal direction with angle parameter A=17.

[0083] For an extended reference sample of reference region index i, the projection can be adapted as follows:

number

number

number

number

number

[0084] Figure 9 shows an example of projection according to an embodiment, where an extended upper reference sample is used as the primary reference and an extended left reference sample is used as the secondary reference, in the case of upper left diagonal prediction with angle parameter A=17.

[0085] When using reference samples from the left and above, the extended reference sample allows for the combination of the nearest reference sample and the extended reference sample. This embodiment of the simple approach from Figure 9 allows for the use of the extended reference sample along the principal direction, i.e., the principal reference, where longer distances of the reference sample may be beneficial in the case of occlusion or edges on the nearest reference sample. However, in the case of secondary references, the correlation between the predicted sample and the nearest reference sample may be higher than the correlation between the extended reference sample and the predicted sample.

[0086] For example, the extended main reference sample

number

number

number

[0087] Figure 10 shows an example of projection of the nearest left-side reference sample as a secondary reference adjacent to an extended upper-side reference sample as the primary reference, in the case of upper-left diagonal prediction using angle parameter A=17. According to the embodiment of Figure 10, the nearest reference sample is used as the source for generating the extended reference sample; that is, the nearest reference sample is mapped to the extended reference sample. Alternatively, an extended reference sample can be mapped to another extended reference sample. See Figure 9.

[0088] A video encoder according to an embodiment such as the video encoder 1000 encodes images of a video into encoded data by block-based predictive coding including in-image prediction, and can be configured to use a plurality of nearest reference samples and a plurality of extended reference samples of the image directly adjacent to the prediction block in order to encode the prediction block of the image for in-image prediction, wherein each extended reference sample of the plurality of extended reference samples is separated from the prediction block by at least one nearest reference sample of the plurality of reference samples. The video decoder is configured to perform boundary filtering in a mode in which extended samples are not used, and in the case of using extended samples, to not use boundary filtering, or to perform boundary filtering on at least a subset of the plurality of nearest reference samples and not use boundary filtering on extended samples.

[0089] The corresponding video decoder decodes the encoded image by encoding the data into video using block-based predictive decoding, which includes in-image prediction. For in-image prediction, it is configured to use multiple nearest reference samples and multiple extended reference samples of the image directly adjacent to the prediction block to decode the prediction block of the image, with each extended reference sample being separated from the prediction block by at least one nearest reference sample from multiple reference samples. The video decoder performs boundary filtering in a mode where extended samples are not used, and when extended samples are used, it is configured to either not use boundary filtering, or to perform boundary filtering on at least a subset of the multiple nearest reference samples and not use boundary filtering on the extended samples.

[0090] Because extended reference samples are not directly adjacent to the prediction samples, discontinuities at certain block boundaries may not be as severe as they would be if the nearest reference sample were used. Therefore, it is beneficial to perform prediction filtering as follows: • If extended reference samples are used, do not perform boundary filtering operations, or • Modify boundary smoothing by using the nearest reference sample instead of the extended reference sample.

[0091] Using extended prediction, it is possible to combine predictions from the nearest reference sample and extended reference samples to obtain a combined prediction. The literature describes a fixed combination of predictions using the nearest reference sample and predictions using extended reference samples with predetermined weights. In this case, both predictions use the same prediction mode for all reference regions, and when the mode is notified, the prediction combination is also notified. On the other hand, this reduces the notification overhead for indicating reference sample regions, but it also loses the flexibility to combine two different prediction modes with two different reference sample regions. Possible combinations that increase flexibility are described in detail below.

[0092] A video encoder according to an embodiment such as video encoder 1000 encodes images of a video into encoded data by block-based predictive coding including in-image prediction. In the in-image prediction, in order to encode a prediction block of an image, multiple nearest reference samples and multiple extended reference samples of the image directly adjacent to the prediction block are determined. Each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample among the multiple reference samples. The prediction of the prediction block is determined using the extended reference samples, the extended reference samples are filtered, and multiple filtered extended reference samples are obtained. The video encoder is configured to combine the prediction and the filtered extended reference samples to obtain a combined prediction of the prediction block.

[0093] A video encoder can be configured to combine predicted samples and extended reference samples that are positioned on the main or secondary diagonal of the sample with respect to the prediction block.

[0094] The video encoder can be configured to combine predicted samples and extended reference samples based on the following decision rules:

number

number

[0095] The normalization coefficient can be determined based on a determination rule.

Number

[0096] The video encoder can be configured to use a combination of the extended corner reference sample of the prediction block and the extended reference samples (r(-1 - i, -1 - i)) arranged in the corner region of the reference samples.

[0097] The video encoder can be configured to obtain a combined prediction based on the following determination rule:

Number

Number

number

[0098] The video encoder can be configured to obtain a predicted p(x,y) based on predictions within the image.

[0099] A video encoder can be configured to use only planar prediction as the in-image prediction method.

[0100] A video encoder can be configured to determine a set of parameters that identify the combination of predictions and filtered extended reference samples for each encoded video block. Alternatively, a video encoder can be configured to use a lookup table containing sets of prediction blocks of different block sizes to determine the set of parameters that identify the combination of predictions and filtered extended reference samples.

[0101] The corresponding video decoder can be configured to decode an encoded image by encoding data into video using block-based predictive decoding, which includes in-image prediction. In the in-image prediction, to decode the prediction block of the image, it may determine multiple nearest reference samples and multiple extended reference samples of the image directly adjacent to the prediction block, each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample of the multiple reference samples, the prediction of the prediction block is determined using the extended reference samples, the extended reference samples are filtered to obtain multiple filtered extended reference samples, and the prediction and the filtered extended reference samples are combined to obtain a combined prediction of the prediction block.

[0102] The video decoder can be configured to combine predicted samples and extended reference samples that are positioned on the main or secondary diagonal of the sample with respect to the prediction block.

[0103] The video encoder can be configured to combine predicted samples and extended reference samples based on the following decision rules:

number

number

[0104] The normalization coefficient can be determined based on the criteria.

number

[0105] The video decoder can be configured to use a combination of an extended corner reference sample of the prediction block and an extended reference sample located in the corner region of the reference sample (r(-1-i,-1-i)).

[0106] The video decoder can be configured to obtain predictions combined based on the following judgment rules:

number

number

number

[0107] The video decoder can be configured to obtain a prediction p(x,y) based on in-image predictions. For example, the video decoder can use only planar predictions as the in-image predictions.

[0108] The video encoder can be configured to determine a set of parameters that identify a combination of predictions and filtered extended reference samples for each decoded video block.

[0109] The video decoder can be configured to determine a set of parameters that identify combinations of predictions and filtered extended reference samples, using a lookup table containing sets of prediction blocks with different block sizes.

[0110] When filtering reference samples, a combination of position-dependent predictions can be obtained by combining predictions using the filtered samples with unfiltered reference samples, based on the position of each sample. According to an exemplary embodiment,

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0111] This known combination of nearest reference samples can be improved by embodiment by using extended reference samples for filtered and unfiltered reference samples. Below are the combinations of predictions using extended reference samples according to embodiment, where the nearest corner sample r(-1,-1) is extended corner sample

number

Number

Number

Number

Number

Number

Number

[0112] Alternatively, or furthermore, different extended prediction modes can be combined to combine the prediction with the sample.

[0113] A video encoder according to an embodiment such as video encoder 1000 encodes images of a video into encoded data by block-based predictive coding including in-image prediction, and for in-image prediction, uses multiple nearest reference samples and / or multiple extended reference samples of the image directly adjacent to the prediction block to encode the prediction block of the image, each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample of the multiple reference samples, and determines a first prediction of the prediction block using a first prediction mode of a set of prediction modes, the first set of prediction modes includes a prediction mode that uses multiple nearest reference samples when there are no extended reference samples, and determines a second prediction of the prediction block using a second prediction mode of a second set of prediction modes, the second set of prediction modes including a subset of the first set of prediction modes is associated with the multiple extended reference samples. The video encoder is configured to encode a first prediction (p0(x,y)) and a second prediction (p i (x, y)) and weighted (w0;w i ) can be combined and configured to obtain a combined prediction (p(x, y)) as the prediction of the prediction block of the encoded data.

[0114] The video encoder can be configured to use a first prediction and a second prediction according to a predetermined combination which is part of a possible combination of a valid first prediction mode and a valid second prediction mode.

[0115] A video encoder can be configured to indicate either a first or second prediction mode without indicating other prediction modes. For example, the first mode can be derived from parameter i based on additional implicit information, such as a specific prediction mode that can only be used in relation to a particular index i or index m.

[0116] As an example of such implicit information, the video encoder may be configured to exclusively use the planar prediction mode as one of the first and second prediction modes.

[0117] The video encoder can be configured to fit a first weight applied to the first prediction of the combined prediction and a second weight applied to the second prediction of the combined prediction based on the block size of the prediction block, and / or to fit the first weight based on the first prediction mode or the second weight based on the second prediction mode.

[0118] The video encoder can be configured to fit a first weight applied to the first prediction of the combined prediction and a second weight applied to the second prediction of the combined prediction based on the position and / or distance within the prediction block.

[0119] The video encoder can be configured to apply a first weight and a second weight based on the following decision rules.

number

[0120] The corresponding video decoder decodes the encoded image by encoding the data into video using block-based predictive decoding, which includes in-image prediction, and for in-image prediction, it uses multiple nearest reference samples and / or multiple extended reference samples of the image directly adjacent to the prediction block to decode the prediction block of the image, and each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample of the multiple reference samples. The first prediction of a prediction block can be determined using a first prediction mode from a set of prediction modes, where the first set of prediction modes includes a prediction mode that uses multiple nearest reference samples when there are no extended reference samples, and the second prediction of a prediction block can be determined using a second prediction mode from a second set of prediction modes, where the second set of prediction modes includes a subset of the prediction modes from the first set, where the subset is associated with multiple extended reference samples. The video decoder can then determine the first prediction (p0(x,y)) and the second prediction (p i (x, y)) and weighted (w0;w i ) can be combined and configured to obtain a combined prediction (p(x, y)) as the prediction of the prediction block of the encoded data.

[0121] The video decoder can be configured to use a first prediction and a second prediction according to a predetermined combination, which is part of a possible combination of a valid first prediction mode and a valid second prediction mode. This makes it possible to obtain low-load notifications.

[0122] The video decoder may be configured to receive a signal indicating a second prediction mode without receiving a signal indicating a first prediction mode, and to derive the first prediction mode from the second prediction mode or parameter information (i).

[0123] The video decoder can be configured to exclusively use the planar prediction mode as one of the first and second prediction modes.

[0124] The video decoder can be configured to fit a first weight applied to the first prediction of the combined prediction and a second weight applied to the second prediction of the combined prediction based on the block size of the prediction block, and / or to fit the first weight based on the first prediction mode or the second weight based on the second prediction mode.

[0125] The video decoder can be configured to adapt a first weight applied to a first prediction of the combined prediction and a second weight applied to a second prediction of the combined prediction based on a position and / or distance within a prediction block.

[0126] The video decoder can be configured to adapt the first weight and the second weight based on the following decision rules.

Number

[0127] When using different reference sample regions, as follows, prediction using the nearest reference sample mode

Number

Number

Number

[0128] Embodiments to alleviate this strict restriction define allowing only specific combinations of modes. This requires notification of a second mode, but limits the number of modes to be notified compared to the first mode. Details of predictive mode notification are provided in relation to mode and reference notification. One promising combination according to embodiments is to use only planes as the second mode, and as a result additional notification of intra-modes becomes obsolete. For example, any intra-mode as the first part of the weighted sum

number

[0129] To fit the weights to the predicted size and mode,

number

number

number

[0130] A video encoder according to an embodiment such as video encoder 1000 encodes images of a video into encoded data by block-based predictive coding including in-image prediction, and for in-image prediction, uses a plurality of nearest reference samples and / or a plurality of extended reference samples of the image directly adjacent to the prediction block to encode the prediction block of the image, wherein each extended reference sample of the plurality of extended reference samples is separated from the prediction block by at least one nearest reference sample of the plurality of reference samples, and can be configured to use a prediction mode which is one of a first set of prediction modes for predicting the prediction block using the nearest reference sample, or a prediction mode which is one of a second set of prediction modes for predicting the prediction block using the extended reference sample, wherein the second set of prediction modes is a subset of the first set of prediction modes, the second subset can be determined by the encoder, and the subset may also include a match between both sets. The video encoder can be configured to notify mode information (m) indicating the prediction mode used to predict the prediction block, then, if the prediction mode is included in a second set of prediction modes, parameter information (i) indicating a subset of extended reference samples used for the prediction mode, and to skip notifying parameter information if the prediction mode used is not included in the second set of prediction modes, thereby enabling the conclusion that parameter i has a predetermined value such as 0.

[0131] A video encoder can be configured to skip notification of parameter information if the mode information indicates DC mode or planar mode.

[0132] The corresponding video decoder decodes an image encoded with encoded data into video by block-based predictive coding including in-image prediction, and for in-image prediction, to decode the prediction block of the image, it uses multiple nearest reference samples and / or multiple extended reference samples of the image directly adjacent to the prediction block, each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample of the multiple reference samples, and it uses a prediction mode that is one of a first set of prediction modes for predicting the prediction block using the nearest reference sample, or a prediction mode that is one of a second set of prediction modes for predicting the prediction block using the extended reference sample, and the second set of prediction modes can be configured to be a subset of the first set of prediction modes. The second set of prediction modes is a subset of the first set of prediction modes and can be determined and / or notified by the encoder. Being a subset can include matching between both sets. The video encoder can be configured to receive mode information (m) indicating the prediction mode used to predict a prediction block, then parameter information (i) indicating a subset of extended reference samples used for the prediction mode, thereby indicating that the prediction mode is included in a second set of prediction modes, and if parameter information is not received, to determine that the prediction mode used is not included in a second set of prediction modes, and to determine the use of the closest reference sample for prediction.

[0133] A video decoder can be configured to determine mode information as indicating the use of DC mode or planar mode when parameter information is not received.

[0134] If the set of allowed intra-prediction modes for an extended reference sample is limited, i.e., limited to a subset of the allowed intra-prediction modes for the nearest reference sample, for example, if the subset is called the limited prediction modes for the extended reference sample, then there are two ways to communicate the mode m and index i: 1. The signal mode m before index i can depend on mode m as follows: a. If mode m is not in the set of allowed modes for the extended reference sample, the notification for index i is skipped, and the prediction mode m is applied to the nearest reference sample (i=0). b. Otherwise, i is notified, and the prediction mode m is applied to the reference sample indicated by i.

[0135] 2. The signal prior to mode m, index i, can depend on index i as follows: If ai indicates that it will use extended reference samples (i>0), the set of modes m that can be indicated is the same as the restricted set.

[0136] b. Otherwise (i=0), the set of permitted modes m that is notified is equal to the set of unrestricted modes.

[0137] For example, if mode m is notified using MPM coding with an index to the most likely mode (MPM) list, modes that are not in the set of permitted modes will not be included in the MPM list.

[0138] That is, according to a second option which can be implemented alternatively or additionally, a video encoder according to an embodiment such as video encoder 1000 encodes images of a video into encoded data by block-based predictive coding including in-image prediction, and for in-image prediction, uses a plurality of reference samples and a plurality of extended reference samples to encode prediction blocks of an image, each extended reference sample of the plurality of extended reference samples is separated from the prediction block by at least one nearest reference sample of the plurality of reference samples, and a first set of prediction modes for predicting the prediction block using the nearest reference sample The system is configured to use one prediction mode, or one of a second set of prediction modes for predicting a prediction block using extended reference samples, wherein the second set of prediction modes is a subset of the first set of prediction modes, and it notifies parameter information (i) indicating a subset of multiple reference samples used for the prediction mode, wherein the subset of multiple reference samples includes only the nearest reference sample or extended reference samples, and then notifies mode information (m) indicating a prediction mode used for predicting a prediction block, wherein the mode information indicates a prediction mode from a subset of modes, and the subset is limited to a set of prediction modes permitted according to parameter information (i).

[0139] For both options, the video decoder is adapted to use the nearest reference sample, as well as extended reference samples of modes included in the second set of prediction modes.

[0140] Furthermore, the video encoder can be adapted so that a first set of prediction modes describes prediction modes that can be used with the nearest reference sample, and a second set of prediction modes describes prediction modes of the first set of prediction modes that can also be used with extended reference samples.

[0141] The range of values ​​for the parameter information, i.e., the domain of values ​​that can be represented by the area index i, can cover the use of only the nearest reference value and the use of different subsets of extended reference values. As illustrated in relation to Figure 11, i can represent the use of only the nearest reference value (i=0), a specific set of extended reference values, or a combination of sets of extended reference values ​​(e.g., lines and / or columns or distances).

[0142] According to the embodiment, different portions of the extended reference sample include different distances to the prediction block.

[0143] The video encoder can be configured to set parameter information to one of a predetermined number of values, which indicate the number and distance of reference samples used in the prediction mode.

[0144] The video encoder can be configured to determine a first set of predictive modes and / or a second set of predictive modes based on the most likely mode coding.

[0145] The corresponding decoder of the second option decodes an encoded image by encoding data into video using block-based predictive decoding including in-image prediction, and for in-image prediction, uses multiple reference samples and multiple extended reference samples to decode the prediction block of the image, each extended reference sample of the multiple extended reference samples being separated from the prediction block by at least one nearest reference sample from the multiple reference samples, and uses a prediction mode which is one of a first set of prediction modes for predicting the prediction block using the nearest reference sample, or uses a prediction mode which is one of a second set of prediction modes for predicting the prediction block using the extended reference sample, the second set of prediction modes being a subset of the first set of prediction modes, and receives parameter information (i) indicating a subset of multiple reference samples used for the prediction mode, the subset of multiple reference samples including only the nearest reference sample or at least one extended reference sample, and then receives mode information (m) indicating a prediction mode used for predicting the prediction block, the mode information indicating a prediction mode from a subset of modes, and the subset being limited to a set of prediction modes permitted according to parameter information (i).

[0146] Decoders relating to the first and / or second options can be adapted, for example, by combining predictions and samples and / or prediction combinations, so that in addition to the nearest reference sample, extended reference samples of modes included in a second set of prediction modes are used.

[0147] A first set of prediction modes can describe prediction modes permitted for use with the nearest reference sample, and a second set of prediction modes can describe prediction modes from the first set of prediction modes that are also permitted for use with extended reference samples.

[0148] As explained regarding encoders, the range of parameter information values ​​covers both using only the nearest reference value and using different subsets of extended reference values.

[0149] Different parts of the extended reference sample can include different distances to the predicted block.

[0150] The video decoder can be configured to set parameter information to one of a predetermined number of values, which indicate the number and distance of reference samples used in the prediction mode.

[0151] A video decoder can be configured to determine a first set of predicted modes and / or a second set of predicted modes based on the most likely mode (MPM) coding, that is, it can generate lists for each, and these lists can be adapted to include only the modes for which each sample is permitted to be used.

[0152] Alternatively, or further, embodiments relating to combining predictions from the nearest reference sample and an extended reference sample may be implemented.

[0153] A video encoder according to an embodiment such as video encoder 1000 may be configured to encode images of a video into encoded data by block-based predictive coding including in-image prediction, to encode prediction blocks of an image for in-image prediction, to use multiple nearest reference samples and / or multiple extended reference samples of the image directly adjacent to the prediction block, each extended reference sample of the multiple extended reference samples being separated from the prediction block by at least one nearest reference sample of the multiple reference samples, to determine a first prediction of the prediction block using a first prediction mode of a set of prediction modes, the first set of prediction modes including a prediction mode that uses multiple nearest reference samples when there are no extended reference samples, to determine a second prediction of the prediction block using a second prediction mode of a second set of prediction modes, the second set of prediction modes including a subset of the first set of prediction modes associated with multiple extended reference samples, and to combine the first and second predictions to obtain a combined prediction as a prediction of the prediction block in the encoded data.

[0154] A prediction block may be a first prediction block, and the video encoder may be configured to predict a second prediction block of video using multiple nearest reference samples associated with the second prediction block if there are no multiple extended reference samples associated with the second prediction block. The video encoder may be configured to signal combination information, such as a bipred flag, indicating whether the prediction in the encoded data is based on a combination of predictions or on a prediction that uses multiple extended reference samples if there are no multiple nearest reference samples.

[0155] The video encoder can be configured to use a first prediction mode as a predetermined prediction mode.

[0156] The video encoder may be configured to select a first prediction mode that is the same as a second prediction mode and uses the closest reference sample when no extended reference sample is available, or to use a pre-configured prediction mode such as a planar prediction mode.

[0157] The corresponding video decoder can be configured to decode an encoded image by encoding data into video using block-based predictive decoding that includes in-image prediction, and for in-image prediction, to decode the prediction block of the image, it may use multiple nearest reference samples and / or multiple extended reference samples of the image directly adjacent to the prediction block, each extended reference sample of the multiple extended reference samples being separated from the prediction block by at least one nearest reference sample of multiple reference samples, and to determine a first prediction of the prediction block using a first prediction mode of a set of prediction modes, the first set of prediction modes including a prediction mode that uses multiple nearest reference samples when there are no extended reference samples, and to determine a second prediction of the prediction block using a second prediction mode of a second set of prediction modes, the second set of prediction modes including a subset of the prediction modes of the first set, the subset being associated with multiple extended reference samples. The video decoder can be configured to combine the first and second predictions to obtain the combined prediction as the prediction of the prediction block in the encoded data.

[0158] The prediction block may be a first prediction block, and the video decoder may be configured to predict a second prediction block of the video using multiple nearest reference samples associated with the second prediction block if there are no multiple extended reference samples associated with the second prediction block. The video decoder may also be configured to receive concatenation information indicating that the prediction in the encoded data is based on a combination of predictions or on a prediction using multiple extended reference samples if there are no multiple nearest reference samples, and to decode the encoded data accordingly.

[0159] The video decoder can be configured to use a first prediction mode as a predetermined prediction mode.

[0160] The video decoder can be configured to select a first prediction mode that is the same as the second prediction mode and uses the closest reference sample when there is no extended reference sample, or to use the first prediction mode as a preset prediction mode such as planar mode.

[0161] If the reference region index i indicates the use of extended reference samples (i>0), the bipred flag and other information, such as the binary information or flags referred to below, are used to indicate whether the prediction using extended reference samples is combined with the prediction using the closest reference samples.

[0162] If the bipred flag indicates a combination of predictions from the nearest (i=0) and extended reference samples (i>0), then the extended reference sample m i The mode is notified before or after the reference region index i, as described here. The mode of the nearest reference sample prediction m0 is fixed, and can be set to a specific mode, such as always being plane, or to the same mode as the extended reference sample.

[0163] As described below, the video encoder according to the embodiment encodes images of a video into encoded data by block-based predictive coding including in-image prediction, and in in-image prediction, uses a plurality of extended reference samples of an image to encode prediction blocks of an image, each extended reference sample of the plurality of extended reference samples is separated from the prediction block by at least one nearest reference sample from a plurality of reference samples that are directly adjacent to the prediction block, and uses a plurality of extended reference samples according to a predetermined set of the plurality of extended reference samples, i.e., a list of area indices can be generated and / or used, and the list can be configured to show a particular subset of reference samples such as only the nearest reference sample, at least one distance of extended reference samples, and / or a combination of different distances of extended reference samples.

[0164] A video encoder can be configured to determine a given set of extended reference samples such that multiple sets differ from one another by the number or combination of lines and / or rows of samples in the image used as a reference sample.

[0165] A video encoder can be configured to determine a predetermined set of multiple extended reference samples based on the block size of the prediction block and / or the prediction mode used to predict the prediction block.

[0166] The video encoder can be configured to determine sets of extended reference samples whose block size is at least a predetermined threshold, and to skip notifications for sets of extended reference samples if the block size falls below the predetermined threshold.

[0167] The predetermined threshold can be a predetermined number of samples along the width or height of the prediction block, and / or a predetermined aspect ratio of the prediction block along the width and height.

[0168] The predetermined number of samples may be any number, but is preferably 8. Alternatively, the aspect ratio may be greater than 1 / 4 and less than 4, based on the number of 8 samples that define one quotient of the aspect ratio.

[0169] A video encoder may be configured to predict a prediction block as a first prediction block using multiple extended reference samples and to predict a second prediction block (which may be part of the same or different image) without using extended reference samples, and the video encoder may be configured to notify a predetermined set of multiple extended reference samples associated with the first prediction block but not a predetermined set of extended reference samples associated with the second prediction block. The predetermined set may be indicated, for example, by a reference region index i.

[0170] The video encoder can be configured to indicate, for each prediction block, the use of the reference sample closest to the one in front of it, along with information indicating one of several specific sets of multiple extended reference samples, and information indicating the in-image prediction mode.

[0171] The video encoder can be configured to signal information indicating an in-image prediction, thereby indicating a prediction mode according to a specific set of multiple extended reference samples, or according to the indicated use of only the closest reference sample.

[0172] The corresponding video decoder decodes the encoded image by encoding the data into video using block-based predictive decoding, which includes in-image prediction. In the in-image prediction, it uses multiple extended reference samples of the image to decode the prediction blocks of the image, and each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample from a plurality of reference samples that are directly adjacent to the prediction block. The decoder is configured to use multiple extended reference samples according to a predetermined set of multiple extended reference samples.

[0173] A video decoder can be configured to determine a given set of extended reference samples such that multiple sets within a set differ from one another in terms of the number or combination of rows of samples in the image used as the reference sample.

[0174] The video decoder can be configured to determine a predetermined set of multiple extended reference samples based on the block size of the prediction block and / or the prediction mode used to predict the prediction block.

[0175] The video encoder can be configured to determine which sets of extended reference samples have a block size of at least a predetermined threshold for the predicted block, and to skip using the set of extended reference samples if the block size falls below the predetermined threshold.

[0176] For example, if the top-left sample of the current block at position (x0, y0) is in the first row of the coding tree block (CTB), the set notification (which can be expressed using parameter information i or similar syntax) can be skipped. The CTB can be considered a basic processing unit where the image is divided and becomes the root of further subdivisions of blocks.

[0177] This can be done by checking whether the vertical y-coordinate y0 is not a multiple of the CTB size. For example, if the CTB size is 64 luminance samples, then in the first row of the CTB, the above blocks in the CTB will have y0=0, and in the second row of the CTB, they will have y0=64, and so on. Thus, for all blocks in the CTB that are above the upper limit of this CTB (which can be checked by, for example, the following modulo operation: y0%CtbSizeY==0), notification of parameter i or similar syntax can be skipped.

[0178] Therefore, a predetermined threshold can be a predetermined number of samples along the width or height of the prediction block, and / or a predetermined aspect ratio of the prediction block along the width and height.

[0179] Therefore, the given sample size can be 8, and / or the aspect ratio may need to be greater than 1 / 4 and less than 4.

[0180] The video decoder can be configured to predict a prediction block as a first prediction block using multiple extended reference samples and to predict a second prediction block without using extended reference samples, and the video decoder can be configured to receive information indicating a predetermined set of multiple extended reference samples associated with the first prediction block, for example using an area index i, and to determine a predetermined set of extended reference samples associated with the second prediction block if there is no signal for each.

[0181] The video decoder can be configured to receive, for each prediction block, information indicating one of several specific sets of multiple extended reference samples, and information indicating the in-image prediction mode, only the use of the reference sample closest to the preceding one.

[0182] The video decoder can thereby be configured to receive information indicating in-image predictions to indicate a prediction mode according to a specific set of multiple extended reference samples, or according to the indicated use of only the nearest reference sample.

[0183] The reference area index can be changed on a per-block basis, so the index can be transmitted in the bitstream of all prediction blocks to which it applies. Embodiments relate to extended reference sample area notifications to enable the decoder to use the correct sample. This area can correspond to a parameter i that indicates an extended reference sample or the nearest reference sample. A predetermined set of extended reference sample lines can be used to trade off the use of notifications and further-distant extended reference samples. For example, only two additional extended reference lines, which are lines with indices 1 and 3, can be used. The set according to embodiments can be I={0,1,3}, where |I|=3. This allows the reference sample area to be extended up to three lines, but only two notifications are required. The set can also be extended to four lines, for example I={0,1,3,4}, in which only the first MaxNumRefAreaIdx element is used, but MaxNumRefAreaIdx can be fixed or notified at the sequence, image, or slice level. The index n to set I can be communicated in a bitstream using entropy coding and truncated single code, as shown in the table in Figure 11, which illustrates an example of truncated single code for a specific set of reference regions of MaxNumRefAreaIdx equal to 3 and 4.

[0184] To account for the different spatial characteristics of different block sizes, the set of reference sample lines can also depend on the predicted block size and / or intra-prediction mode. In another embodiment, the set includes only one additional row at index 2 for small blocks and only two rows at indices 1 and 3 for larger blocks. An embodiment of set selection that depends on the intra-prediction mode is having different sets of reference sample lines for prediction directions between horizontal and vertical.

[0185] The notification of reference region indices can also be restricted to larger block sizes by selecting an empty set for smaller block sizes that do not require notification. This essentially disables the use of extended reference samples for smaller block sizes. In one embodiment, the use of extended reference samples can be restricted to blocks where both the width W and height H are 8 or greater. In addition, blocks where one side is less than a quarter of the other, such as a 32x8 block or an 8x32 block based on the aforementioned symmetry, can also be excluded from the use of extended reference samples. Figures 12a and 12b show this embodiment where several block sizes have already been excluded for in-image prediction (Figure 12a) in general, as indicated by reference numeral 1102, and the blocks are shaded accordingly. In this embodiment, it is assumed that in-image prediction slices and inter-image prediction slices will allow for combinations of different block sizes. Sample 1104 and the accordingly shaded block are not allowed for extended reference samples, i.e., i>0. Figures 12a and 12b show examples of limitations on extended reference samples for intra-image and inter-image predictive slices. As can be seen, block size, and in particular block aspect ratio-dependent block tolerances or limitations, can be symmetric with respect to quotients W / H and H / W.

[0186] If other intra-prediction modes exist that do not use extended reference samples, the reference area index is only notified when a prediction mode that uses extended reference samples is notified. Alternatively, the reference area index can be notified before all other intra-mode information is notified. If the reference area index i indicates extended reference samples (i>0), notifications for modes that do not use extended reference samples (e.g., template matching or trained predictors) can be skipped. This can also skip notification information for certain transformations that do not apply to the prediction residuals of predictions that use extended reference samples.

[0187] The following describes embodiments that refer to considerations for parallel coding.

[0188] The video encoder according to the embodiment can be configured to encode a plurality of prediction blocks into encoded data by block-based predictive coding including in-image prediction, to use a plurality of extended reference samples of the image to encode the prediction blocks of the plurality of prediction blocks for in-image prediction, each extended reference sample of the plurality of extended reference samples is separated from the prediction block by at least one nearest reference sample of a plurality of reference samples directly adjacent to the prediction block, and the video encoder can be configured to determine that the extended reference samples are at least a partial part of the adjacent prediction blocks of the plurality of prediction blocks, to determine that the adjacent prediction blocks have not yet been predicted, and to notify information indicating that the extended prediction samples are associated with the prediction block and have been placed as unavailable samples for the adjacent prediction blocks.

[0189] A video encoder can be configured to encode an image by parallel coding lines of blocks according to a wavefront approach, predicting predicted blocks based on angle predictions, and determining extended reference samples used to predict the predicted blocks so that they are placed in already predicted blocks of the image. According to the wavefront approach, encoding or decoding a second line can follow the decoding of a first line separated by, for example, one block. For example, in a 45° vertical angle mode, starting from the second line, up to one block in the upper right can be decoded and encoded.

[0190] A video encoder can be associated with a prediction block and configured to notify adjacent prediction blocks of augmented prediction samples that are irregularly positioned as unavailable or available samples at the sequence level, i.e., at the sequence, image level, or slice level of an image, where a slice is a portion of an image.

[0191] The video encoder can be associated with a prediction block and can be configured to notify information indicating extended prediction samples that are placed in adjacent prediction blocks as unavailable samples, along with information indicating parallel coding of the image.

[0192] The corresponding decoder can be configured to decode an image encoded with encoded data into video by block-based predictive decoding including in-image prediction, wherein for each image, multiple prediction blocks are decoded, and multiple extended reference samples of the image are used to decode the prediction blocks of the multiple prediction blocks for in-image prediction, wherein each extended reference sample of the multiple extended reference samples is separated from the prediction block by at least one nearest reference sample of multiple reference samples directly adjacent to the prediction block, and the video decoder is configured to determine that the extended reference samples are at least a partial part of the adjacent prediction blocks of the multiple prediction blocks, determine that the adjacent prediction blocks have not yet been predicted, and receive information indicating that the extended prediction samples are associated with the prediction block and placed as unavailable samples for the adjacent prediction blocks.

[0193] The video decoder can be configured to decode the image by parallel decoding lines of blocks according to a wavefront approach and to predict predicted blocks based on angle predictions, and the video decoder can be configured to determine an extended reference sample used to predict the predicted blocks so that they are placed in already predicted blocks of the image.

[0194] A video decoder may be associated with a prediction block and configured to receive information indicating an extended prediction sample that is unavailable or available at the sequence level, image level, or slice level for adjacent prediction blocks.

[0195] The video decoder may be configured to receive information indicating extended prediction samples that are associated with a prediction block and placed in adjacent prediction blocks as unavailable samples, along with information indicating parallel decoding of the image.

[0196] In angular image prediction, a reference sample is copied to the current prediction region along a specified direction. If this direction points to the upper right, the required reference sample region also shifts to the right as the distance to the boundary of the prediction region increases. Figure 13 shows an embodiment of vertical angular prediction at a 45-degree angle for a W×H prediction block. It can be seen that the nearest reference sample region (blue) extends H samples to the upper right of the current W×H block. When the prediction is extended to a more distant reference sample region, the extended reference samples (green) extend H+1, H+2, ... samples to the upper right of the current W×H block.

[0197] For example, when a square coding tree unit (CTU) is used as the basic processing unit, the maximum intra-prediction block size can be equal to the maximum block size, i.e., the CTU block size N × N. Figures 14a and 14b show embodiments from Figure 13 where the W × H prediction block is equal to an N × N CTU block size. Correspondingly, extended reference samples 10641 and / or 10642 span the upper right CTU (CTU2) and allow at least one sample to reach the next CTU (CTU3). That is, some of the extended reference samples 10641 and / or 10642 for the 45° vertical angle prediction can be placed in the unprocessed CTU3. If CTU3 and CTU5 are processed in parallel, these samples can be made unavailable.

[0198] Figures 14a and 14b show examples of the nearest reference sample 1062 and extended reference sample 1064 required for diagonal vertical image prediction.

[0199] In Figure 14b, we can see two CTU lines. If we encode the two lines in parallel using a wavefront-like approach, once CTU2 is encoded and encoding of CTU3 begins, encoding of the second CTU line can begin from CTU5. In this case, since some reference samples 1104 are located inside CTU3, the 45-degree vertical angle prediction using extended reference samples cannot be used for CTU5. The following approach can solve this problem: 1. As described in relation to this embodiment, extended references within the following basic processing areas are marked as unavailable. This allows them to be treated like other unavailable reference samples, for example, at the boundaries of an image, slice, or tile, or when constrained internal predictions are used that do not allow the use of samples from inter-image prediction areas as references to in-image prediction areas.

[0200] 2. To prevent the reference sample from being extended to multiple basic processing units in the upper right of the current region, the use of extended reference samples is permitted only in region and in-image prediction modes.

[0201] Both approaches can be switched via high-level flags at the sequence, image, or slice level. In this way, the encoder can inform the decoder to apply limitations when necessary in the encoder's parallel processing scheme. If not necessary, the encoder can inform the decoder that no limitations will be applied.

[0202] Another approach is to combine the notification of both limitations with the notification of the parallel coding scheme. For example, if wavefront parallelism is enabled and notified in the bitstream, the extended reference sample limit would also be applied.

[0203] Several advantageous embodiments are described below.

[0204] 1. In one embodiment, intra-prediction uses two additional reference lines with reference sample area indices i=1 and i=3 (see Figure 3). Only angular prediction as described herein is permitted for these extended reference sample lines. Notification of reference sample area index i is performed as described in MaxNumRefAreaIdx=3 in relation to Figure 11 and is notified before notification of the intra-prediction mode. If the reference area index is not equal to 0, i.e., an extended reference line is used, the DC and in-planar prediction modes are not used. To avoid unnecessary notifications, the DC and planar modes are excluded from notification of the intra-prediction mode. The intra-prediction mode can be encoded using an index to a list of most likely modes (MPMs). The MPM list contains a fixed number of prediction mode candidates derived from adjacent blocks. Since the DC or planar mode can be used if adjacent blocks are encoded using the nearest reference line (i=0), the MPM list derivation process is modified to exclude the DC and planar modes. If deleted, DC mode and Planar mode can be replaced by Horizontal, Vertical, and Bottom-Left Diagonal mode to fill the list (see Figures 4a to 4e). This is done in a way that avoids redundancy; for example, if the first candidate mode derived from an adjacent block is Vertical and the second candidate mode is DC, DC mode is replaced by Horizontal mode instead of Vertical mode because it is not yet in the list. Another way to prevent unnecessary notifications of DC and Planar modes when i > 0 is to notify the reference region index i after the prediction mode and adjust the notification of i in the prediction mode. If the intra-prediction mode is equal to DC or Planar, index i is not notified. However, if the intra-prediction mode is notified using an index to the MPM list, the aforementioned method takes precedence over this method because it introduces analysis dependencies. Intra-prediction modes that require the derivation of the MPM list must be reconstructed so that i can be analyzed. On the other hand, the derivation of the MPM list refers to the prediction modes of adjacent blocks that must be reconstructed before i can be analyzed.This undesirable parsing dependency is resolved by notifying i before the prediction mode and modifying the derivation of the MPM list accordingly, as described above. The video encoder may be configured to determine a list of multiple most likely prediction modes based on the use of multiple nearest reference samples or multiple extended reference samples for the prediction mode, and the video encoder may be configured to replace prediction modes that are restricted for the reference samples used with the modes allowed in the prediction mode. The corresponding video decoder may be configured to determine a list of most likely prediction modes based on the use of multiple nearest reference samples or multiple extended reference samples for the prediction mode, and the video decoder may be configured to replace prediction modes that are restricted for the reference samples used with the modes allowed in the prediction mode.

[0205] 2. In another embodiment, the intra-prediction of the extended reference sample is further restricted to be applied only to luminance samples. That is, the video encoder may be configured to apply prediction using the extended reference sample to an image containing only luminance information. Thus, the video decoder may be configured to apply prediction using the extended reference sample to an image containing only luminance information.

[0206] 3. In another embodiment, extended reference samples (i>0) (see W+H in Figure 3) that exceed the width and height of the nearest reference sample (i=0) are not generated by using already reconstructed samples (if available), but rather by padding from the last sample, for example, r(23,-1-i) for the upper right sample in Figure 3, and r(-1-i,23) for the lower left sample. This reduces memory access to the extended reference sample line. That is, a video encoder can be configured to generate extended reference samples that exceed the width and / or height of the nearest reference sample along the first and second image directions by padding from the nearest extended reference sample. Thus, a video decoder can be configured to generate extended reference samples that exceed the width and / or height of the nearest reference sample along the first and second image directions by padding from the nearest extended reference sample.

[0207] 4. In another embodiment, only the angle mode per second may be used for the extended reference sample. As a result, the derivation of the MPM list is modified to exclude these modes in addition to the DC mode and planar mode when i>0. That is, the video encoder can be configured to predict predictions using an angle prediction mode that uses only a subset of angles from the possible angles of the angle prediction mode, and to exclude unused angles from notifying the decoder of the encoded information. Thus, the video decoder can be configured to predict predictions using an angle prediction mode that uses only a subset of angles from the possible angles of the angle prediction mode, and to exclude unused angles from the prediction.

[0208] 5. In another embodiment, the number of additional reference sample lines is increased to 3 (see Figure 11 with MaxNumRefAreaIdx=4). That is, the extended reference samples can be placed in at least 2 lines and rows in addition to the nearest reference samples, preferably at least 3 lines and rows. Such configurations can be applied to video encoders and decoders. A video encoder can be configured to use a particular set of multiple extended reference samples to predict a predicted block, and the video encoder is configured to select a particular set from the multiple sets to include the lowest similarity of the image content when compared to the multiple nearest reference samples extended by the set, i.e., to use related reference samples having the same reference area index. Thus, a video decoder can be configured to use a particular set of multiple extended reference samples to predict a predicted block, and the video decoder is configured to select a particular set from the multiple sets to include the lowest similarity of the image content when compared to the multiple nearest reference samples extended by the set.

[0209] 6. In another embodiment, instead of index i, a flag is provided indicating whether an extended reference sample is to be used. If the flag indicates that an extended reference sample is to be used, index i (i>0) is derived by calculating the similarity between reference line i>0 and reference line i=0 (e.g., using the sum of absolute differences). The index i with the lowest similarity is selected. The idea behind this is that the higher the correlation between the extended reference sample (i>0) and the normal reference sample (i=0), the higher the correlation of the predicted result, so there is no additional benefit to using the extended reference sample in the prediction. That is, the video encoder can be configured to indicate the use of an extended reference sample using a flag or other, possibly binary, information. Thus, the video decoder can be configured to receive information indicating the use of an extended reference sample by such a flag.

[0210] 7. In another embodiment, a quadratic transform, such as an inseparable quadratic transform (NSST), may be applied after the first transform of the intraprediction residuals. For extended reference sample lines, no quadratic transform is performed, and all notifications related to NSST are disabled if i > 0. That is, the video encoder may be configured to selectively use only extended reference samples or the nearest reference samples, and the video encoder may be configured to transform the residuals obtained by predicting the predicted blocks using the first transform procedure to obtain a first transform result, and to transform the first transform result using the second transform procedure to obtain a second transform result when the extended reference samples are not used to predict the predicted blocks. This may also affect notifications of whether the second transform is used. That is, if it is necessary to notify of the use of the second transform, the notification may be skipped when extended reference samples are used. The video encoder may be configured to notify of the use of a quadratic transform, or implicitly notify of the non-use of a quadratic transform when indicating the use of extended reference samples, and not include information related to the results of the quadratic transform in the encoded data.

[0211] Therefore, the video decoder can be configured to selectively use only extended reference samples or the nearest reference samples, and the video decoder is configured to transform the residuals obtained by predicting the predicted blocks using a first transformation procedure to obtain a first transformation result, and to transform the first transformation result using a second transformation procedure to obtain a second transformation result when the extended reference samples are not used to predict the predicted blocks. The video decoder can be configured to receive information indicating the use of a quadratic transformation, or to derive the non-use of a quadratic transformation when it indicates the use of extended reference samples, and not receive information related to the results of a quadratic transformation of the encoded data.

[0212] 8. In another embodiment, prediction using extended reference samples (i>0) is combined with planar prediction using the nearest reference sample (i=0). As outlined in relation to the different combinations of extended prediction modes, namely, using extended reference samples, the weighting can be fixed (e.g., 0.5 and 0.5), dependent on block size, or specified at the slice, image, or sequence level. When extended reference samples are used (i>0), an additional flag indicates whether combined prediction is applied. That is, the video encoder can be configured. Thus, the video decoder can be configured.

[0213] 9. In another embodiment, to reduce notification overhead, notification of combined predictions from the above is omitted. Instead, the decision of whether to apply combined predictions is derived from an analysis of the nearest reference sample (i=0). One possible analysis is the flatness of the nearest reference sample. If the nearest reference sample signal is flat (no edges), the combination is applied; if it contains high frequencies and edges, the combination is not applied. That is, the video encoder can be configured. Thus, the video decoder can be configured.

[0214] While several embodiments have been described in the context of the apparatus, it is clear that these embodiments also represent descriptions of the corresponding methods, where blocks or apparatus correspond to method steps or features of method steps. Similarly, embodiments described in the context of method steps also represent descriptions of the corresponding blocks, items, or functions of the corresponding apparatus.

[0215] Depending on specific implementation requirements, embodiments of the present invention may be implemented in hardware or software. Implementation can be carried out using digital storage media such as floppy disks, DVDs, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memory, which store electronically readable control signals and cooperate (or can cooperate) with a computer system that is programmable to perform each method.

[0216] Some embodiments of the present invention include a data carrier having an electronically readable control signal that can cooperate with a programmable computer system so that one of the methods described herein can be performed.

[0217] Generally, embodiments of the present invention can be implemented as a computer program product comprising program code, which operates to perform one of the methods when the computer program product is executed on a computer. The program code may be stored, for example, in a machine-readable carrier.

[0218] Other embodiments include a computer program stored in a machine-readable carrier for performing one of the methods described herein.

[0219] In other words, embodiments of the method of the present invention are computer programs having program code for performing one of the methods of the present invention when the computer program is executed on a computer.

[0220] Therefore, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) on which a computer program for performing one of the methods described herein is recorded.

[0221] Therefore, a further embodiment of the method of the present invention is a data stream or sequence of signals representing a computer program for performing one of the methods described herein. The data stream or sequence of signals may be configured to be transmitted over a data communication connection, such as the Internet.

[0222] Further embodiments include processing means configured or adapted to perform one of the methods described herein, such as a computer or a programmable logic device.

[0223] Further embodiments include a computer on which a computer program for performing one of the methods described herein is installed.

[0224] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the method herein. In some embodiments, a field-programmable gate array can cooperate with a microprocessor to perform one of the methods herein. Generally, the method is preferably performed by any hardware device.

[0225] The embodiments described above are merely illustrative of the principles of the present invention. Modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Therefore, the invention is intended to be limited only by the imminent claims and not by the descriptions of the embodiments herein or by the specific details presented herein.

Claims

1. A method for video decoding, For the prediction block, decode the reference sample index whose value is greater than 0, Identifying the prediction mode for the prediction block, and that the prediction mode is an angle prediction mode. Determining whether a first reference sample is available at a first position within the reference line, and that the reference line is spatially separated from the prediction block. If it is determined that the first reference sample at the first position is unavailable, the availability of each reference sample included in the reference line is determined sequentially, starting with the second reference sample adjacent to the first reference sample, until it is determined that the third reference sample at the second position is available. Setting the value of the first reference sample to the value of the third reference sample at the second position, To generate at least one additional reference sample having the value of a further reference sample located at the end of the aforementioned reference line, Decoding the prediction block using the identified prediction mode, the value of the first reference sample, and the value of the further reference sample located at the end of the reference line, The method, including the method described above.

2. After the value of the first reference sample is set, the value of at least one reference sample located between the first position and the second position is set to the value of the first reference sample. The method according to claim 1, further comprising:

3. Identifying a reference line from multiple reference lines based on the syntax elements of the bitstream of encoded video data, The method according to claim 1, further comprising:

4. The aforementioned reference line is an extended reference line separated from the prediction block by at least one reference line, The aforementioned method, Determining whether an adjacent reference line directly adjacent to the prediction block is available, If it is determined that the adjacent reference line is available, then it is determined whether the first reference sample at the first position in the extended reference line is available, The method according to claim 1, further comprising:

5. Filtering the reference sample of the aforementioned reference line, Decoding the prediction block using the value of the reference sample of the reference line and the value of at least one additional reference sample, The method according to claim 1, further comprising:

6. The reference line is spatially separated from the prediction block by at least one sample position. The method according to claim 1.

7. A video decoder, For the prediction block, decode the reference sample index whose value is greater than 0. Identify the prediction mode for the prediction block, and the prediction mode is an angle prediction mode. Determine whether a first reference sample is available at a first position within the reference line, and that the reference line is spatially separated from the prediction block. If it is determined that the first reference sample at the first position is unavailable, the availability of each reference sample included in the reference line is sequentially determined, starting with the second reference sample adjacent to the first reference sample, until it is determined that the third reference sample at the second position is available. The value of the first reference sample is set to the value of the third reference sample at the second position. Generate at least one additional reference sample having the value of a further reference sample located at the end of the aforementioned reference line. The prediction block is decoded using the identified prediction mode, the value of the first reference sample, and the value of the further reference sample located at the end of the reference line. The decoder is configured in such a way.

8. After the value of the first reference sample is set, the value of at least one reference sample located between the first position and the second position is set to the value of the first reference sample. The decoder according to claim 7, further configured as follows.

9. From multiple reference lines, the reference lines are identified based on the syntax elements of the bitstream of the encoded video data. The decoder according to claim 7, further configured as follows.

10. The aforementioned reference line is an extended reference line, The decoder mentioned above is Determine whether an adjacent reference line directly adjacent to the prediction block is available. If it is determined that the adjacent reference line is available, then it is determined whether the first reference sample at the first position in the extended reference line is available. The decoder according to claim 7, further configured as follows.

11. Filter the reference samples of the aforementioned reference line, The prediction block is decoded using the value of the reference sample of the reference line and the value of the at least one additional reference sample. The decoder according to claim 7, further configured as follows.

12. The reference line is spatially separated from the prediction block by at least one sample position. The decoder according to claim 7.

13. A non-temporary computer-readable storage medium, When executed, at least one processor, For the prediction block, decode the reference sample index whose value is greater than 0. The prediction mode for the prediction block is identified, and the prediction mode is an angle prediction mode. Determine whether a first reference sample is available at a first position within the reference line, and the reference line is spatially separated from the prediction block. If it is determined that the first reference sample at the first position is unavailable, the system will sequentially determine whether each reference sample included in the reference line is available, starting with the second reference sample adjacent to the first reference sample, until it is determined that the third reference sample at the second position is available. The value of the first reference sample is set to the value of the third reference sample at the second position. To generate at least one additional reference sample having the value of a further reference sample located at the end of the aforementioned reference line, The prediction block is decoded using the identified prediction mode, the value of the first reference sample, and the value of the further reference sample located at the end of the reference line. A non-temporary, computer-readable storage medium that stores instructions.

14. When executed, at least one of the processors, After the value of the first reference sample is set, the value of at least one reference sample located between the first position and the second position is set to the value of the first reference sample. A non-temporary computer-readable storage medium according to claim 13, further storing instructions.

15. When executed, at least one of the processors, Based on the syntax elements of the bitstream of encoded video data, the reference lines are identified from multiple reference lines. A non-temporary computer-readable storage medium according to claim 13, further storing instructions.

16. The aforementioned reference line is an extended reference line, The non-temporary computer-readable storage medium, when executed, is connected to at least one processor, Determine whether an adjacent reference line directly adjacent to the prediction block is available. If it is determined that the adjacent reference line is available, then it is determined whether the first reference sample at the first position in the extended reference line is available. A non-temporary computer-readable storage medium according to claim 13, further storing instructions.

17. When executed, at least one of the processors, Filter the reference sample of the aforementioned reference line, The prediction block is decoded using the value of the reference sample of the reference line and the value of the at least one additional reference sample. A non-temporary computer-readable storage medium according to claim 13, further storing instructions.

18. The reference line is spatially separated from the prediction block by at least one sample position. A non-temporary computer-readable storage medium according to claim 13.