Extended reference intra-picture prediction

By employing extended reference samples and advanced filtering techniques, the video codec improves parallel processing and prediction accuracy in video encoders and decoders, addressing inefficiencies in existing codecs like HEVC.

JP7805403B2Active Publication Date: 2026-01-23FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024113317
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-06-29
Filing Date
2024-07-16
Publication Date
2026-01-23
Estimated Expiration
2039-06-28

AI Technical Summary

Technical Problem

Existing video codecs like HEVC do not efficiently support parallel processing capabilities in video encoders and decoders for intra-picture prediction, particularly in handling reference samples for predictive blocks.

Method used

The use of extended reference samples, which are not directly adjacent to the predictive block, for intra-picture prediction, along with methods to determine availability, replace unavailable samples, and apply filtering techniques to enhance prediction accuracy and efficiency.

Benefits of technology

Enhances the efficiency and accuracy of video encoding and decoding by utilizing extended reference samples, allowing for more effective parallel processing and improved prediction modes, even when direct reference samples are unavailable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007805403000103
    Figure 0007805403000103
  • Figure 0007805403000104
    Figure 0007805403000104
  • Figure 0007805403000105
    Figure 0007805403000105
Patent Text Reader

Abstract

To provide a video encoder, a video decoder, and a video encoding and decoding method for encoding images of a video into encoded data by block-based predictive coding, including intra-image prediction.SOLUTION: For intra-picture prediction 1001, a video encoder 1000 uses a plurality of nearest reference samples and a plurality of extended reference samples of an image directly adjacent to a predictive block to encode a predictive block of the image, separates each extended reference sample of the plurality of extended reference samples from the predictive block by at least one nearest reference sample of the plurality of reference samples, sequentially determines the availability or unavailability of each of the plurality of nearest reference samples, replaces the nearest reference sample determined to be unavailable by a replacement sample, and uses the replacement sample for intra-picture prediction.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to video coding, in particular hybrid video coding including intra-picture prediction. The present invention further relates to a video encoder, a video decoder and methods for video encoding and decoding, respectively. [Background technology]

[0002] H.265 / HEVC is a video codec that already provides tools for enhancing or enabling parallel processing in the encoder and / or decoder. For example, HEVC supports subdivision of an image into an array of tiles that are coded independently of each other. Another concept supported by HEVC is related to WPP, whereby rows or CTUs of an image can be processed in parallel (i.e., in stripes) from left to right, provided that a minimum CTU offset is respected in the processing of consecutive CTU lines. However, it would be preferable to have a video codec at hand that more efficiently supports parallel processing capabilities in a video encoder and / or video decoder. Summary of the Invention [Problem to be solved by the invention]

[0003] It is therefore an object of the present invention to provide a video codec that allows more efficient processing in the encoder and / or decoder with respect to reference samples used to predict a predictive block. [Means for solving the problem]

[0004] This object is achieved by the subject matter of the independent claims of the present application.

[0005] According to an embodiment, a video encoder is configured to encode an image of a video into coded data by block-based predictive coding, where the block-based predictive coding includes intra-image prediction. For intra-image prediction, the video encoder is configured to use multiple extended reference samples of the image to encode a predictive block of the image, where each extended reference sample of the multiple extended reference samples is separated from the predictive block by at least one nearest reference sample of the multiple reference samples that is directly adjacent to the predictive block. The video encoder is further configured to sequentially determine the availability or unavailability of each of the multiple nearest reference samples and replace the nearest reference sample determined to be unavailable with a replacement sample. The video encoder is configured to use the replacement sample for intra-image prediction. This allows the concept of prediction using the nearest reference sample to be used even when such a sample is unavailable, which may occur, for example, when a buffer / memory has a line or row of samples but the memory does not actually have a column of samples and therefore is unavailable.

[0006] According to an embodiment, a video encoder is configured to encode an image of a video into coded data by block-based predictive coding, the block-based predictive coding including intra-image prediction. The video encoder is configured to use a plurality of extended reference samples of an image for encoding a predictive block of the image in the intra-image prediction, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples that is directly adjacent to the predictive block. The video encoder is further configured to filter at least a subset of the plurality of extended reference samples using a bilateral filter to obtain a plurality of filtered extended reference samples, and to use the plurality of filtered extended reference samples for the intra-image prediction.

[0007] According to an embodiment, a video encoder is configured to encode an image of a video into coded data by block-based predictive coding, the block-based predictive coding including intra-image prediction. The video encoder is configured to use multiple extended reference samples of the image to encode a predictive block of the image in the intra-image prediction, each extended reference sample of the multiple extended reference samples being separated from the predictive block by at least one nearest reference sample of the multiple reference samples that is directly adjacent to the predictive block, the multiple nearest reference samples being arranged along a first image direction of the predictive block and along a second image direction of the predictive block, and to map at least some of the nearest reference samples arranged along the second direction to the extended reference samples arranged along the first direction such that the mapped reference samples exceed an extension of the predictive block along the first image direction. The video encoder is configured to use the mapped extended reference samples for prediction.

[0008] According to an embodiment, a video encoder is configured to encode an image of a video into coded data by block-based predictive coding, the block-based predictive coding including intra-image prediction. The video encoder is configured to use, for the intra-image prediction, a plurality of nearest reference samples of images directly neighboring the predictive block and a plurality of extended reference samples to encode a predictive block of the image, each extended reference sample of the plurality of reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples. The video encoder is configured to boundary filter in a mode in which the extended samples are not used and not use boundary filtering when the extended samples are used, or the video encoder is configured to boundary filter at least a subset of the plurality of nearest reference samples and not use boundary filtering for the extended samples.

[0009] According to an embodiment, a video encoder is configured to encode an image of a video into coded data by block-based predictive coding, the block-based predictive coding including intra-image prediction. For intra-image prediction, the video encoder is configured to determine a plurality of closest reference samples of images directly neighboring the predictive block and a plurality of extended reference samples to encode a predictive block of the image, where each extended reference sample of the plurality of reference samples is separated from the predictive block by at least one closest reference sample of the plurality of extended reference samples. The video encoder is further configured to determine a prediction of the predictive block using the extended reference samples, filter the extended reference samples to obtain a plurality of filtered extended reference samples, and combine the prediction and the filtered extended reference samples to obtain a combined prediction of the predictive block.

[0010] According to an embodiment, a video encoder is configured to encode an image of a video into coded data by block-based predictive coding, the block-based predictive coding including intra-image prediction. The video encoder is configured to use, for intra-image prediction, a plurality of nearest reference samples and / or a plurality of extended reference samples of images directly neighboring the predictive block to encode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, determine a first prediction of the predictive block using a first prediction mode of a set of prediction modes, the first prediction mode including a prediction mode that uses the plurality of nearest reference samples in the absence of an extended reference sample, and determine a second prediction of the predictive block using a second prediction mode of the second set of prediction modes, the second prediction mode including a subset of the prediction modes of the first set, the subset being associated with the plurality of extended reference samples. The video encoder is configured to: i (x, y)) and weighted (w0;w i) to obtain a combined prediction (p(x, y)) as a prediction for a predictive block of encoded data.

[0011] According to an embodiment, a video encoder encodes an image of a video into coded data by block-based predictive coding including intra-image prediction, and for the intra-image prediction, uses a plurality of nearest reference samples and / or a plurality of extended reference samples of images directly neighboring the predictive block to encode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, e.g., in the absence of an extended reference sample, uses a prediction mode that is one of a first set of prediction modes for predicting the predictive block using the nearest reference sample, or uses a prediction mode for predicting the predictive block using the extended reference sample. The decoder is configured to use a prediction mode that is one of a second set of modes, the second set of prediction modes being a subset of the first set of prediction modes, signal mode information (m) indicating a prediction mode used to predict a prediction block, and then, if the prediction mode is included in the second set of prediction modes, signal parameter information (i) indicating a subset of extended reference samples used for the prediction mode, and skip signaling the parameter information if the prediction mode used is not included in the second set of prediction modes, thereby enabling a decoder to conclude that a specific value of the parameter is selected or determined, wherein a predetermined property enables skipping the signaling, i.e., the absence of a signal is given a useful meaning. For example, its absence can indicate that the nearest reference sample should be used.

[0012] According to an embodiment, a video encoder encodes an image of a video into coded data by block-based predictive coding including intra-image prediction. For the intra-image prediction, the encoder uses a plurality of reference samples, including a nearest reference sample of an image immediately adjacent to the predictive block, and a plurality of extended reference samples to encode a predictive block of the image, each extended reference sample being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples. The encoder is configured to use a prediction mode that is one of a first set of prediction modes for predicting the predictive block using the nearest reference sample or a second set of prediction modes for predicting the predictive block using the extended reference sample, the second set of prediction modes being a subset of the first set of prediction modes. The video encoder can generate the first set and / or the second set using available reference data and / or can determine the set using information derived from the image. The second set being a subset of the first set includes cases where both sets are equal. The video encoder is configured to signal parameter information indicating a subset of a plurality of reference samples used for a prediction mode, the subset of the plurality of reference samples including only the nearest reference sample or the extended reference sample, and then signal mode information (m) indicating a prediction mode to be used to predict a prediction block, the mode information indicating a prediction mode from the subset of modes, the subset being restricted to a set of prediction modes allowed according to the parameter information (i). Based on the association of the used reference sample, i.e., nearest or extended, only the prediction mode associated with the indicated reference sample is applied, thereby enabling identification of the restricted set.

[0013] According to an embodiment, a video encoder is configured to encode an image of a video into coded data by block-based predictive coding including intra-image prediction, and for the intra-image prediction, to encode a predictive block of the image using a plurality of nearest reference samples and / or a plurality of extended reference samples of images directly neighboring the predictive block, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, to determine a first prediction of the predictive block using a first prediction mode of a set of prediction modes, the first set of prediction modes including a prediction mode that uses the plurality of nearest reference samples in the absence of an extended reference sample, to determine a second prediction of the predictive block using a second prediction mode of a second set of prediction modes, the second set including a subset of the prediction modes of the first set being associated with the plurality of extended reference samples, and to combine the first prediction and the second prediction to obtain a combined prediction as a prediction of the predictive block in the coded data.

[0014] According to an embodiment, a video encoder is configured to encode an image of a video into coded data by block-based predictive coding including intra-image prediction, wherein in the intra-image prediction, the encoder uses a plurality of extended reference samples of the image to encode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples directly neighboring the predictive block, and the encoder uses the plurality of extended reference samples according to a predetermined set of the plurality of extended reference samples. The plurality of extended reference samples can, for example, be included in a list of area indexes identified by an identifier.

[0015] According to an embodiment, a video encoder is configured to encode a plurality of predictive blocks into coded data by block-based predictive coding including intra-picture prediction, and for the intra-picture prediction, to use a plurality of extended reference samples of an image to encode a predictive block of the plurality of predictive blocks, where each extended reference sample of the plurality of extended reference samples is separated from the predictive block by at least one nearest reference sample among the plurality of reference samples directly neighboring the predictive block. The video encoder is configured to determine the extended reference sample to be at least a partial part of a neighboring predictive block of the plurality of predictive blocks, determine that the neighboring predictive block has not yet been predicted, and signal information associated with the predictive block and indicating the extended predictive sample located as an unavailable sample in the neighboring predictive block.

[0016] According to an embodiment, a video decoder is configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, wherein in the intra-image prediction, a plurality of extended reference samples of the image are used to encode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples that is directly adjacent to the predictive block, sequentially determine the availability or unavailability of each of the plurality of nearest reference samples, replace the nearest reference sample determined to be unavailable with the replacement sample, and use the replacement sample for the intra-image prediction.

[0017] According to an embodiment, a video decoder is configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, in which the video decoder uses a plurality of extended reference samples of the image to decode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample among the plurality of reference samples directly adjacent to the predictive block, filter at least a subset of the plurality of extended reference samples using a bilateral filter to obtain a plurality of filtered extended reference samples, and use the plurality of filtered extended reference samples for the intra-image prediction.

[0018] According to an embodiment, a video decoder is configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, and in the intra-image prediction, to use a plurality of extended reference samples of the image to decode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample among the plurality of reference samples directly adjacent to the predictive block, the plurality of nearest reference samples being arranged along a first image direction of the predictive block and along a second image direction of the predictive block, to map at least some of the nearest reference samples arranged along the second direction to the extended reference samples arranged along the first direction such that the mapped reference samples exceed an extension of the predictive block along the first image direction, and to use the mapped extended reference samples for prediction.

[0019] According to an embodiment, a video decoder is configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, and for the intra-image prediction, to decode a predictive block of the image, use a plurality of nearest reference samples and a plurality of extended reference samples of images directly neighboring the predictive block, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples. The video decoder is configured to perform boundary filtering in a mode in which no extended samples are used, and to not use boundary filtering when extended samples are used, or to border filter at least a subset of the plurality of nearest reference samples and not use boundary filtering for the extended samples.

[0020] According to an embodiment, a video decoder is configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, and in the intra-image prediction, to decode a predictive block of the image, determine a plurality of nearest reference samples and a plurality of extended reference samples of an image directly adjacent to the predictive block, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, determine a prediction of the predictive block using the extended reference samples, filter the extended reference samples to obtain a plurality of filtered extended reference samples, and combine the prediction and the filtered extended reference samples to obtain a combined prediction of the predictive block.

[0021] According to an embodiment, a video decoder is configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, and for the intra-image prediction, to decode a predictive block of the image, the decoder uses a plurality of nearest reference samples and / or a plurality of extended reference samples of images immediately adjacent to the predictive block, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, determine a first prediction of the predictive block using a first prediction mode of a set of prediction modes, the first set of prediction modes including a prediction mode that uses the plurality of nearest reference samples in the absence of an extended reference sample, and determine a second prediction of the predictive block using a second prediction mode of a second set of prediction modes, the second set of prediction modes including a subset of the prediction modes of the first set, the subset being associated with the plurality of extended reference samples. The video decoder is configured to weightedly combine the first prediction and the second prediction to obtain a combined prediction as a prediction of the predictive block in the encoded data.

[0022] According to an embodiment, a video decoder is configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, and for the intra-image prediction, to decode a predictive block of the image, use a plurality of nearest reference samples and / or a plurality of extended reference samples of images directly adjacent to the predictive block, wherein each extended reference sample of the plurality of extended reference samples is separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, and use a prediction mode that is one of a first set of prediction modes for predicting the predictive block using the nearest reference samples or a prediction mode that is one of a second set of prediction modes for predicting the predictive block using the extended reference samples, wherein the second set of prediction modes is a subset of the first set of prediction modes, receive mode information (m) indicating a prediction mode used to predict the predictive block, and then receive parameter information (i) indicating a subset of the extended reference samples used for the prediction mode, thereby indicating that the prediction mode is included in the second set of prediction modes, and if no parameter information is received, determine that the used prediction mode is not included in the second set of prediction modes and determine the use of the nearest reference sample for the prediction.

[0023] According to an embodiment, a video decoder is configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, and for the intra-image prediction, to decode a predictive block of the image, use a plurality of reference samples and a plurality of extended reference samples including a nearest reference sample of an image directly adjacent to the predictive block, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, use a prediction mode that is one of a first set of prediction modes for predicting the predictive block using the nearest reference sample or use a prediction mode that is one of a second set of prediction modes for predicting the predictive block using the extended reference samples, the second set of prediction modes being a subset of the first set of prediction modes, receive parameter information (i) indicating a subset of the plurality of reference samples to be used for the prediction mode, the subset of the plurality of reference samples including only the nearest reference sample or at least one extended reference sample, and then receive mode information (m) indicating a prediction mode to be used to predict the predictive block, wherein the mode information indicates a prediction mode from the subset of modes, the subset being restricted to a set of prediction modes allowed according to the parameter information (i).

[0024] According to an embodiment, a video decoder is configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, and for the intra-image prediction, to decode a predictive block of the image, the decoder uses a plurality of nearest reference samples and / or a plurality of extended reference samples of images immediately neighboring the predictive block, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, determine a first prediction of the predictive block using a first prediction mode of a set of prediction modes, the first set of prediction modes including a prediction mode that uses the plurality of nearest reference samples in the absence of an extended reference sample, determine a second prediction of the predictive block using a second prediction mode of a second set of prediction modes, and a second set of prediction modes including a subset of the prediction modes of the first set associated with the plurality of extended reference samples. The video decoder is configured to combine the first prediction and the second prediction to obtain a combined prediction as a prediction of the predictive block in the encoded data.

[0025] According to an embodiment, a video decoder is configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, and in the intra-image prediction, to use a plurality of extended reference samples of the image to decode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples directly adjacent to the predictive block, and to use the plurality of extended reference samples according to a predetermined set of the plurality of extended reference samples.

[0026] According to an embodiment, a video decoder is configured to decode images encoded by encoding data into video by block-based predictive decoding including intra-image prediction, where for each image, a plurality of predictive blocks are decoded, and for the intra-image prediction, a plurality of extended reference samples of the image are used to decode a predictive block of the plurality of predictive blocks, where each extended reference sample of the plurality of extended reference samples is separated from the predictive block by at least one nearest reference sample of the plurality of reference samples directly neighboring the predictive block. The video decoder is configured to determine the extended reference sample to be at least a partial part of a neighboring predictive block of the plurality of predictive blocks, determine that the neighboring predictive block has not yet been predicted, and receive information associated with the predictive block and indicating the extended predictive sample located as an unavailable sample in the neighboring predictive block.

[0027] Further embodiments relate to methods for encoding and decoding video, and to computer program products.

[0028] With regard to the aforementioned embodiments of the present patent application, it should be noted that same can be combined such that two or more of the aforementioned embodiments, such as all embodiments, are implemented in a video codec at the same time.

[0029] Further advantageous embodiments of the present patent application are the subject of the dependent claims, preferred embodiments of which are described below with reference to the figures therein. [Brief explanation of the drawings]

[0030] [Figure 1] 1 shows a schematic block diagram of a video encoder according to an embodiment, comprising a decoder according to an embodiment; [Figure 2] 1 shows a schematic flowchart of a method for encoding a video stream according to an embodiment; [Figure 3] 1 shows examples of immediately adjacent (nearest) reference samples and extended reference samples used in an embodiment. [Figure 4a] 10 illustrates an example of five intra-picture prediction angles for a 4x2 block of prediction samples according to an embodiment. [Figure 4b] 10 illustrates an example of five intra-picture prediction angles for a 4x2 block of prediction samples according to an embodiment. [Figure 4c] 10 illustrates an example of five intra-picture prediction angles for a 4x2 block of prediction samples according to an embodiment. [Figure 4d] 10 illustrates an example of five intra-picture prediction angles for a 4x2 block of prediction samples according to an embodiment. [Figure 4e] 10 illustrates an example of five intra-picture prediction angles for a 4x2 block of prediction samples according to an embodiment. [Figure 4f] 1 shows a schematic diagram to illustrate the direction of angle prediction used in an embodiment; [Figure 4g] 1 shows a table to illustrate an example of the dependency of the number of taps used in the filter, which number depends on the block size and prediction mode of the prediction block. [Figure 4h] 1 shows a table to illustrate an example of the dependency of the number of taps used in the filter, which number depends on the block size and prediction mode of the prediction block. [Figure 5a] 1 illustrates an embodiment relating to angle prediction using angular parameter definitions. [Figure 5b] 1 illustrates an embodiment relating to angle prediction using angular parameter definitions. [Figure 5c] 1 illustrates an embodiment relating to angle prediction using angular parameter definitions. [Figure 6a] 10 illustrates the derivation of a vertical offset associated with a mapping of reference samples according to an embodiment. [Figure 6b] 10 illustrates the derivation of a vertical offset associated with a mapping of reference samples according to an embodiment. [Figure 6c]10 illustrates the derivation of a vertical offset associated with a mapping of reference samples according to an embodiment. [Figure 7a] 10 illustrates the derivation of a horizontal offset according to an embodiment. [Figure 7b] 10 illustrates the derivation of a horizontal offset according to an embodiment. [Figure 7c] 10 illustrates the derivation of a horizontal offset according to an embodiment. [Figure 8] The upper left corner of the diagonal shows an embodiment and the use of the nearest reference sample according to the embodiment. [Figure 9] 10 illustrates an example of projection of an extended left reference sample as a side reference adjacent to an extended top reference sample as a primary reference in the case of top-left diagonal prediction according to an embodiment. [Figure 10] 10 illustrates an example of a projection of the nearest left reference sample as a side reference next to an extended top reference sample according to an embodiment. [Figure 11] 1 illustrates an exemplary truncated single code for a particular set of reference regions according to an embodiment. [Figure 12a] 1 shows a schematic diagram of available block sizes according to an embodiment; [Figure 12b] 1 shows a schematic diagram of available block sizes according to an embodiment; [Figure 13] 10 shows a schematic diagram of an embodiment of vertical angle prediction with a prediction block angle of 45 degrees. [Figure 14a] 10 illustrates examples of nearest reference samples and extended reference samples required for diagonal vertical intra prediction according to an embodiment. [Figure 14b] 10 illustrates examples of nearest reference samples and extended reference samples required for diagonal vertical intra prediction according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0031] Equal or equivalent elements or elements having equal or equivalent functionality are indicated in the following description by equal or equivalent reference signs, even if they occur in different figures.

[0032] In the following description, numerous details are set forth to provide a more thorough explanation of embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form, rather than in detail, to avoid obscuring the embodiments of the present invention. Furthermore, features of different embodiments described below may be combined with each other unless otherwise specified.

[0033] Hybrid video coding uses intra-picture prediction to encode regions of image samples by generating a prediction signal from available neighboring samples, i.e., reference samples. The prediction signal is subtracted from the original signal to obtain a residual signal. This residual signal, or prediction error, is further transformed, scaled, quantized, and entropy coded, as shown, for example, in FIG. 1 , which shows a schematic block diagram of a video encoder 1000 according to an embodiment, which is a hybrid video encoder with an intra-picture prediction block 1001. The video encoder 1000 is configured to receive an input video signal 1002 containing multiple images, where the sequence of images forms a video. The video encoder 1000 includes a block 1004 configured to divide the signal 1002 into regions of samples, i.e., to form blocks from the input video signal 1002. A controller 1006 of the video encoder 1000 is configured to control the block 1004 and a decoder 1008, which may be part of the encoder 1000. A decoder for receiving and decoding the output bitstream 1012 and the generated output video signal 1014, i.e., the encoded data, can be implemented accordingly. In particular, the transform, scaling, and quantization block 1016, together with the block 1018 for motion estimation of the signal 1022, which is the input video signal 1002 divided into blocks by the block 1004, can provide information on both the quantized transform coefficients and the motion information to enable entropy coding of the output bitstream 1012.

[0034] The quantized transformed coefficients are then scaled and inverse transformed to produce a reconstructed residual signal, prior to a potential in-loop filtering operation. This signal can be added back to the prediction signal to obtain a reconstruction that is also usable at the decoder. The reconstructed signal can be used to predict subsequent samples in coding order within the same image.

[0035] For details of intra-picture prediction, see FIG. 2. First, reference samples used for prediction are generated based on the reconstructed samples in block 1042. This step also includes, for example, replacing neighboring samples that are unavailable at picture, slice, or tile boundaries. Second, in block 1044, the reference samples can be filtered to eliminate discontinuities in the reference signal. Third, in block 1046, prediction samples are calculated using the reference samples according to an intra-picture prediction mode. The prediction mode describes how the prediction signal is generated from the reference samples, for example, by averaging them in DC mode or copying them along one prediction angle in angle prediction mode. The encoder must determine which intra-picture prediction mode to select, and the selected intra-picture prediction mode is signaled in the bitstream to the decoder through entropy coding. At the decoder side, the intra-picture prediction mode is extracted from the bitstream through entropy decoding. Fourth, and perhaps finally, in block 1048, the prediction samples can also be filtered to smooth the signal. In other words, FIG. 2 shows a flowchart of an intra-picture prediction process or method. Generally, the correlation between samples in an image decreases as the distance increases. Therefore, directly adjacent samples are generally suitable as reference samples for predicting the area of ​​a sample. However, there are cases where directly adjacent reference samples represent edges or objects in uniform regions (occlusions). In these cases, the correlation between the sample to be predicted (uniform or textured region) and the directly adjacent reference samples (edges) is low. Extended reference intra-picture prediction solves this problem by incorporating more distant reference samples that are not directly adjacent. Although the concept of extending the nearest reference sample is known, several novel improvements to all parts of the intra-picture prediction process and notification are defined in embodiments of the present invention and are described below.

[0036] Extended reference intra-picture prediction allows for generating a prediction signal for a sample region using extended reference samples. Extended reference samples are available reference samples that are not directly adjacent. Below, improved reference sample generation, filtering, prediction, and prediction filtering using extended reference samples according to embodiments are described in more detail. Special cases that combine prediction using extended reference samples with prediction using directly adjacent samples or unfiltered reference samples are described later. Various methods according to embodiments are then described for improving prediction modes and extended reference region signaling for extended reference samples. Furthermore, embodiments for facilitating parallel coding using extended reference samples are described.

[0037] For the generation of reference samples, current video coding standards predict the current block using directly neighboring samples. In addition to the nearest directly neighboring samples, the use of multiple reference lines has been proposed in the literature. The additional reference lines used in intra-picture prediction are further referred to as extended reference samples in the following. Then, an improved method for replacing unavailable extended reference samples according to an embodiment is described.

[0038] An example showing the nearest reference sample line and three extended reference sample lines of a predicted 16x8 block is shown in Figure 3, which shows examples of directly adjacent (nearest) reference sample 1062 and extended reference samples 10641, 10642 and 10643.

[0039] The nearest reference sample 1062 and extended reference samples 10641, 10642, and 10643 are located adjacent to the predicted block 1066 and are arranged along two directions in the image: direction x, and direction y, which is perpendicular to direction x. Along direction x, the predicted block includes an extension W with samples ranging from 0 to W-1. Along direction y, the predicted block includes an extension H with samples ranging from 0 to H-1.

[0040] A reference region with index i may indicate the distance between the respective reference samples, i.e., the nearest reference sample with index i=0, i.e., the reference samples located directly adjacent and where the extended reference sample is spaced from the prediction block 1066 by at least the nearest reference sample 1062. For example, the reference region index i may indicate the extension of the distance between the prediction block 1066 and the respective reference sample 1062 or 1064. As an example, increasing the parameter x along the direction x may be referred to as moving right, and conversely, decreasing x may be referred to as moving left.

[0041] Alternatively, or in addition, decreasing the index i along the negative direction y can be referred to as moving upward or towards the top of the image, and increasing the parameter y can be referred to as moving downward or towards the bottom of the image. Terms such as top, bottom, left, and right are used to simplify understanding of the present invention. According to other embodiments, such terms can be modified, changed, or substituted for any other direction without limiting the scope of the present embodiments. As an example, the reference samples 1062 and / or 1064 located to the left from the prediction block, i.e., with x<0, can be referred to as left reference samples. Assuming that the top left corner of the prediction block 1066 has position 0,0, y The reference samples 1062 and / or 1064 positioned to have a .intg.<0 may be referred to as top reference samples. The identified reference samples, left reference samples, and top reference samples may be referred to as corner reference samples. Thus, the reference samples beyond the extension W along the x-direction may be referred to as right reference samples, and the reference samples beyond the extension H of the prediction block 1066 may be referred to as bottom reference samples.

[0042] To indicate which reference sample to use for prediction, each line of reference samples is associated with a reference region index i. The nearest reference sample is given index i=0, the next line of extended reference samples i=1, etc. Using the notation in Figure 3, the reference samples above are

number

number

number

number

[0043] A video encoder according to an embodiment, such as video encoder 1000, can be configured to encode images of a video into coded data by block-based predictive coding, where block-based predictive coding includes intra-image prediction. In intra-image prediction, the video encoder can use multiple extended reference samples of an image to encode a predictive block of the image, where each extended reference sample of the multiple extended reference samples is separated from the predictive block by at least one nearest reference sample of the multiple reference samples that is directly adjacent to the predictive block. The video encoder can sequentially determine the availability or unavailability of each of the multiple nearest reference samples and replace the nearest reference sample determined to be unavailable with a substitution sample. The video encoder can use the substitution sample for intra-image prediction.

[0044] To determine availability or unavailability, the video encoder may sequentially check the samples according to the sequence and determine the replacement sample as a copy of the last extended reference sample determined to be available in the sequence, and / or may determine the replacement sample as a copy of the next extended reference sample determined to be available in the sequence.

[0045] The video encoder may further determine availability or unavailability sequentially according to the sequence, and determine the replacement sample based on a combination of extended reference samples that are determined to be available and for which the reference sample is determined to be unavailable, are placed in the sequence before the extended reference sample is determined to be available, and are placed in sequence after the reference sample is determined to be unavailable.

[0046] Alternatively, or in addition, the video encoder may be configured to use multiple closest reference samples and multiple extended reference samples of images immediately adjacent to the predictive block to encode the predictive block of the image for intra-image prediction, each extended reference sample of the multiple extended reference samples being separated from the predictive block by at least one closest reference sample of the multiple reference samples, and to determine availability or unavailability of each of the multiple extended reference samples. The video encoder may signal use of the multiple extended reference samples if a portion of the available extended reference samples of the multiple extended reference samples is equal to or greater than a predetermined threshold, and may skip signaling use of the multiple extended reference samples if a portion of the available extended reference samples of the multiple extended reference samples is below the predetermined threshold.

[0047] Thus, the video decoder 1008 or a respective decoder, such as a video decoder for regenerating a video stream, may be configured to decode an image encoded with the encoding data into video by block-based predictive decoding including intra-image prediction, where the intra-image prediction uses multiple extended reference samples of the image to encode a predictive block of the image, each extended reference sample of the multiple extended reference samples being separated from the predictive block by at least one nearest reference sample among the multiple reference samples directly adjacent to the predictive block, and sequentially determine the availability or unavailability of each of the multiple extended reference samples. The video decoder may replace the extended reference sample determined to be unavailable with a replacement sample and use the replacement sample for the intra-image prediction.

[0048] The video decoder may further be configured to determine availability or unavailability in order according to the sequence, and to determine the replacement sample as a copy of the last extended reference sample determined to be available in the sequence, and / or to determine the replacement sample as a copy of the next extended reference sample determined to be available in the sequence.

[0049] Further, the video decoder may be configured to determine availability or unavailability in order according to the sequence, and to determine the replacement sample based on a combination of extended reference samples where the reference sample is determined to be available and the reference sample is determined to be unavailable, and where the extended reference sample is placed in the sequence before the reference sample is determined to be available, and where the reference sample is determined to be unavailable.

[0050] Alternatively, or in addition, the video decoder may be configured to, for intra-image prediction, use a plurality of nearest reference samples and a plurality of extended reference samples of images immediately adjacent to the predictive block to decode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, determine availability or unavailability of each of the plurality of extended reference samples, notify received information indicating use of the plurality of extended reference samples if a portion of the available extended reference samples of the plurality of extended reference samples is equal to or greater than a predetermined threshold, indicating use of the plurality of extended reference samples, and skip use of the plurality of extended reference samples if there is no information, i.e., if a portion of the available extended reference samples of the plurality of extended reference samples is below the predetermined threshold.

[0051] If neighboring reference samples are unavailable, then, according to embodiments, extended reference sample substitution can be performed, e.g., the unavailable sample is replaced with the nearest available neighboring sample, a combination of the two nearest neighbors, or, if no neighboring samples are available, e.g., 2 bitdepth-1 For example, reference samples are unavailable if they lie outside the boundaries of a picture, slice, or tile, or if constrained intra prediction is used, which does not allow using samples from inter-picture prediction regions as reference for intra-picture prediction regions.

[0052] For example, if the current block to be predicted is at the left image boundary, the reference samples at the left and upper left corners are unavailable. In this case, the reference samples at the left and upper left corners are replaced by the first available top reference sample. This first available top reference sample is the first sample, i.e.,

number

number

[0053] When using constrained intra prediction, one or more neighboring blocks may not be available because they were coded using inter-picture prediction. For example, the H samples on the left

number

number

[0054] In one embodiment, the availability check process for each reference sample is performed sequentially, e.g., from bottom left to top right, or vice versa, and the first unavailable sample along this direction is replaced by the last available sample. If there is no previous available sample, the unavailable sample is replaced by the next available sample. In an embodiment starting from the bottom right, the W bottom right sample

number

number

number

[0055] If it is determined that most of the extended reference area samples are unavailable, then using the extended reference sample provides no advantage over using the closest reference sample. Thus, notification of the reference area index can be saved and the notification can be limited to blocks where at least half of the extended reference samples are available.

[0056] A video encoder according to an embodiment, such as video encoder 1000, may be configured to encode an image of a video into coded data by block-based predictive coding including intra-image prediction, wherein in the intra-image prediction, a plurality of extended reference samples of the image are used to encode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample among the plurality of reference samples directly adjacent to the predictive block, filtering at least a subset of the plurality of extended reference samples using a bilateral filter to obtain a plurality of filtered extended reference samples, and using the plurality of filtered extended reference samples for the intra-image prediction.

[0057] A video encoder according to this embodiment may be configured to combine a plurality of filtered extended reference samples with a plurality of unfiltered extended reference samples to obtain a plurality of combined reference values, and the video encoder is configured to use the plurality of combined reference values ​​for intra-image prediction.

[0058] Alternatively, or in addition, the video encoder may be configured to filter the plurality of extended reference samples using one of a 3-tap filter, a 5-tap filter, and a 7-tap filter.

[0059] The video encoder may further be configured to select an angular prediction mode to predict the prediction block, where the 3-tap, 5-tap, and 7-tap filters are configured as bilateral filters, and the video encoder may be configured to select one of the 3-tap, 5-tap, and 7-tap filters based on an angle used for the angular prediction, where the angle is positioned between the horizontal and vertical directions of the angular prediction modes, and / or the video decoder may be configured to select one of the 3-tap, 5-tap, and 7-tap filters based on a block size of the prediction block. As shown in FIG. 4f, the angle ε may represent the angle of the direction of the angular prediction used to predict the prediction block 1066 relative to the horizontal boundary 1072 and / or the vertical boundary 1074 of the prediction block 1066, measured toward the diagonal 1076 between the horizontal and vertical directions. That is, the angle of the angular prediction is at most 45°. The larger the angle ε, the greater the number of taps available for use in the filter. Alternatively, or in addition, the block size may define a basis or dependency for selecting the filter.

[0060] A corresponding video decoder may be configured to decode an image encoded with the encoding data into video by block-based predictive decoding including intra-image prediction, where in the intra-image prediction, use a plurality of extended reference samples of the image to decode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample among the plurality of reference samples directly adjacent to the predictive block, filter at least a subset of the plurality of extended reference samples using a bilateral filter to obtain a plurality of filtered extended reference samples, and use the plurality of filtered extended reference samples for the intra-image prediction.

[0061] The video decoder can be further configured to combine the plurality of filtered extended reference samples with the plurality of unfiltered extended reference samples to obtain a plurality of combined reference values, and the video decoder is configured to use the plurality of combined reference values ​​for intra-picture prediction.

[0062] Alternatively, or in addition, the video decoder may be configured to filter the multiple extended reference samples using one of a 3-tap filter, a 5-tap filter, and a 7-tap filter. As described for the encoder, the 3-tap filter, the 5-tap filter, and the 7-tap filter may be configured as bilateral filters, and the video decoder may be configured to predict the prediction block using an angular prediction mode and select one of the 3-tap filter, the 5-tap filter, and the 7-tap filter to be used based on an angle used for the angular prediction, the angle being positioned between the horizontal direction and the vertical direction of the angular prediction mode, and / or the video decoder may be configured to select one of the 3-tap filter, the 5-tap filter, and the 7-tap filter to be used based on a block size of the prediction block.

[0063] For example, instead of a bilateral filter, a 3-tap FIR filter can be used, which allows filtering only the nearest reference samples (even if no bilateral filter is used) and leaving the extended reference samples unfiltered.

[0064] When the sample area is large, discontinuities in the reference samples can occur, distorting the prediction. The state-of-the-art solution to this is to apply a linear smoothing filter to the reference samples. In the case of discontinuities, which can be detected by comparing them to a predefined threshold, stronger smoothing can be applied. This usually involves generating reference samples by interpolating between corner reference samples.

[0065] However, a linear smoothing filter can also remove edge structures that need to be preserved. Applying a bilateral filter to the extended reference samples according to the reference sample filtering embodiment can prevent undesired smoothing of sharp edges. Because bilateral filtering is more effective for large blocks and intra-prediction angles that deviate from the exact horizontal and exact vertical directions, the decision on whether to apply a filter and its length can depend on the block size and / or prediction mode. An example design can incorporate dependencies as shown in FIG. 4g, which shows a dependency for block sizes smaller than 64×64 and equal to or larger than 64×64 in terms of W×H, while FIG. 4h shows a different dependency for block sizes smaller than 64×64 and equal to or larger than 64×64 in terms of W×H. In the embodiment of FIG. 4g, the intra-picture prediction mode can be, for example, one of planar mode, DC mode, near-horizontal mode, near-vertical mode, or different angles in angular mode. In the embodiment of FIG. 4h, angles identified as more horizontal and more vertical can be further selected, for example, a larger value of angle ε shown in FIG. 4f when compared to near-horizontal or near-vertical. As can be seen, a larger block size can result in a larger number of taps to facilitate filtering of a larger amount of data, and furthermore, increasing ε can also result in an increase in taps. According to Figure 4h, an embodiment can apply a small 3-tap filter to the near-horizontal and near-vertical modes, increasing the filter length as the distance from the horizontal and vertical directions increases. While shown as depending on both the prediction mode and the block size, the choice of filter, or at least the number of taps, can instead depend only on one of the two and / or additional parameters.

[0066] In another embodiment of reference sample filtering, intra-picture prediction using filtered reference samples can be combined with unfiltered reference samples using position-dependent weighting, as described in connection with position-dependent prediction combination. In this case, the reference samples of prediction using filtered reference samples can use different reference sample filtering than in the case of uncombined prediction. For example, the filtering can be selected from a set of 3-tap, 5-tap, and 7-tap filters.

[0067] A video encoder according to an embodiment, such as video encoder 1000, may be configured to encode images of a video into coded data by block-based predictive coding including intra-image prediction, where in the intra-image prediction, a plurality of extended reference samples of the image are used to encode a predictive block of the image, where each extended reference sample of the plurality of extended reference samples is separated from the predictive block by at least one nearest reference sample among the plurality of reference samples directly adjacent to the predictive block, where the plurality of nearest reference samples are arranged along a first image direction of the predictive block and along a second image direction of the predictive block, where at least some of the nearest reference samples arranged along the second direction are mapped to extended reference samples arranged along the first direction such that the mapped reference samples exceed an extension of the predictive block along the first image direction, and where the mapped extended reference samples are used for prediction.

[0068] The video encoder can be further configured to map the portion of the nearest reference sample according to a prediction mode used to predict the predictive block. The video encoder can be configured to map the portion of the nearest reference sample according to a direction used in the prediction mode for predicting the predictive block.

[0069] A corresponding video decoder decodes an image encoded with the encoding data into video by block-based predictive decoding including intra-image prediction, wherein in the intra-image prediction, a plurality of extended reference samples of the image are used to decode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample among the plurality of reference samples immediately neighboring the predictive block, and the plurality of nearest reference samples are arranged along a first image direction of the predictive block and along a second image direction of the predictive block; The method may be configured to map at least some of the closest reference samples arranged along the second direction to extended reference samples arranged along the first direction such that the mapped reference samples exceed the extension of the prediction block along the first image direction, and to use the mapped extended reference samples for prediction.

[0070] The video decoder can be configured to map portions of the nearest reference samples according to a prediction mode used to predict the predictive block. The video decoder can be configured to map portions of the nearest reference samples according to a direction used in the prediction mode for predicting the predictive block.

[0071] In principle, all intra-picture predictions that use directly neighboring reference samples can be adapted to use extended reference samples. Three predictions have been adopted in the literature:

[0072] ·plane DC ·angle In the following, each prediction is described in detail for the extended reference samples to explain the embodiments of the present invention.

[0073] Planar prediction is a bilinear interpolation of W × H predicted samples from the boundaries shown in Figure 3. Since the right and bottom boundaries have not yet been reconstructed, the right boundary samples are obtained by subtracting the sample in the top right corner from the

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0074] DC prediction calculates the average of the reference sample value DC and all predicted samples in the W × H block.

number

number

number

number

number

[0075] According to an embodiment, a finer predicted angular granularity can be used so that the directions can point to non-integer sample positions. These can be generated by interpolation using the nearest integer sample positions. The interpolation can be performed using a simple bilinear filter or a more advanced filter, such as a cubic or Gaussian 4-tap filter.

[0076] For prediction angles between horizontal and vertical, e.g., the top left angle of the diagonal and all angles in between, a prediction is generated using reference samples from the left and top. To facilitate such generation, reference sample projection can be used. For all vertical angles between the top left diagonal and the vertical, the top reference sample can be displayed as the primary reference and the left sample as the secondary reference. For all horizontal angles between the top left horizontal line and the diagonal, the left reference sample can be displayed as the primary reference and the top reference sample as the secondary reference. To simplify the calculation, i.e., to avoid switching between the primary and secondary reference sample calculations, the secondary reference samples are projected along the prediction angle to extend the line of the primary reference sample. To the left of the top primary reference

number

number

[0077] 5a-5c illustrate embodiments relating to angle prediction using a definition of the angle parameter A given at 1 / 32 sample accuracy for the top-left diagonal prediction angle and two other vertical angles. The following embodiments assume 1 / 32 sample accuracy and a vertical prediction angle between vertical and the top-left diagonal, although any other value can be implemented. The 33 prediction angle ranges are defined by the angle parameter A ranging from 0 (vertical) to 32 (top-left diagonal) as shown in FIGS. 5a-5c.

number

[0078] Angle Parameter

number

number

number

number

number

number

number

[0079] Figures 7a to 7c show the height of the predicted block for the top-left horizontal prediction angles A = 32, 17, and 1.

number

number

[0080] In this embodiment, the vertical offset is rounded to use the nearest integer reference sample instead of the sub-sample position, which simplifies the calculation. However, it is also possible to project the reference sample at the interpolated sub-sample side.

[0081] Figure 8 shows the upper left corner of the diagonal.

number

number

number

[0082] FIG. 8 shows an example of a projection of the nearest left reference sample as a side reference next to the nearest top reference sample as a primary reference for a left-top horizontal direction with angle parameter A=17 according to an embodiment.

[0083] For an extended reference sample with reference region index i, the projection can be fitted as follows:

number

number

number

number

number

[0084] FIG. 9 shows an example of projection of an extended left reference sample as a secondary reference next to an extended top reference sample as a primary reference in the case of top-left diagonal prediction with angle parameter A=17 according to an embodiment.

[0085] When using reference samples from the left and top, the extended reference sample allows for combining the closest reference sample with the extended reference sample. This embodiment of the simple approach from Figure 9 allows for utilizing extended reference samples along the main direction, i.e., the primary reference, and the longer distance of the reference sample can be beneficial in cases of occlusion or edges on the closest reference sample. However, in the case of secondary references, the correlation between the predicted sample and the closest reference sample can be higher than between the extended reference sample and the predicted sample.

[0086] For example, the extended primary reference sample

number

number

number

[0087] Figure 10 shows an example of projection of the nearest left reference sample as a secondary reference next to the extended top reference sample as a primary reference in the case of left-top diagonal prediction with angle parameter A=17. According to the embodiment of Figure 10, the nearest reference sample is used as a source for generating the extended reference sample, that is, the nearest reference sample is mapped to the extended reference sample. Alternatively, the extended reference sample can also be mapped to the extended reference sample. See Figure 9.

[0088] A video encoder according to an embodiment, such as video encoder 1000, may be configured to encode images of a video into coded data by block-based predictive coding, including intra-image prediction, and for intra-image prediction, to use a plurality of nearest reference samples and a plurality of extended reference samples of images immediately neighboring the predictive block to encode a predictive block of the image, wherein each extended reference sample of the plurality of extended reference samples is separated from the predictive block by at least one nearest reference sample of the plurality of reference samples. The video decoder is configured to perform boundary filtering in a mode in which extended samples are not used, and to not use boundary filtering when extended samples are used, or to border filter at least a subset of the plurality of nearest reference samples and not use boundary filtering on the extended samples.

[0089] A corresponding video decoder is configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, and for the intra-image prediction, to use a plurality of nearest reference samples and a plurality of extended reference samples of images immediately neighboring the predictive block to decode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples. The video decoder is configured to perform boundary filtering in a mode in which the extended samples are not used, and when the extended samples are used, to either not use boundary filtering or to boundary filter at least a subset of the plurality of nearest reference samples and not use boundary filtering on the extended samples.

[0090] Because the extended reference samples are not directly adjacent to the predicted samples, discontinuities at a particular block boundary may not be as severe as if the nearest reference samples were used. Therefore, it is beneficial to perform predictive filtering according to: Do not perform boundary filtering operations when extended reference samples are used, or · Modify boundary smoothing by using the nearest reference sample instead of the extended reference sample.

[0091] Using extended prediction, predictions from the nearest reference sample and the extended reference sample can be combined to obtain a combined prediction. The literature describes a fixed combination of a prediction using the nearest reference sample and a prediction using the extended reference sample with a predetermined weight. In this case, both predictions use the same prediction mode for all reference regions, and signaling the mode also signals the prediction combination. While this reduces the signaling overhead for indicating the reference sample region, it also eliminates the flexibility to combine two different prediction modes with two different reference sample regions. Possible combinations to increase flexibility are described in detail below.

[0092] A video encoder according to an embodiment, such as video encoder 1000, may be configured to encode images of a video into coded data by block-based predictive coding including intra-image prediction, where the intra-image prediction involves determining a plurality of nearest reference samples and a plurality of extended reference samples of images immediately neighboring the predictive block to encode a predictive block of the image, where each extended reference sample of the plurality of extended reference samples is separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, determining a prediction of the predictive block using the extended reference samples, filtering the extended reference samples, and obtaining a plurality of filtered extended reference samples. The video encoder is configured to combine the prediction and the filtered extended reference samples to obtain a combined prediction of the predictive block.

[0093] The video encoder can be configured to combine prediction samples and enhanced reference samples that are located on the main or sub-diagonal of samples with respect to the prediction block.

[0094] The video encoder may be configured to combine the prediction samples and the enhanced reference samples based on the following decision rules:

number

number

[0095] The normalization factor can be determined based on a determination rule.

number

[0096] The video encoder may be configured to use a combination of extended corner reference samples of a predictive block and extended reference samples (r(-1-i,-1-i)) located in corner regions of the reference samples.

[0097] The video encoder may be configured to obtain the combined prediction based on the following decision rules:

number

number

number

[0098] The video encoder may be configured to obtain the prediction p(x,y) based on intra-picture prediction.

[0099] The video encoder can be configured to use only planar prediction as intra-picture prediction.

[0100] The video encoder can be configured to determine, for each coded video block, a parameter set that identifies a combination of a prediction and a filtered extended reference sample. The video encoder can be configured to determine, using a lookup table that includes a set of different block sizes of prediction blocks, a parameter set that identifies a combination of a prediction and a filtered extended reference sample.

[0101] A corresponding video decoder may be configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, and in the intra-image prediction, to decode a predictive block of the image, determine a plurality of nearest reference samples and a plurality of extended reference samples of images immediately adjacent to the predictive block, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, determine a prediction of the predictive block using the extended reference samples, filter the extended reference samples to obtain a plurality of filtered extended reference samples, and combine the prediction and the filtered extended reference samples to obtain a combined prediction of the predictive block.

[0102] The video decoder may be configured to combine prediction samples and enhanced reference samples that are located on the main or sub-diagonal of samples for the prediction block.

[0103] The video encoder may be configured to combine the prediction samples and the enhanced reference samples based on the following decision rules:

number

number

[0104] The normalization factor can be determined based on a determination rule.

number

[0105] The video decoder may be configured to use a combination of extended corner reference samples of the predictive block and extended reference samples located in the corner regions of the reference samples (r(-1-i,-1-i)).

[0106] The video decoder may be configured to obtain the combined prediction based on the following decision rules:

number

number

number

[0107] The video decoder may be configured to obtain the prediction p(x,y) based on intra-picture prediction. The video decoder may, for example, use only planar prediction as intra-picture prediction.

[0108] The video encoder may be configured to determine, for each decoded video block, a parameter set that identifies a combination of prediction and filtered enhanced reference samples.

[0109] The video decoder may be configured to determine a parameter set that identifies a combination of prediction and filtered enhanced reference samples using a lookup table that includes a set of different block sizes of prediction blocks.

[0110] When filtering reference samples, predictions using filtered samples can be combined with unfiltered reference samples based on the position of each sample to obtain position-dependent prediction combinations.

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0111] This known combination of nearest reference samples can be improved in accordance with embodiments by using extended reference samples for the filtered and unfiltered reference samples. The following illustrates a combination of predictions using extended reference samples in accordance with embodiments, where the nearest corner sample r(-1,-1) is used as the extended corner sample r(-1,-1) as follows:

number

number

number

number

number

number

number

[0112] Alternatively, or in addition, different enhanced prediction modes can be combined to combine predictions with samples.

[0113] A video encoder according to an embodiment, such as video encoder 1000, is configured to encode an image of a video into coded data by block-based predictive coding including intra-image prediction, and for the intra-image prediction, to encode a predictive block of the image using a plurality of nearest reference samples and / or a plurality of extended reference samples of images immediately neighboring the predictive block, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, determine a first prediction of the predictive block using a first prediction mode of a set of prediction modes, the first set of prediction modes including a prediction mode that uses the plurality of nearest reference samples in the absence of an extended reference sample, determine a second prediction of the predictive block using a second prediction mode of a second set of prediction modes, and a second set of prediction modes including a subset of the prediction modes of the first set associated with the plurality of extended reference samples. i (x, y)) and weighted (w0;w i ) to obtain a combined prediction (p(x, y)) as a prediction for a predictive block of encoded data.

[0114] The video encoder may be configured to use the first prediction and the second prediction according to a predetermined combination that is a subset of possible combinations of valid first prediction modes and valid second prediction modes.

[0115] The video encoder can be configured to signal either the first prediction mode or the second prediction mode without signaling the other prediction mode. For example, the first mode can be derived from the parameters i based on additional implicit information, such as a particular prediction mode that can only be used in association with a particular index i or index m.

[0116] As an example of such implicit information, the video encoder may be configured to exclusively use a planar prediction mode as one of the first prediction mode and the second prediction mode.

[0117] The video encoder may be configured to adapt the first weight applied to the first prediction of the joint prediction and the second weight applied to the second prediction of the joint prediction based on a block size of the prediction block, and / or to adapt the first weight based on the first prediction mode or the second weight based on the second prediction mode.

[0118] The video encoder may be configured to adapt a first weight applied to a first prediction of the joint prediction and a second weight applied to a second prediction of the joint prediction based on a position and / or distance within the prediction block.

[0119] The video encoder may be configured to adapt the first weight and the second weight based on the following decision rule:

number

[0120] A corresponding video decoder decodes an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, and for the intra-image prediction, uses a plurality of nearest reference samples and / or a plurality of extended reference samples of images immediately adjacent to the predictive block to decode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples; The video decoder may be configured to determine a first prediction of the predictive block using a first prediction mode of a set of prediction modes, the first set of prediction modes including a prediction mode that uses multiple nearest reference samples when there are no extended reference samples, and to determine a second prediction of the predictive block using a second prediction mode of a second set of prediction modes, the second set of prediction modes including a subset of the prediction modes of the first set, the subset being associated with multiple extended reference samples. i (x, y)) and weighted (w0;w i ) to obtain a combined prediction (p(x, y)) as a prediction for a predictive block of encoded data.

[0121] The video decoder may be configured to use the first prediction and the second prediction according to a predetermined combination that is a subset of possible combinations of valid first prediction modes and valid second prediction modes, which allows for low-load notification.

[0122] The video decoder may be configured to receive a signal indicating the second prediction mode without receiving a signal indicating the first prediction mode, and to derive the first prediction mode from the second prediction mode or the parameter information (i).

[0123] The video decoder may be configured to exclusively use the planar prediction mode as one of the first prediction mode and the second prediction mode.

[0124] The video decoder may be configured to adapt a first weight applied to a first prediction of the joint prediction and a second weight applied to a second prediction of the joint prediction based on a block size of the prediction block, and / or to adapt the first weight based on a first prediction mode or the second weight based on a second prediction mode.

[0125] The video decoder may be configured to adapt a first weight applied to a first prediction of the joint prediction and a second weight applied to a second prediction of the joint prediction based on a position and / or distance within the prediction block.

[0126] The video decoder may be configured to adapt the first weight and the second weight based on the following decision rule:

number

[0127] When using different reference sample regions, the prediction using the nearest reference sample mode is as follows:

number

number

number

[0128] To relax this strict restriction, embodiments define that only certain combinations of modes are allowed. This requires the signaling of a second mode, but limits the number of modes signaled compared to the first mode. Details of prediction mode signaling are provided in the context of mode and reference signaling. One promising combination according to embodiments is to use only planar as the second mode, thereby making additional signaling of intra modes obsolete. For example, any intra mode as the first part of the weighted sum can be used.

number

[0129] To adapt the weights to the prediction size and mode,

number

number

number

[0130] A video encoder according to an embodiment, such as video encoder 1000, may be configured to encode an image of a video into coded data by block-based predictive coding including intra-image prediction, and for intra-image prediction, to encode a predictive block of the image using a plurality of nearest reference samples and / or a plurality of extended reference samples of images directly adjacent to the predictive block, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, and to use a prediction mode that is one of a first set of prediction modes for predicting the predictive block using the nearest reference samples, or to use a prediction mode that is one of a second set of prediction modes for predicting the predictive block using the extended reference samples, wherein the second set of prediction modes is a subset of the first set of prediction modes, and the second subset may be determined by the encoder, and the subset may also include a match between both sets. The video encoder may be configured to report mode information (m) indicating a prediction mode used to predict a prediction block, and then, if the prediction mode is included in a second set of prediction modes, report parameter information (i) indicating a subset of extended reference samples used for the prediction mode, and skip reporting the parameter information if the used prediction mode is not included in the second set of prediction modes, thereby enabling a conclusion that parameter i has a predetermined value, such as 0.

[0131] The video encoder may be configured to skip reporting the parameter information if the mode information indicates a DC mode or a planar mode.

[0132] A corresponding video decoder may be configured to decode an image encoded with the coding data by block-based predictive coding including intra-image prediction into video, and for intra-image prediction, to decode a predictive block of the image using a plurality of nearest reference samples and / or a plurality of extended reference samples of images immediately adjacent to the predictive block, where each extended reference sample of the plurality of extended reference samples is separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, and to use a prediction mode that is one of a first set of prediction modes for predicting the predictive block using the nearest reference samples, or a prediction mode that is one of a second set of prediction modes for predicting the predictive block using the extended reference samples, where the second set of prediction modes is a subset of the first set of prediction modes and may be determined and / or signaled by the encoder. Being a subset may include a match between both sets. The video encoder may be configured to receive mode information (m) indicating a prediction mode used to predict the prediction block, then receive parameter information (i) indicating a subset of extended reference samples used for the prediction mode, thereby indicating that the prediction mode is included in a second set of prediction modes, and if no parameter information is received, determine that the used prediction mode is not included in the second set of prediction modes and determine the use of the nearest reference sample for prediction.

[0133] The video decoder may be configured to determine the mode information as indicating the use of DC mode or planar mode when no parameter information is received.

[0134] If the set of allowed intra prediction modes of the extended reference sample is restricted, i.e., a subset of the allowed intra prediction modes of the nearest reference sample, e.g., if the subset is called the restricted prediction modes of the extended reference sample, there are two ways to signal the mode m and index i: 1. Signal mode m before index i, so that the index signaling can depend on mode m as follows: a. If mode m is not in the set of allowed modes of the extended reference sample, the notification of index i is skipped and predicted mode m is applied to the nearest reference sample (i=0). b. Otherwise, i is signaled and prediction mode m is applied to the reference sample indicated by i.

[0135] 2. Signal index i before mode m, so that the mode signal can depend on index i as follows: If ai indicates that an extended reference sample (i>0) is used, the set of modes m that can be signaled is the same as the restricted one.

[0136] b. Otherwise (i=0), the set of advertised allowed modes m is equal to the set of unrestricted modes.

[0137] For example, if mode m is signaled using Most Probable Mode (MPM) coding using an index into an MPM list, modes that are not in the set of allowed modes will not be included in the MPM list.

[0138] That is, according to a second option that may be implemented alternatively or additionally, a video encoder according to an embodiment, such as video encoder 1000, encodes an image of a video into coded data by block-based predictive coding including intra-image prediction, and for the intra-image prediction, uses a plurality of reference samples and a plurality of extended reference samples to encode a predictive block of the image, the plurality of reference samples including nearest reference samples of images directly adjacent to the predictive block, each extended reference sample being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, and a first set of prediction modes for predicting the predictive block using the nearest reference sample. a prediction mode that is one of a first set of prediction modes, or a prediction mode that is one of a second set of prediction modes for predicting the prediction block using extended reference samples, the second set of prediction modes being a subset of the first set of prediction modes, signaling parameter information (i) indicating a subset of the plurality of reference samples to be used for the prediction mode, the subset of the plurality of reference samples including only the nearest reference samples or the extended reference samples, and then signaling mode information (m) indicating a prediction mode to be used for predicting the prediction block, the mode information indicating a prediction mode from the subset of modes, the subset being configured to be restricted to the set of prediction modes allowed according to the parameter information (i).

[0139] For both options, the video decoder is adapted such that, in addition to the nearest reference sample, extended reference samples of the modes included in the second set of prediction modes are used.

[0140] Further, the video encoder can be adapted such that a first set of prediction modes describes prediction modes that are capable of being used with the nearest reference sample, and a second set of prediction modes describes prediction modes of the first set of prediction modes that are also capable of being used by the extended reference sample.

[0141] The range of values ​​of the parameter information, i.e., the domain of values ​​that can be represented by the area index i, can cover the use of only the nearest reference value and the use of different subsets of extended reference values. As described in connection with Figure 11, i can represent the use of only the nearest reference value (i = 0), a specific set of extended reference values, or a combination of the set of extended reference values ​​(e.g., line and / or column or distance).

[0142] According to an embodiment, different parts of the extended reference samples contain different distances to the prediction block.

[0143] The video encoder may be configured to set the parameter information to one of a predetermined number of values, which indicates the number and distance of reference samples to be used for the prediction mode.

[0144] The video encoder may be configured to determine the first set of prediction modes and / or the second set of prediction modes based on a most probable mode coding.

[0145] a second optional corresponding decoder is configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, and for the intra-image prediction, to decode a predictive block of the image, use a plurality of reference samples and a plurality of extended reference samples including a nearest reference sample of an image immediately adjacent to the predictive block, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, use a prediction mode that is one of a first set of prediction modes for predicting the predictive block using the nearest reference sample or use a prediction mode that is one of a second set of prediction modes for predicting the predictive block using the extended reference samples, the second set of prediction modes being a subset of the first set of prediction modes, receive parameter information (i) indicating a subset of the plurality of reference samples used for the prediction mode, the subset of the plurality of reference samples including only the nearest reference sample or at least one extended reference sample, and then receive mode information (m) indicating a prediction mode to be used to predict the predictive block, the mode information indicating a prediction mode from the subset of modes, the subset being restricted to a set of prediction modes allowed according to the parameter information (i).

[0146] A decoder according to the first and / or second option can be adapted such that, in addition to the nearest reference sample, extended reference samples of modes included in the second set of prediction modes are used, for example by combining predictions with samples and / or combinations of predictions.

[0147] The first set of prediction modes may describe prediction modes that are allowed for use with the nearest reference sample, and the second set of prediction modes may describe prediction modes in the first set of prediction modes that are also allowed for use with the extended reference sample.

[0148] As described for the encoder, the range of values ​​for the parameter information covers the use of only the nearest reference value and the use of a different subset of the extended reference values.

[0149] Different portions of the extended reference samples may contain different distances to the prediction block.

[0150] The video decoder may be configured to set the parameter information to one of a predetermined number of values, which indicates the number and distance of reference samples to be used for the prediction mode.

[0151] The video decoder may be configured to determine a first set of prediction modes and / or a second set of prediction modes based on Most Probable Mode (MPM) coding, i.e., generate respective lists, which may be adapted to include only modes that are allowed to be used for the respective sample.

[0152] Alternatively, or in addition, embodiments relating to combining predictions from the nearest reference sample and the extended reference sample can be implemented.

[0153] A video encoder according to an embodiment, such as video encoder 1000, may be configured to encode an image of a video into encoded data by block-based predictive coding including intra-image prediction, and for the intra-image prediction, use a plurality of nearest reference samples and / or a plurality of extended reference samples of images immediately neighboring the predictive block to encode a predictive block of the image, wherein each extended reference sample of the plurality of extended reference samples is separated from the predictive block by at least one nearest reference sample of the plurality of reference samples, determine a first prediction of the predictive block using a first prediction mode of a set of prediction modes, the first set of prediction modes including a prediction mode that uses the plurality of nearest reference samples in the absence of an extended reference sample, determine a second prediction of the predictive block using a second prediction mode of a second set of prediction modes, the second set of prediction modes including a subset of the prediction modes of the first set associated with the plurality of extended reference samples, and combine the first prediction and the second prediction to obtain a combined prediction as a prediction of the predictive block in the encoded data.

[0154] The predictive block may be a first predictive block, and the video encoder may be configured to predict a second predictive block of the video using a plurality of nearest reference samples associated with the second predictive block in the absence of a plurality of extended reference samples associated with the second predictive block. The video encoder may be configured to signal combination information, such as a bipred flag, indicating that prediction in the encoded data is based on a combination of predictions or based on a prediction using a plurality of extended reference samples in the absence of a plurality of nearest reference samples.

[0155] The video encoder may be configured to use the first prediction mode as the predetermined prediction mode.

[0156] The video encoder may be configured to select the first prediction mode as the same mode as the second prediction mode and to use the nearest reference sample in the absence of an extended reference sample, or to use the first prediction mode as a pre-set prediction mode, such as a planar prediction mode.

[0157] A corresponding video decoder may be configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, and for the intra-image prediction, to decode a predictive block of the image using multiple nearest reference samples and / or multiple extended reference samples of images immediately neighboring the predictive block, where each extended reference sample of the multiple extended reference samples is separated from the predictive block by at least one nearest reference sample of the multiple reference samples, determine a first prediction of the predictive block using a first prediction mode of a set of prediction modes, the first set of prediction modes including a prediction mode that uses the multiple nearest reference samples in the absence of an extended reference sample, and determine a second prediction of the predictive block using a second prediction mode of a second set of prediction modes, the second set of prediction modes including a subset of the prediction modes of the first set, the subset being associated with the multiple extended reference samples. The video decoder may be configured to combine the first prediction and the second prediction to obtain a combined prediction as a prediction of the predictive block in the encoded data.

[0158] The predictive block may be a first predictive block, and the video decoder may be configured to predict a second predictive block of the video using a plurality of nearest reference samples associated with the second predictive block in the absence of a plurality of extended reference samples associated with the second predictive block. The video decoder may be further configured to receive combination information indicating that a prediction in the encoded data is based on a combination of predictions or a prediction using a plurality of extended reference samples in the absence of a plurality of nearest reference samples, and to decode the encoded data accordingly.

[0159] The video decoder may be configured to use the first prediction mode as the predetermined prediction mode.

[0160] The video decoder can be configured to select the first prediction mode as the same mode as the second prediction mode and to use the nearest reference sample in the absence of an extended reference sample, or to use the first prediction mode as a preset prediction mode, such as a planar mode.

[0161] If the reference region index i indicates the use of an extended reference sample (i>0), a respective piece of information, such as a binary information or flag, hereafter referred to as a biprediction flag (bipred flag), is used to signal whether the prediction using the extended reference sample is combined with the prediction using the nearest reference sample.

[0162] If the bipred flag indicates a combination of predictions from the nearest (i=0) and extended reference samples (i>0), then the extended reference sample m i The mode of m is signaled before or after the reference region index i, as described here. The mode of the nearest reference sample prediction m is fixed, e.g., always set to a specific mode such as planar, or set to the same mode as the extended reference sample.

[0163] As described below, a video encoder according to an embodiment encodes an image of a video into coded data by block-based predictive coding including intra-image prediction, wherein in the intra-image prediction, a plurality of extended reference samples of the image are used to encode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample among the plurality of reference samples directly adjacent to the predictive block, and the plurality of extended reference samples are used according to a predetermined set of the plurality of extended reference samples, i.e., a list of area indices can be generated and / or used, and the list can be configured to indicate a particular subset of reference samples, such as only the nearest reference sample, at least one distance of the extended reference samples, and / or a combination of different distances of the extended reference samples.

[0164] The video encoder may be configured to determine multiple predetermined sets of extended reference samples such that multiple sets in the sets differ from each other by the number or combination of lines and / or rows of samples of the image used as reference samples.

[0165] The video encoder may be configured to determine a predetermined set of multiple enhanced reference samples based on a block size of the predictive block and / or a prediction mode used to predict the predictive block.

[0166] The video encoder may be configured to determine a set of multiple extended reference samples for which the block size of the predictive block is at least a predetermined threshold, and to skip reporting the set of multiple extended reference samples if the block size is below the predetermined threshold.

[0167] The predetermined threshold may be a predetermined number of samples along the width or height of the prediction block and / or a predetermined aspect ratio of the prediction block along the width and height.

[0168] The predetermined number of samples may be any number, but is preferably 8. Alternatively, or in addition, the aspect ratio may be greater than 1 / 4 and less than 4, perhaps based on the number of 8 samples that define a quotient of the aspect ratio.

[0169] The video encoder may be configured to predict a predictive block as a first predictive block using a plurality of extended reference samples and to predict a second predictive block (which may be part of the same or a different image) without using the extended reference samples, and the video encoder may be configured to signal a predetermined set of the plurality of extended reference samples associated with the first predictive block and not signal a predetermined set of the extended reference samples associated with the second predictive block. The predetermined set may be indicated, for example, by a reference region index i.

[0170] The video encoder may be configured to signal, for each predictive block, information indicating one of a particular set of extension reference samples and the use of the nearest reference sample only before information indicating the intra-picture prediction mode.

[0171] The video encoder may be configured to signal information indicating intra-picture prediction, thereby indicating a prediction mode according to an indicated particular number of sets of extended reference samples, or according to an indicated use of only the nearest reference sample.

[0172] The corresponding video decoder is configured to decode an image encoded by encoding data into video by block-based predictive decoding including intra-image prediction, and in the intra-image prediction, use a plurality of extended reference samples of the image to decode a predictive block of the image, each extended reference sample of the plurality of extended reference samples being separated from the predictive block by at least one nearest reference sample of the plurality of reference samples directly adjacent to the predictive block, and use the plurality of extended reference samples according to a predetermined set of the plurality of extended reference samples.

[0173] The video decoder may be configured to determine multiple predetermined sets of extended reference samples such that multiple sets in the sets differ from each other by the number or combination of rows and / or rows of samples of the image used as reference samples.

[0174] The video decoder may be configured to determine a predetermined set of multiple enhanced reference samples based on a block size of the predictive block and / or a prediction mode used to predict the predictive block.

[0175] The video encoder may be configured to determine a set of multiple extended reference samples for which the block size of the predictive block is at least a predetermined threshold, and to skip using the set of multiple extended reference samples if the block size is below the predetermined threshold.

[0176] For example, if the top-left sample of the current block at position (x0,y0) is in the first row in the coding tree block (CTB), then the set notification (parameter information i or a similar syntax can be used) can be skipped. The CTB can be seen as the fundamental processing unit into which the image is divided and which is the root of further block subdivisions.

[0177] This can be done by checking whether the vertical y coordinate y0 is not a multiple of the CTB size. For example, if the CTB size is 64 luma samples, then in the first CTB row, the above blocks in the CTB have y0=0, in the second CTB row, they have y0=64, etc. Therefore, if all blocks in the CTB are at the upper limit of this CTB (which can be checked, for example, by the following modulo operation: y0%CtbSizeY==0), then the notification of parameter i or a similar syntax can be skipped.

[0178] Thus, the predetermined threshold may be a predetermined number of samples along the width or height of the prediction block and / or a predetermined aspect ratio of the prediction block along the width and height.

[0179] Thus, the predetermined number of samples may be 8, and / or the aspect ratio may need to be greater than 1 / 4 and less than 4.

[0180] The video decoder may be configured to predict a predictive block as a first predictive block using a plurality of extended reference samples and predict a second predictive block without using the extended reference samples, and the video decoder is configured to receive information indicating a predetermined set of a plurality of extended reference samples associated with the first predictive block, for example, using an area index i, and to determine the predetermined set of extended reference samples associated with the second predictive block in the absence of a respective signal.

[0181] The video decoder can be configured to receive, for each predictive block, information indicating one of a particular set of extension reference samples and the use of the nearest reference sample only before the information indicating the intra-picture prediction mode.

[0182] The video decoder can thereby be configured to receive information indicating intra-picture prediction to indicate a prediction mode according to an indicated particular number of sets of extended reference samples or according to an indicated use of only the nearest reference sample.

[0183] Because the reference area index can change from block to block, the index can be transmitted in the bitstream for all applicable prediction blocks. An embodiment relates to extended reference sample area signaling to enable a decoder to use the correct sample. This area can correspond to a parameter i, which indicates the extended reference sample or the nearest reference sample. To trade off signaling and the use of more distant extended reference samples, a predetermined set of extended reference sample lines can be used. For example, only two additional extended reference lines, lines with indexes 1 and 3, can be used. In accordance with an embodiment, the set can be I={0,1,3}, where |I|=3. This allows the reference sample area to be extended up to three rows, but only two signals are required. The set can also be extended to four rows, e.g., I={0,1,3,4}, and only the first MaxNumRefAreaIdx element is used, but MaxNumRefAreaIdx can be fixed or signaled at the sequence, picture, or slice level. The index n into set I can be signaled in the bitstream using entropy coding and truncated single codes, as shown in the table shown in Figure 11, which shows examples of truncated single codes for specific reference area sets with MaxNumRefAreaIdx equal to 3 and 4.

[0184] To account for different spatial characteristics of different block sizes, the set of reference sample lines can also depend on the prediction block size and / or intra-prediction mode. In another embodiment, the set includes only one additional row with index 2 for small blocks and only two rows for large blocks with indexes 1 and 3. An embodiment of the set selection depending on the intra-prediction mode is to have different sets of reference sample lines for prediction directions between horizontal and vertical.

[0185] The signaling of reference region indices can also be limited to larger block sizes by selecting an empty set for small block sizes that do not require signaling. This can essentially disable the use of extended reference samples for small block sizes. In one embodiment, the use of extended reference samples can be limited to blocks whose width W and height H are both equal to or greater than 8. In addition, blocks whose one side is less than or equal to one-quarter of the other, such as 32×8 blocks or 8×32 blocks based on the aforementioned symmetry, can also be excluded from the use of extended reference samples. Figures 12a and 12b illustrate this embodiment, in which some block sizes are already excluded for intra prediction (Figure 12a), as indicated by reference numeral 1102, and the blocks are shaded accordingly. In this embodiment, it is assumed that intra and inter prediction slices allow for different block size combinations. Sample 1104 and the correspondingly shaded blocks are not allowed for extended reference samples, i.e., i>0. 12a and 12b show examples of restrictions on the extended reference samples of intra-predicted slices and inter-predicted slices. As can be seen, the allowance or restriction of blocks depending on the block size, especially the aspect ratio of the blocks, can be symmetric with respect to the quotients W / H and H / W.

[0186] If there are other intra prediction modes that do not use extended reference samples, the reference area index is signaled only if a prediction mode that uses extended reference samples is signaled. Another method is to signal the reference area index before all other intra mode information is signaled. If the reference area index i indicates an extended reference sample (i>0), signaling of mode information that does not use extended reference samples (e.g., template matching, trained predictors, etc.) can be skipped. This can also skip signaling information of certain transforms that are not applied to the prediction residual of predictions that use extended reference samples.

[0187] In the following, reference is made to embodiments that refer to parallel encoding considerations.

[0188] A video encoder according to an embodiment may be configured to encode a plurality of predictive blocks into coded data by block-based predictive coding including intra-image prediction, and to use a plurality of extended reference samples of an image to encode predictive blocks of the plurality of predictive blocks for the intra-image prediction, wherein each extended reference sample of the plurality of extended reference samples is separated from the predictive block by at least one nearest reference sample of the plurality of reference samples directly adjacent to the predictive block, and the video encoder may be configured to determine the extended reference sample to be at least a partial part of an adjacent predictive block of the plurality of predictive blocks, determine that the adjacent predictive block has not yet been predicted, and notify information indicating the extended predictive sample associated with the predictive block and located as an unavailable sample in the adjacent predictive block.

[0189] A video encoder may be configured to encode an image by parallel coding lines of blocks according to a wavefront approach, predict a predicted block based on angle prediction, and determine extended reference samples used to predict the predicted block so that the extended reference samples are located in an already predicted block of the image. According to the wavefront approach, the encoding or decoding of a second line may follow the decoding of a first line separated by one block, for example. For example, in a 45° vertical angle mode, starting from the second line, a maximum of one block from the top right may be decoded and coded.

[0190] The video encoder can be configured to inform neighboring prediction blocks of extended prediction samples associated with the prediction block, which are arranged differently as unavailable or available samples at a sequence level, i.e., a sequence of images, at a picture level, or at a slice level, where a slice is a portion of an image.

[0191] The video encoder may be configured to signal information associated with the predictive block and indicating parallel coding of the image together with information indicating extended predictive samples located in adjacent predictive blocks as unavailable samples.

[0192] A corresponding decoder can be configured to decode images encoded with the encoding data into video by block-based predictive decoding including intra-image prediction, where for each image, a plurality of predictive blocks are decoded, and for intra-image prediction, use a plurality of extended reference samples of the image to decode a predictive block of the plurality of predictive blocks, where each extended reference sample of the plurality of extended reference samples is separated from the predictive block by at least one nearest reference sample of the plurality of reference samples directly adjacent to the predictive block, and the video decoder is configured to determine the extended reference sample to be at least a partial part of a neighboring predictive block of the plurality of predictive blocks, determine that the neighboring predictive block has not yet been predicted, and receive information indicating the extended predictive sample associated with the predictive block and located as an unavailable sample in the neighboring predictive block.

[0193] The video decoder can be configured to decode an image by parallel decoding lines of blocks according to a wavefront approach and predict a prediction block based on angular prediction, and the video decoder is configured to determine extended reference samples used to predict the prediction block to be placed in already predicted blocks of the image.

[0194] The video decoder can be configured to receive information associated with a predictive block and indicating extended prediction samples that are atypically arranged in neighboring predictive blocks as unavailable samples or available samples at the sequence level, picture level, or slice level.

[0195] The video decoder may be configured to receive information associated with the predictive block and indicating parallel decoding of the image together with information indicating extended predictive samples located in adjacent predictive blocks as unavailable samples.

[0196] In angular intra prediction, reference samples are copied into the current prediction area along a specified direction. If this direction points to the upper right, the required reference sample area also shifts to the right as the distance to the prediction area boundary increases. Figure 13 shows an embodiment of vertical angular prediction at a 45-degree angle for a W×H prediction block. It can be seen that the closest reference sample area (blue) extends H samples to the upper right of the current W×H block. When extending the prediction to a more distant reference sample area, the extended reference samples (green) extend H+1, H+2,... samples to the upper right of the current W×H block.

[0197] For example, when a square coding tree unit (CTU) is used as the basic processing unit, the maximum intra-prediction block size may be equal to the maximum block size, i.e., the CTU block size N×N. Figures 14a and 14b show an embodiment from Figure 13 in which the W×H prediction block is equal to the CTU block size of N×N. Correspondingly, the extended reference samples 10641 and / or 10642 span the upper right CTU (CTU2) and allow at least one sample to reach the next CTU (CTU3). That is, some of the extended reference samples 10641 and / or 10642 of the 45° vertical angle prediction may be placed in the unprocessed CTU3. When CTU3 and CTU5 are processed in parallel, these samples may be unavailable.

[0198] 14a and 14b show examples of the nearest reference samples 1062 and extended reference samples 1064 required for diagonal vertical intra prediction.

[0199] In Figure 14b, two CTU lines can be seen. If two lines are coded in parallel using a wavefront-like approach, once CTU2 is coded and coding of CTU3 begins, coding of the second CTU line can start with CTU5. In this case, some reference samples 1104 are located inside CTU3, so 45 degree vertical angle prediction with extended reference samples cannot be used for CTU5. The following approach can solve this problem: 1. As described in connection with this embodiment, mark the extended references in the next basic processing domain as unavailable so that they can be treated like other unavailable reference samples, for example at picture, slice or tile boundaries or when constrained intra prediction is used which does not allow using samples from inter-picture prediction domains as references for intra-picture prediction domains.

[0200] 2. Allow the use of extended reference samples only for regional and intra-picture prediction modes, so that the reference samples are not extended to multiple elementary processing units to the top right of the current region.

[0201] Both approaches can be enabled by a high-level flag at the sequence, picture, or slice level. In this way, the encoder can inform the decoder to apply the constraint if the encoder's parallel processing scheme requires it, or if not, the encoder can inform the decoder that the constraint is not applied.

[0202] Another approach is to combine the signaling of both limitations with the signaling of the parallel coding scheme, e.g., if wavefront parallelism is enabled and signaled in the bitstream, the extended reference sample limitation also applies.

[0203] In the following, some advantageous embodiments are described.

[0204] 1. In one embodiment, intra prediction uses two additional reference lines with reference sample area indices i=1 and i=3 (see FIG. 3). These extended reference sample lines allow only angular prediction as described herein. The signaling of the reference sample area index i is performed as described in connection with FIG. 11 for MaxNumRefAreaIdx=3 and is signaled before signaling the intra prediction mode. If the reference area index is not equal to 0, i.e., if extended reference lines are used, DC and intra prediction modes are not used. To avoid unnecessary signaling, DC and planar modes are excluded from the signaling of intra prediction modes. The intra prediction mode can be coded using an index into a list of most probable modes (MPM). The MPM list contains a fixed number of candidate prediction modes derived from neighboring blocks. If the neighboring block is coded using the nearest reference line (i=0), DC or planar modes can be used, so the derivation process of the MPM list is modified to exclude DC and planar modes. If removed, DC and planar modes can be replaced with horizontal, vertical, and bottom-left diagonal modes to fill the list (see Figures 4a through 4e). This is done in a way that avoids redundancy. For example, if the first candidate mode derived from a neighboring block is vertical and the second candidate mode is DC, DC mode is replaced with horizontal mode instead of vertical mode because it is not already in the list. Another way to prevent unnecessary reporting of DC and planar modes when i>0 is to signal the reference region index i after the prediction mode and coordinate the reporting of i with the prediction mode. If the intra-prediction mode is equal to DC or planar, index i is not signaled. However, if the intra-prediction mode is signaled using an index into the MPM list, this method introduces a parsing dependency, so the previously mentioned method is preferred over this method. In order to be able to analyze i, the intra-prediction mode needs to be reconstructed, which requires the derivation of the MPM list. On the other hand, the derivation of the MPM list references the prediction mode of the neighboring block, which needs to be reconstructed before parsing i.This undesirable parsing dependency is resolved by signaling i before the prediction mode and modifying the derivation of the MPM list accordingly, as described above. The video encoder can be configured to determine a list of multiple most likely prediction modes based on the use of multiple nearest reference samples or the use of multiple extended reference samples for the prediction mode, where the video encoder is configured to replace prediction modes that are restricted with respect to the used reference samples by modes that are allowed in the prediction mode. A corresponding video decoder can be configured to determine a list of most likely prediction modes based on the use of multiple nearest reference samples or the use of multiple extended reference samples for the prediction mode, where the video decoder is configured to replace prediction modes that are restricted with respect to the used reference samples by modes that are allowed in the prediction mode.

[0205] 2. In another embodiment, intra-prediction of extended reference samples is further restricted to be applied only to luma samples. That is, the video encoder can be configured to apply prediction using extended reference samples to images that include only luma information. Accordingly, the video decoder can be configured to apply prediction using extended reference samples to images that include only luma information.

[0206] 3. In another embodiment, extended reference samples (i>0) (see W+H in FIG. 3 ) that exceed the width and height of the nearest reference sample (i=0) are not generated by using already reconstructed samples (if available), but are generated by padding from the last sample, such as r(23,-1-i) for the top-right sample in FIG. 3 and r(-1-i,23) for the bottom-left sample. This reduces memory accesses for extended reference sample lines. That is, a video encoder can be configured to generate extended reference samples that exceed the width and / or height of the nearest reference sample along the first and second image directions by padding from the nearest extended reference sample. Accordingly, a video decoder can be configured to generate extended reference samples that exceed the width and / or height of the nearest reference sample along the first and second image directions by padding from the nearest extended reference sample.

[0207] 4. In another embodiment, only per-second angular modes may be used for the extended reference samples. As a result, the derivation of the MPM list is modified to exclude these modes in addition to DC mode and planar mode for i>0. That is, the video encoder may be configured to predict predictions using angular prediction modes that use only a subset of angles from the possible angles of the angular prediction modes and to exclude unused angles from reporting encoding information to the decoder. Thus, the video decoder may be configured to predict predictions using angular prediction modes that use only a subset of angles from the possible angles of the angular prediction modes and to exclude unused angles from the prediction.

[0208] 5. In another embodiment, the number of additional reference sample lines is increased to three (see FIG. 11 with MaxNumRefAreaIdx=4). That is, the extended reference samples can be arranged at least two lines and rows in addition to the nearest reference samples, preferably at least three lines and rows. Such a configuration can be applied to a video encoder and a decoder. A video encoder can be configured to use a specific set of extended reference samples to predict a predictive block, and the video encoder is configured to select a specific set from the sets to include the lowest similarity in image content when compared to the nearest reference samples extended by the set, i.e., to use related reference samples with the same reference area index, for example. Thus, a video decoder can be configured to use a specific set of extended reference samples to predict a predictive block, and the video decoder is configured to select a specific set from the sets to include the lowest similarity in image content when compared to the nearest reference samples extended by the set.

[0209] 6. In another embodiment, instead of the index i, a flag indicating whether an extended reference sample is used is signaled. If the flag indicates that an extended reference sample is used, the index i (i>0) is derived by calculating the similarity between the reference line i>0 and the reference line i=0 (e.g., using the sum of absolute differences). The index i that results in the lowest similarity is selected. The idea behind this is that the higher the correlation between the extended reference sample (i>0) and the normal reference sample (i=0), the higher the correlation of the resulting prediction, so there is no additional benefit in using the extended reference sample in the prediction. That is, the video encoder can be configured to signal the use of the extended reference sample using a flag or other, possibly binary, information. Therefore, the video decoder can be configured to receive information indicating the use of the extended reference sample via such a flag.

[0210] 7. In another embodiment, a secondary transform, such as a non-separable secondary transform (NSST), may be applied after the first transform of the intra-prediction residual. For extended reference sample lines, the secondary transform is not performed, and all notifications related to the NSST are disabled when i>0. That is, the video encoder may be configured to selectively use only extended reference samples or nearest reference samples, and may be configured to transform the residual obtained by predicting the predictive block using a first transform procedure to obtain a first transform result, and to transform the first transform result using a second transform procedure to obtain a second transform result when extended reference samples are not used to predict the predictive block. This may also affect the notification of whether the second transform is used. That is, if the use of the second transform needs to be notified, the notification may be skipped when extended reference samples are used. The video encoder may be configured to signal the use of the secondary transform, or to implicitly signal non-use of the secondary transform when indicating the use of extended reference samples, and not include information related to the result of the secondary transform in the encoded data.

[0211] Therefore, the video decoder can be configured to selectively use only the extended reference samples or the nearest reference samples, and the video decoder is configured to transform a residual obtained by predicting the predictive block using a first transform procedure to obtain a first transform result, and to transform the first transform result using a second transform procedure to obtain a second transform result when the extended reference samples are not used to predict the predictive block. The video decoder can be configured to receive information indicating the use of a secondary transform, or to derive no use of a secondary transform when indicating the use of an extended reference sample, and to not receive information related to the result of the secondary transform of the encoded data.

[0212] 8. In another embodiment, prediction using extended reference samples (i>0) is combined with planar prediction using the nearest reference sample (i=0). As outlined in relation to the combination of different extended prediction modes, that is, using extended reference samples, weighting can be fixed (e.g., 0.5 and 0.5), dependent on block size, or signaled at the slice, image, or sequence level. When extended reference samples are used (i>0), an additional flag indicates whether combined prediction is applied. That is, the video encoder can be configured. The video decoder can be configured accordingly.

[0213] 9. In another embodiment, to reduce notification overhead, the notification of combined prediction from above is omitted. Instead, the decision of whether to apply combined prediction is derived based on the analysis of the nearest reference sample (i=0). One possible analysis may be the flatness of the nearest reference sample. If the nearest reference sample signal is flat (no edges), combining is applied, and if it contains high frequencies and edges, combining is not applied. That is, the video encoder can be configured. The video decoder can be configured accordingly.

[0214] While some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or apparatus corresponds to a method step or feature of a method step, and similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or function of a corresponding apparatus.

[0215] Depending on particular implementation requirements, embodiments of the present invention can be implemented in hardware or software, using a digital storage medium such as, for example, a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM or flash memory, on which electronically readable control signals are stored and which cooperate (or can cooperate) with a programmable computer system to perform the respective method.

[0216] Some embodiments of the present invention comprise a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.

[0217] Generally, embodiments of the present invention can be implemented as a computer program product comprising program code that operates to perform one of the methods when the computer program product is run on a computer, and the program code may for example be stored on a machine-readable carrier.

[0218] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0219] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0220] A further embodiment of the inventive method is, therefore, a data carrier (or digital storage medium, or computer readable medium) having recorded thereon the computer program for performing one of the methods described herein.

[0221] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, the data stream or sequence of signals being adapted to be transferred via a data communication connection, such as, for example, the Internet.

[0222] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0223] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0224] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0225] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented by way of illustration and description of the embodiments herein.

Claims

1. 1. A method for video decoding, comprising: determining whether a reference sample at a first position within a reference line is available, the reference line being spatially separated from the prediction block; if it is determined that the reference sample at the first position is unavailable, determining whether each reference sample included in the reference line is available, starting from a reference sample adjacent to the reference sample at the first position, until it is determined that a reference sample at a second position is available; setting the value of the reference sample at the first location to the value of the reference sample at the second location; generating at least one additional reference sample having a value of a reference sample located at the end of the reference line; decoding the prediction block using the value of the reference sample at the first position and the value of the reference sample located at the end of the reference line; Including, the reference line is an extended reference line; The method comprises: determining whether adjacent reference lines directly adjacent to the predicted block are available; if it is determined that the adjacent reference line is available, determining whether the reference sample at the first position within the extended reference line is available; The method further comprises:

2. after the value of the reference sample at the first position is set, setting a value of at least one reference sample located between the first position and the second position to the value of the reference sample at the first position; The method of claim 1 further comprising:

3. identifying the reference line from a plurality of reference lines based on syntax elements of a bitstream of encoded video data; The method of claim 1 further comprising:

4. filtering the reference samples of the reference line; decoding the prediction block using the filtered values ​​of the reference samples of the reference line and values ​​of the at least one additional reference sample; The method of claim 1 further comprising:

5. and decoding the prediction block further comprises using at least one angular prediction mode to decode the prediction block. The method of claim 1.

6. and decoding the predictive block further includes determining a list of most probable prediction modes based on a prediction mode used to predict at least one reference sample in the reference line. The method of claim 1.

7. the reference line is spatially separated from the prediction block by at least one sample position; The method of claim 1.

8. 1. A decoder for video decoding, comprising: determining whether a reference sample at a first position within a reference line is available, the reference line being spatially separated from the prediction block; if it is determined that the reference sample at the first position is unavailable, determining whether each reference sample included in the reference line is available, starting with a reference sample adjacent to the reference sample at the first position, until it is determined that a reference sample at a second position is available; setting the value of the reference sample at the first location to the value of the reference sample at the second location; generating at least one additional reference sample having a value of a reference sample located at the end of the reference line; decoding the prediction block using the value of the reference sample at the first position and the value of the reference sample located at the end of the reference line; It is configured as follows: the reference line is an extended reference line; The decoder determining whether adjacent reference lines directly adjacent to the predicted block are available; if it is determined that the adjacent reference line is available, determining whether the reference sample at the first position within the extended reference line is available; The decoder is further configured to:

9. after the value of the reference sample at the first position is set, setting a value of at least one reference sample located between the first position and the second position to the value of the reference sample at the first position; 9. The decoder of claim 8, further configured to:

10. identifying the reference line from a plurality of reference lines based on syntax elements of a bitstream of encoded video data; 9. The decoder of claim 8, further configured to:

11. filtering the reference samples of the reference lines; decoding the prediction block using the filtered values ​​of the reference samples of the reference line and values ​​of the at least one additional reference sample; 9. The decoder of claim 8, further configured to:

12. To decode the prediction block, the decoder is further configured to use at least one angular prediction mode to decode the prediction block.

9. A decoder according to claim 8.

13. To decode the predictive block, the decoder is further configured to determine a list of most probable prediction modes based on a prediction mode used to predict at least one reference sample in the reference line.

9. A decoder according to claim 8.

14. the reference line is spatially separated from the prediction block by at least one sample position; 9. A decoder according to claim 8.

15. A non-transitory computer-readable storage medium, comprising: When executed, the method causes at least one processor to: determining whether a reference sample at a first position within a reference line is available, the reference line being spatially separated from the prediction block; if it is determined that the reference sample at the first position is unavailable, determining whether each reference sample included in the reference line is available, starting from a reference sample adjacent to the reference sample at the first position, until it is determined that a reference sample at a second position is available; setting the value of the reference sample at the first location to the value of the reference sample at the second location; generating at least one additional reference sample having a value of a reference sample located at the end of the reference line; decoding the prediction block using the value of the reference sample at the first position and the value of the reference sample located at the end of the reference line; Memorize the command, the reference line is an extended reference line; The non-transitory computer-readable storage medium, when executed, causes the at least one processor to: determining whether adjacent reference lines directly adjacent to the predicted block are available; if it is determined that the adjacent reference line is available, determining whether the reference sample at the first position within the extended reference line is available; A non-transitory computer-readable storage medium further storing instructions.

16. When executed, the at least one processor: after the value of the reference sample at the first position is set, setting the value of at least one reference sample located between the first position and the second position to the value of the reference sample at the first position; 16. The non-transitory computer-readable storage medium of claim 15, further storing instructions.

17. When executed, the at least one processor: identifying the reference line from a plurality of reference lines based on syntax elements of a bitstream of encoded video data; 16. The non-transitory computer-readable storage medium of claim 15, further storing instructions.

18. When executed, the at least one processor: filtering the reference samples of the reference lines; decoding the prediction block using the filtered values ​​of the reference samples of the reference line and the values ​​of the at least one additional reference sample; 16. The non-transitory computer-readable storage medium of claim 15, further storing instructions.

19. The instructions that, when executed, cause the at least one processor to decode the predictive block, when executed, cause the at least one processor to: using at least one angular prediction mode for decoding the prediction block; 20. The non-transitory computer-readable storage medium of claim 15, further comprising instructions.

20. The instructions that, when executed, cause the at least one processor to decode the predictive block, when executed, cause the at least one processor to: determining a list of most probable prediction modes based on a prediction mode used to predict at least one reference sample in the reference line; 20. The non-transitory computer-readable storage medium of claim 15, further comprising instructions.

21. the reference line is spatially separated from the prediction block by at least one sample position; 16. The non-transitory computer-readable storage medium of claim 15.

Citation Information

Patent Citations

  • Method and device for encoding and decoding image

    US20140362906A1

  • Intra-picture prediction using non-adjacent reference lines of sample values

    WO2017190288A1

  • Image processing device and image processing method

    WO2018070267A1

  • Intra-prediction with multiple reference lines

    WO2018205950A1

  • Method and apparatus for multiple line intra prediction in video compression

    WO2019155451A1