Apparatuses and methods for encoding and decoding a video using prediction refinement
Patent Information
- Application Number
- CA3324075
- Authority / Receiving Office
- CA · CA
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-15
- Filing Date
- 2025-03-13
- Publication Date
- 2025-09-18
AI Technical Summary
Existing video coding technologies face challenges in achieving precise prediction signals, leading to inefficient entropy encoding and suboptimal rate-distortion performance due to large residual signals.
The proposed solution involves refining the prediction signal by deriving an estimation for the prediction signal's precision in an extension region of a picture, comparing it with the actual reconstruction, and adjusting the prediction based on the deviation to minimize residual errors.
This approach results in a more precise prediction, reducing residual signals and improving the rate-distortion performance during encoding, thereby enhancing coding efficiency.
Abstract
Description
[0001] Apparatuses and methods for encoding and decoding a video using prediction refinement
[0002] Description
[0003] Embodiments of the present invention provide an apparatus for encoding a video, an apparatus for decoding a video, a method for encoding a video, a method for decoding a video, and a data stream comprising a video.
[0004] For example, embodiments of the present invention may apply to hybrid video coding. For example, embodiments may have the purpose to improve the prediction signal, thus achieving a better coding performance.
[0005] In block-based predictive video coding, a block of a picture of a video is encoded by determining a prediction signal for the block based on one or more previously coded blocks of the video, and encoding a residual between the block and the prediction signal derived for the block. The better the prediction of the block matches the block, the smaller is the residual signal, eventually resulting in a higher number of zeros when encoding the block, e.g., after transforming the block using a linear transform. In turn, a high number of zeros allows for an efficient entropy encoding, resulting in a low bitrate or a good rate-distortion- relation.
[0006] Accordingly, it is desirable to provide a video coding concept that provides a precise prediction signal.
[0007] This objective is achieved by the subject-matter of the independent claims.
[0008] Embodiments of the present invention rely on the idea to refine a prediction signal for a portion of a picture by deriving an estimation for a precision of the prediction signal. To this end, the derivation of the prediction signal may be extended to an extension region, which is located in an already reconstructed portion of the picture, and the obtained prediction for the extension region may be compared with the actual reconstruction of the extension region, thereby providing an estimate for a precision of the prediction in terms of a deviation between the reconstruction of the extension region and the prediction of the extension region. The prediction signal for the extension region may then be refined based on this deviation. This refinement may result in a more precise prediction of the portion, thereby providing a smaller residual, and eventually, a better rate-distortion relation in encoding the portion into the data stream.
[0009] An embodiment of the present invention provides an apparatus for decoding a video from a data stream. The apparatus is configured for reconstructing a picture (e.g., a current picture) of the video in units of portions (e.g., blocks, e.g., rectangular blocks), into which the picture is subdivided, wherein the apparatus is configured for reconstructing a predetermined portion (e.g., a currently reconstructed portion) of the picture by: obtaining a prediction residual of the predetermined portion (e.g., deriving the prediction residual from the data stream, e.g., in form of a residual signal or in form of an indication that the prediction residual is zero); obtaining prediction parameters for the predetermined portion (e.g., from the data stream), and using the prediction parameters for deriving, based on a previously reconstructed portion of the video (e.g., a previously reconstructed picture of the video or a previously reconstructed portion of the current picture), a prediction signal for the predetermined portion and an extension prediction signal for an extension region which spatially neighbors the predetermined portion, the extension region being part of a previously reconstructed portion of the picture; deriving a difference signal (or correction signal) based on the extension prediction signal and based on reconstructed samples of the extension region; deriving a refined prediction signal for the predetermined portion based on the prediction signal for the predetermined portion and the difference signal; and reconstructing the predetermined portion based on the refined prediction signal and the prediction residual (e.g., combining the refined prediction signal with a residual signal of the predetermined portion, e.g., to obtain reconstructed samples of the portion).
[0010] For example, predicting the extension region using the same prediction parameters as for predicting the predetermined portion for deriving the extension prediction signal, and comparing the extension prediction signal to the reconstructed samples of the extension region, provides a measure for a prediction error when applying the prediction parameters to the extension region. Embodiments of the invention rely on the finding that, as the extension region neighbors the predetermined portion to be reconstructed, the prediction error in the extension region may be a good estimate for a prediction error in the predetermined portion, and that it is therefore suitable for refining the prediction for the predetermined portion.
[0011] According to an embodiment, the apparatus is configured for deriving the difference signal by deriving, for a sample position of the extension region (e.g., for each sample position of the extension region), a weighted difference between a reconstructed sample value of the respective sample position (e.g., a sample value of the reconstructed portion of the picture at the respective sample position) and a predicted sample value of the extension prediction signal for the extension region (e.g., to obtain a difference sample value of the difference signal).
[0012] According to an embodiment, the apparatus is configured for deriving the difference signal by deriving, for a sample position of the extension region (e.g., for each sample position of the extension region), a weighted difference between a reconstruction value obtained from one or more reconstructed sample values positioned within a region (e.g., a first region) around the respective sample position (e.g., by averaging the one or more reconstructed sample values of the region) and a prediction value obtained from one or more predicted sample values of the extension prediction signal for the extension region, the one or more predicted sample values being positioned within a region (e.g., a second region, or the first region) around the respective sample position (e.g., by averaging the one or more predicted sample values of the region).
[0013] According to an embodiment, the weighted difference is an equal-weighted difference.
[0014] According to an embodiment, weights of the weighted difference depend on the sample position, for which the weighted difference is derived. For example, the extent, to which the prediction of the extension region provides a good estimate for a prediction error of the predetermined portion may depend on the sample position.
[0015] According to an embodiment, the apparatus is configured for deriving the refined prediction signal by forming, for a predetermined sample position (or one of the sample positions, e.g., predetermined in terms of a currently considered one) of the predetermined portion (e.g., for each sample position of the predetermined portion), a combination (e.g. a sum, e.g., a weighted sum or an equal weighted sum) of a sample value of the prediction signal for the predetermined sample position and the output of a function of the difference signal.
[0016] According to an embodiment, the function depends on one or more of the predetermined sample position, a shape of the predetermined portion, a size of the predetermined portion, a coding mode for the predetermined portion, and a prediction mode for the predetermined portion (e.g., a manner according to which the prediction signal is derived). According to an embodiment, the apparatus is configured for deriving the refined prediction signal by forming, for a predetermined sample position of the predetermined portion (e.g., for each sample position of the predetermined portion), a combination of a sample value of the prediction signal for the predetermined sample position and one or more difference signal values of the difference signal (E.g., a weighted sum of a sample value of the prediction signal for the predetermined sample position and one or more difference signal values of the difference signal .E.g., the weighted sum comprises, or consists of, a sample value of the prediction signal for the predetermined sample position and one or more difference signal values of the difference signal).
[0017] According to an embodiment, the apparatus is configured for deriving the refined prediction signal by forming, for a predetermined sample position of the predetermined portion (e.g., for each sample position of the predetermined portion), a combination of a sample value of the prediction signal for the predetermined sample position and a weighted sum of one or more difference signal values of the difference signal wherein the weighted sum comprises all difference signal values of the difference signal (e.g., difference signal values for all sample positions of the extension region).
[0018] According to an embodiment, the apparatus is configured for deriving weights according to which each of the one or more difference signal values of the combination contribute to the refined prediction signal at the predetermined sample position (e.g., weights for the one or more difference signal values in the weighted sum) in dependence on a sample position of the respective difference signal value and / or a shape of the predetermined portion and / or a size of the predetermined portion and / or an indication derived from the data stream.
[0019] According to an embodiment, the apparatus is configured for deriving a syntax element from the data stream, which indicates weights according to which each of the one or more difference signal values of the combination contribute to the refined prediction signal at the predetermined sample position (e.g., weights for the one or more difference signal values in the weighted sum) (e.g., the syntax element indicates one out of a set of sets of weights to be applied for the one or more difference signal values, or the syntax element indicates one out of a set of functions for deriving the weights to be applied for the one or more difference signal values, (E.g., the weights may depend on further parameters, such as position, block size, block shape)). According to an embodiment, the syntax element is signaled in the data stream individually for the predetermined portion. (E.g., the apparatus derives the syntax element individually for portions, to which the prediction refinement is applied.)
[0020] According to an embodiment, the apparatus is configured for deriving weights, according to which each of the one or more difference signal values of the combination contribute to the refined prediction signal at the predetermined sample position, (e.g., weights for the one or more difference signal values in the weighted sum) in dependence on a distance of respective sample positions of the difference signal values from a reference position (e.g., a corner, e.g., a top left corner) of the predetermined portion.
[0021] According to an embodiment, the apparatus is configured for deriving weights, according to which each of the one or more difference signal values of the combination contribute to the refined prediction signal at the predetermined sample position (e.g., weights for the one or more difference signal values in the weighted sum), in a manner that a smaller or equal weight (e.g., in terms of absolute value) is assigned to a difference signal value positioned at a higher distance from the reference position of the predetermined portion.
[0022] According to an embodiment, the apparatus is configured for deriving a weight according to which a difference signal value of the combination contributes to the refined prediction signal at the predetermined sample position (e.g., weights for the one or more difference signal values in the weighted sum), which difference signal value has a sample position in a horizontally adjacent portion of the predetermined portion, in a manner that a smaller or equal weight (e.g., in terms of absolute value) is assigned to a difference signal value positioned at a higher distance in vertical direction from the reference position of the predetermined portion. Additionally or alternatively, the apparatus is configured for deriving a weight according to which a difference signal value of the combination contributes to the refined prediction signal at the predetermined sample position (e.g., weights for the one or more difference signal values in the weighted sum), which difference signal value has a sample position in a vertically adjacent portion of the predetermined portion, in a manner that a smaller or equal weight (e.g., in terms of absolute value) is assigned to a difference signal value positioned at a higher distance in horizontal direction from the reference position of the predetermined portion.
[0023] According to an embodiment, the apparatus is configured for deriving a weight according to which a difference signal value of the combination contributes to the refined prediction signal at the predetermined sample position (e.g., weights for the one or more difference signal values in the weighted sum), which difference signal value has a sample position in a horizontally adjacent portion of the predetermined portion, using a first function, which depends on a distance in vertical direction of the sample position of the difference signal value from the reference position of the predetermined portion in a manner that a smaller or equal weight (e.g., in terms of absolute value) is assigned to a difference signal value positioned at a higher distance in vertical direction from the reference position. According to this embodiment, the apparatus is configured for deriving a weight according to which a difference signal value of the combination contributes to the refined prediction signal at the predetermined sample position (e.g., weights for the one or more difference signal values in the weighted sum), which difference signal value has a sample position in a vertically adjacent portion of the predetermined portion, using a second function, which depends on a distance in horizontal direction of the sample position of the difference signal value from the reference position of the predetermined portion in a manner that a smaller or equal weight (e.g., in terms of absolute value) is assigned to a difference signal value positioned at a higher distance in horizontal direction from the reference position. According to this embodiment, the first function equals the second function, or wherein the first function differs from the second function.
[0024] According to an embodiment, the first function and the second function depend on one or more of a size, a shape of the predetermined portion, a coding mode for the predetermined portion, and a prediction mode for the predetermined portion (e.g., a manner according to which the prediction signal is derived).
[0025] According to an embodiment, the apparatus is configured for deriving the refined prediction signal by deriving the difference signal (e.g., “d”) by deriving, for a sample position of the extension region (e.g., for each sample position of the extension region), a difference signal value of the difference signal based on a weighted difference between one or more reconstructed samples of the extension region on the one hand and one or more predicted sample values of the extension prediction signal for the extension region on the other hand.
[0026] According to an embodiment, the apparatus is configured for deriving the refined prediction signal by: deriving, for a predetermined sample position of the predetermined portion (e.g., for each sample position of the predetermined portion), a refinement sample value of a refinement signal based on (e.g., based on, e.g. by forming a weighted sum of) one or more difference signal values of the difference signal; and deriving a sample value of the refined prediction signal at the predetermined sample position based on a combination (e.g. a sum, e.g., a weighted sum or an equal weighted sum) of a predicted sample value of the prediction signal for the predetermined sample position and the refinement sample value for the predetermined sample position.
[0027] According to an embodiment, the apparatus is configured for deriving the refinement sample value of the refinement signal by forming a first weighted sum of a plurality of difference signal values of the difference signal. According to this embodiment, weights of the first weighted sum depend on one or more of the predetermined sample position, a size of the predetermined portion, a shape of the predetermined portion, a coding mode of the predetermined portion, a prediction mode of the predetermined portion, and a dedicated syntax element derived from the data stream which indicates the weights of the weighted sum (e.g., a syntax element dedicated for the predetermined portion).
[0028] According to an embodiment, the apparatus is configured for deriving the refined prediction signal by combining the refinement signal and the prediction signal in a position-wise manner (e.g., in a manner according to which each predicted sample value of the prediction signal is combined with a respective collocated (e.g., located at the same sample position within the predetermined portion) refinement sample value of the refinement signal).
[0029] According to an embodiment, the apparatus is configured for deriving the sample value of the refined prediction signal at the predetermined sample position by forming a second weighted sum of the predicted sample value and the refinement sample value for the predetermined sample position.
[0030] According to an embodiment, weights of the second weighted sum depend on one or more of the predetermined sample position, a size of the predetermined portion, a shape of the predetermined portion, a coding mode of the predetermined portion, a prediction mode of the predetermined portion, and a dedicated syntax element derived from the data stream which indicates the weights of the weighted sum (e.g., a syntax element dedicated for the predetermined portion).
[0031] According to an embodiment, the one or more difference signal values for deriving the refinement signal value for the predetermined sample position comprise (or consist of) one or more difference sample values which are equally positioned with respect to a horizontal direction as the predetermined sample position of the predetermined portion, and / or one or more difference sample values which are equally positioned with respect to a vertical direction as the predetermined sample position of the predetermined portion.
[0032] According to an embodiment, the apparatus is configured for filtering the difference signal and deriving the refined prediction signal based on the filtered difference signal.
[0033] According to an embodiment, the filtering comprises a smoothing and / or a low-pass- filtering.
[0034] According to an embodiment, the apparatus is configured for sub-sampling the difference signal and deriving the refined prediction signal based on the sub-sampled difference signal.
[0035] According to an embodiment, the video is a color video, wherein the picture comprises multiple channels (e.g., color channels, e.g., Y, Cb, Cr), and wherein the apparatus is configured for reconstructing each of the channels by reconstructing a predetermined portion (e.g., a currently reconstructed portion) of the respective channel (e.g., a predetermined portion, a currently decoded one of the portions of the channel) by: obtaining a prediction residual of the predetermined portion of the channel; obtaining prediction parameters for the predetermined portion of the channel (e.g., from the data stream), and using the prediction parameters of the predetermined portion of the channel for deriving, based on a previously reconstructed portion of the channel (e.g., a previously reconstructed picture of the channel or a previously reconstructed portion of the current picture of the channel), a prediction signal for the predetermined portion of the channel; and reconstructing the predetermined portion of the channel based on the prediction signal and the prediction residual of the channel (e.g., to obtain reconstructed samples of the portion). According to this embodiment, the apparatus is configured for performing, for one or more or all of the multiple channels, a refinement of the prediction signal of the predetermined portion of the respective channel by: using the prediction parameters obtained for the predetermined portion of the channel for deriving, based on the previously reconstructed portion of the channel, an extension prediction signal for an extension region of the predetermined portion of the channel which extension region spatially neighbors the predetermined portion of the channel, the extension region being part of a previously reconstructed portion of the channel; deriving a difference signal (or difference signal or correction signal) for the predetermined portion of the channel based on the extension prediction signal of the predetermined portion of the channel and based on reconstructed samples of the extension region of the predetermined portion of the channel; and refining the prediction signal of the predetermined portion of the channel using the difference signal of the predetermined portion of the channel (and combining the refined prediction signal for the predetermined portion of the channel with the residual signal of the predetermined portion of the channel).
[0036] According to an embodiment, the apparatus is configured to perform the refinement of the respective prediction signals of the one or more or all channels independently from each other (e.g., independently from the refinement of the prediction signals of further channels).
[0037] According to an embodiment, the apparatus is configured for performing the refinement of the prediction signal of a predetermined one of the channels by deriving the difference signal for the predetermined portion of the predetermined channel further based on the extension prediction signal derived for one or more further ones of the channels
[0038] According to an embodiment, the apparatus is configured for deriving the difference signal for the predetermined portion of the predetermined channel further based on a difference signal and / or a prediction signal and / or a residual signal and / or a reconstructed signal derived for the one or more further ones of the channels.
[0039] According to an embodiment, the extension region comprises (or consists of): one or more columns of sample positions of a horizontally adjacent portion of the predetermined portion, the one or more columns being horizontally adjacent to the predetermined portion, and / or one or more rows of sample positions of a vertically adjacent portion of the predetermined portion, the one or more rows being horizontally adjacent to the predetermined portion, and / or one or more samples of a horizontally and vertically adjacent portion of the predetermined portion.
[0040] According to an embodiment, the previously reconstructed portion of the video (based on which the prediction signal is derived), comprises (or is part of) one or more previously reconstructed pictures.
[0041] According to an embodiment, the apparatus is, configured for deriving the prediction signal for the predetermined portion using temporal prediction. According to an embodiment, the apparatus is configured for: activating or deactivating, (e.g., individually) for each of portions of the picture (e.g., a set of the portions into which the picture is subdivided), a refinement of a prediction signal of the respective portion; and for each of portions, for which the refinement of the prediction signal is activated, reconstructing the respective portion by combining a refined prediction signal of the respective portion with a prediction residual of the respective portion; and for each of portions, for which the refinement of the prediction signal is deactivated, reconstructing the respective portion by combining the prediction signal of the respective portion with a prediction residual of the respective portion.
[0042] According to an embodiment, the apparatus is configured for, for each of portions, for which the refinement of the prediction signal is activated: deriving an extension prediction signal for an extension region which spatially neighbors the respective portion, the extension region being part of a previously reconstructed portion of the picture; deriving a difference signal based on the extension prediction signal and based on reconstructed samples of the extension region; and deriving the refined prediction signal for the respective portion based on the prediction signal for the portion and the difference signal.
[0043] According to an embodiment, the apparatus is configured for: for each of a first set of portions of the picture, (e.g., individually) activating or deactivating the refinement of the prediction signal of the respective portion in dependence on an indication (e.g., specific to the respective portion) derived from the data stream; and / or for each of a second set of portions of the picture, activating the refinement of the prediction signal of the respective portion; and / or for each of a third set of portions of the picture, deactivating the refinement of the prediction signal of the respective portion; and / or for each of a fourth set of portions of the picture, inferring whether to activate or deactivate the refinement of the prediction signal of the respective portion (e.g., in dependence on a reference portion, e.g., a merge candidate, assigned to the respective portion).
[0044] According to an embodiment, the apparatus is configured for assigning a portion of the picture to one of the first set of portions, the second set of portions, the third set of portions, or the fourth set of portions based on a prediction mode assigned to the respective portion.
[0045] According to an embodiment, the apparatus is configured for activating or deactivating, (e.g., individually) for each of portions of the picture (e.g., a set of the portions into which the picture is subdivided) the refinement of the prediction signal of the respective portion in dependence on: an indication derived from the data stream (e.g., a flag, e.g., referred to as refine flag) which indicates whether or not to activate or deactivate the derivation of the refinement of the prediction signal for the respective portion; and / or a prediction mode of the respective portion (e.g., the prediction mode may be a mode according to which the prediction signal may be obtained) (e.g., the prediction modes may include a temporal prediction mode, e.g., referred to as inter-prediction mode, e.g., a regular inter-prediction mode, an affine inter-prediction mode, and a geometric mode, a merge mode, e.g., a regular merge mode, or a combined inter / intra prediction (CUP) mode, and or one or more intraprediction modes).
[0046] According to an embodiment, the apparatus is configured for, if the prediction mode of the respective portion is a first temporal prediction mode (e.g., a regular inter-prediction mode), derive an indication from the data stream which indicates whether or not to activate or deactivate the derivation of the refinement of the prediction signal for the respective portion.
[0047] According to an embodiment, the apparatus is configured for, if the prediction mode of the respective portion is a merge mode (e.g., a mode according to which prediction parameters for the respective block are inherited from a previously reconstructed portion, e.g. a specified previously reconstructed portion (e.g., referred to as specified merging candidate)), deriving an indication which indicates whether or not to activate or deactivate the derivation of the refinement of the prediction signal for the respective portion from a respective indication for a previously reconstructed portion (e.g. the specified previously reconstructed portion).
[0048] According to an embodiment, the apparatus is configured for, if the prediction mode of the respective portion is not the first temporal prediction mode and not the merge mode, deactivate the refinement of the prediction signal of the respective portion.
[0049] A further embodiment of the invention provides an apparatus for encoding a video into a data stream, configured for encoding a picture (e.g., a current picture) of the video in units of portions (e.g., blocks, e.g., rectangular blocks), into which the picture is subdivided. The apparatus is configured for encoding a predetermined portion (e.g., a currently reconstructed portion) of the picture (e.g., a predetermined portion, a currently decoded one of the portions of the picture) by: deriving (e.g., based on a previously encoded portion of the video) prediction parameters for the predetermined portion, and using the prediction parameters for deriving, based on a previously encoded portion of the video (e.g., a previously encoded and reconstructed portion of the video, e.g., a previously reconstructed picture of the video or a previously reconstructed portion of the current picture), a prediction signal for the predetermined portion and an extension prediction signal for an extension region which spatially neighbors the predetermined portion, the extension region being part of a previously encoded portion of the picture; deriving a difference signal (or correction signal) based on the extension prediction signal and based on previously encoded and reconstructed samples of the extension region (E.g., the encoder may reconstruct the previously encoded portion and use the previously encoded and reconstructed portion as a basis for deriving the prediction signal and the extension prediction signal. Thus, when referred to reconstructed sample values, in particular with respect to the encoder, these may be encoded and reconstructed sample values, e.g., sample values of a reconstruction of the previously encoded portion of the picture); deriving a refined prediction signal for the predetermined portion based on the prediction signal for the predetermined portion and the difference signal; and using the refined prediction signal to derive a prediction residual of the predetermined portion (e.g., in form of a residual signal or in form of an indication that the prediction residual is zero).
[0050] A further embodiment of the invention provides a method for decoding a video from a data stream, the method comprising reconstructing a picture of the video in units of portions, into which the picture is subdivided, wherein the method comprises reconstructing a predetermined portion of the picture by: obtaining a prediction residual of the predetermined portion; obtaining prediction parameters for the predetermined portion, and using the prediction parameters for deriving, based on a previously reconstructed portion of the video, a prediction signal for the predetermined portion and an extension prediction signal for an extension region which spatially neighbors the predetermined portion, the extension region being part of a previously reconstructed portion of the picture; deriving a difference signal based on the extension prediction signal and based on reconstructed samples of the extension region; deriving a refined prediction signal for the predetermined portion based on the prediction signal for the predetermined portion and the difference signal; and reconstructing the predetermined portion based on the refined prediction signal and the prediction residual.
[0051] A further embodiment of the invention provides a method for encoding a video into a data stream, wherein the method comprises encoding a picture of the video in units of portions, into which the picture is subdivided, wherein the method comprises encoding a predetermined portion of the picture by: deriving prediction parameters for the predetermined portion, and using the prediction parameters for deriving, based on a previously encoded portion of the video, a prediction signal for the predetermined portion and an extension prediction signal for an extension region which spatially neighbors the predetermined portion, the extension region being part of a previously encoded portion of the picture; deriving a difference signal based on the extension prediction signal and based on previously encoded and reconstructed samples of the extension region; deriving a refined prediction signal for the predetermined portion based on the prediction signal for the predetermined portion and the difference signal; and using the refined prediction signal to derive a prediction residual of the predetermined portion.
[0052] A further embodiment of the invention provides a data stream comprising a video, the video being encoded into the data stream using the disclosed method for encoding a video.
[0053] Advantageous implementations are defined by the subject-matter of the dependent claims.
[0054] Embodiments of the present disclosure are described in more detail below with respect to the figures, among which:
[0055] Fig. 1 illustrates an encoder providing a framework for an embodiment,
[0056] Fig. 2 illustrates a decoder providing a framework for an embodiment,
[0057] Fig. 3 illustrates a subdivision of a picture into blocks according to an embodiment,
[0058] Fig. 4 illustrates a decoder according to an embodiment,
[0059] Fig. 5 illustrates the extension region according to an embodiment,
[0060] Fig. 6 illustrates a refinement module according to an embodiment,
[0061] Fig. 7 illustrates a prediction refinement using the spatial reference samples of a bipredicted inter block according to an embodiment,
[0062] Fig. 8 position-dependent weights for various block widths / heights according to an embodiment, and Fig. 9 illustrates an encoder according to an embodiment,
[0063] Fig. 10 illustrates an encoder according to another embodiment.
[0064] Embodiments of the present invention are now described in more detail with reference to the accompanying drawings, in which the same or similar elements or elements that have the same or similar functionality have the same reference signs assigned or are identified with the same name. In the following description, a plurality of details is set forth to provide a thorough explanation of embodiments of the disclosure. However, it will be apparent to one skilled in the art that other embodiments may be implemented without these specific details. In addition, features of the different embodiments described herein may be combined with each other, unless specifically noted otherwise.
[0065] The following description of the figures starts with a presentation of a description of an encoder and a decoder of a block-based predictive codec for coding pictures of a video in order to form an example for a coding framework into which embodiments of the present invention may be built in. The respective encoder and decoder are described with respect to Fig. 1 , Fig. 2, and Fig. 3. Thereinafter the description of embodiments of the concept of the present invention is presented along with a description as to how such concepts could be built into the encoder and decoder of Fig. 1 , and Fig. 2, respectively, although the embodiments described with the subsequent Figures and following, may also be used to form encoders and decoders not operating according to the coding framework underlying the encoder and decoder of Fig. 1 , and Fig. 2.
[0066] Fig. 1 shows an apparatus for predictively coding a picture 12 into a data stream 14 exemplarily using transform-based residual coding. The apparatus, or encoder, is indicated using reference sign 10. Fig. 2 shows a corresponding decoder 20, i.e. an apparatus 20 configured to predictively decode the picture 12’ from the data stream 14 also using transform-based residual decoding, wherein the apostrophe has been used to indicate that the picture 12’ as reconstructed by the decoder 20 deviates from picture 12 originally encoded by apparatus 10 in terms of coding loss introduced by a quantization of the prediction residual signal. Fig. 1 and Fig. 2 exemplarily use transform based prediction residual coding, although embodiments of the present application are not restricted to this kind of prediction residual coding. This is true for other details described with respect to Fig. 1 , and Fig. 2, too, as will be outlined hereinafter. The encoder 10 is configured to subject the prediction residual signal to spatial-to-spectral transformation and to encode the prediction residual signal, thus obtained, into the data stream 14. Likewise, the decoder 20 is configured to decode the prediction residual signal from the data stream 14 and subject the prediction residual signal thus obtained to spectral- to-spatial transformation.
[0067] Internally, the encoder 10 may comprise a prediction residual signal former 22 which generates a prediction residual 24 so as to measure a deviation of a prediction signal 26 from the original signal, i.e. from the picture 12. The prediction residual signal former 22 may, for instance, be a subtractor which subtracts the prediction signal from the original signal, i.e. from the picture 12. The encoder 10 then further comprises a transformer 28 which subjects the prediction residual signal 24 to a spatial-to-spectral transformation to obtain a spectral-domain prediction residual signal 24’ which is then subject to quantization by a quantizer 32, also comprised by the encoder 10. The thus quantized prediction residual signal 24” is coded into bitstream 14. To this end, encoder 10 may optionally comprise an entropy coder 34 which entropy codes the prediction residual signal as transformed and quantized into data stream 14. The prediction signal 26 is generated by a prediction stage 36 of encoder 10 on the basis of the prediction residual signal 24” encoded into, and decodable from, data stream 14. To this end, the prediction stage 36 may internally, as is shown in Fig. 1 , comprise a dequantizer 38 which dequantizes prediction residual signal 24” so as to gain spectral-domain prediction residual signal 24”’, which corresponds to signal 24’ except for quantization loss, followed by an inverse transformer 40 which subjects the latter prediction residual signal 24’” to an inverse transformation, i.e. a spectral-to- spatial transformation, to obtain prediction residual signal 24””, which corresponds to the original prediction residual signal 24 except for quantization loss. A combiner 42 of the prediction stage 36 then recombines, such as by addition, the prediction signal 26 and the prediction residual signal 24”” so as to obtain a reconstructed signal 46, i.e. a reconstruction of the original signal 12. Reconstructed signal 46 may correspond to signal 12’. A prediction module 44 of prediction stage 36 then generates the prediction signal 26 on the basis of signal 46 by using, for instance, spatial prediction, i.e. intra-picture prediction, and / or temporal prediction, i.e. inter-picture prediction.
[0068] Likewise, decoder 20, as shown in Fig. 2, may be internally composed of components corresponding to, and interconnected in a manner corresponding to, prediction stage 36. In particular, entropy decoder 50 of decoder 20 may entropy decode the quantized spectral- domain prediction residual signal 24” from the data stream, whereupon dequantizer 52, inverse transformer 54, combiner 56 and prediction module 58, interconnected and cooperating in the manner described above with respect to the modules of prediction stage 36, recover the reconstructed signal on the basis of prediction residual signal 24” so that, as shown in Fig. 2, the output of combiner 56 results in the reconstructed signal, namely picture 12’.
[0069] Although not specifically described above, it is readily clear that the encoder 10 may set some coding parameters including, for instance, prediction modes, motion parameters and the like, according to some optimization scheme such as, for instance, in a manner optimizing some rate and distortion related criterion, i.e. coding cost. For example, encoder 10 and decoder 20 and the corresponding modules 44, 58, respectively, may support different prediction modes such as intra-coding modes and inter-coding modes. The granularity at which encoder and decoder switch between these prediction mode types may correspond to a subdivision of picture 12 and 12’, respectively, into coding segments or coding blocks. In units of these coding segments, for instance, the picture may be subdivided into blocks being intra-coded and blocks being inter-coded. Intra-coded blocks are predicted on the basis of a spatial, already coded / decoded neighborhood of the respective block as is outlined in more detail below. Several intra-coding modes may exist and be selected for a respective intra-coded segment including directional or angular intra- coding modes according to which the respective segment is filled by extrapolating the sample values of the neighborhood along a certain direction which is specific for the respective directional intra-coding mode, into the respective intra-coded segment. The intra- coding modes may, for instance, also comprise one or more further modes such as a DC coding mode, according to which the prediction for the respective intra-coded block assigns a DC value to all samples within the respective intra-coded segment, and / or a planar intra- coding mode according to which the prediction of the respective block is approximated or determined to be a spatial distribution of sample values described by a two-dimensional linear function over the sample positions of the respective intra-coded block with driving tilt and offset of the plane defined by the two-dimensional linear function on the basis of the neighboring samples. Compared thereto, inter-coded blocks may be predicted, for instance, temporally. For inter-coded blocks, motion vectors may be signaled within the data stream, the motion vectors indicating the spatial displacement of the portion of a previously coded picture of the video to which picture 12 belongs, at which the previously coded / decoded picture is sampled in order to obtain the prediction signal for the respective inter-coded block. This means, in addition to the residual signal coding comprised by data stream 14, such as the entropy-coded transform coefficient levels representing the quantized spectral- domain prediction residual signal 24”, data stream 14 may have encoded thereinto coding mode parameters for assigning the coding modes to the various blocks, prediction parameters for some of the blocks, such as motion parameters for inter-coded segments, and optional further parameters such as parameters for controlling and signaling the subdivision of picture 12 and 12’, respectively, into the segments. The decoder 20 uses these parameters to subdivide the picture in the same manner as the encoder did, to assign the same prediction modes to the segments, and to perform the same prediction to result in the same prediction signal.
[0070] Fig. 3 illustrates the relationship between the reconstructed signal, i.e. the reconstructed picture 12’, on the one hand, and the combination of the prediction residual signal 24”” as signaled in the data stream 14, and the prediction signal 26, on the other hand. As already denoted above, the combination may be an addition. The prediction signal 26 is illustrated in Fig. 3 as a subdivision of the picture area into intra-coded blocks which are illustratively indicated using hatching, and inter-coded blocks which are illustratively indicated nothatched. The subdivision may be any subdivision, such as a regular subdivision of the picture area into rows and columns of square blocks or non-square blocks, or a multi-tree subdivision of picture 12 from a tree root block into a plurality of leaf blocks of varying size, such as a quadtree subdivision or the like, wherein a mixture thereof is illustrated in Fig. 3 in which the picture area is first subdivided into rows and columns of tree root blocks which are then further subdivided in accordance with a recursive multi-tree subdivisioning into one or more leaf blocks.
[0071] Again, data stream 14 may have an intra-coding mode coded thereinto for intra-coded blocks 80, which assigns one of several supported intra-coding modes to the respective intra-coded block 80. For inter-coded blocks 82, the data stream 14 may have one or more motion parameters coded thereinto. Generally speaking, inter-coded blocks 82 are not restricted to being temporally coded. Alternatively, inter-coded blocks 82 may be any block predicted from previously coded portions beyond the current picture 12 itself, such as previously coded pictures of a video to which picture 12 belongs, or picture of another view or an hierarchically lower layer in the case of encoder and decoder being scalable encoders and decoders, respectively.
[0072] The prediction residual signal 24”” in Fig. 3 is also illustrated as a subdivision of the picture area into blocks 84. These blocks might be called transform blocks in order to distinguish same from the coding blocks 80 and 82. In effect, Fig. 3 illustrates that encoder 10 and decoder 20 may use two different subdivisions of picture 12 and picture 12’, respectively, into blocks, namely one subdivisioning into coding blocks 80 and 82, respectively, and another subdivision into transform blocks 84. Both subdivisions might be the same, i.e. each coding block 80 and 82, may concurrently form a transform block 84, but Fig. 3 illustrates the case where, for instance, a subdivision into transform blocks 84 forms an extension of the subdivision into coding blocks 80, 82 so that any border between two blocks of blocks 80 and 82 overlays a border between two blocks 84, or alternatively speaking each block 80, 82 either coincides with one of the transform blocks 84 or coincides with a cluster of transform blocks 84. However, the subdivisions may also be determined or selected independent from each other so that transform blocks 84 could alternatively cross block borders between blocks 80, 82. As far as the subdivision into transform blocks 84 is concerned, similar statements are thus true as those brought forward with respect to the subdivision into blocks 80, 82, i.e. the blocks 84 may be the result of a regular subdivision of picture area into blocks (with or without arrangement into rows and columns), the result of a recursive multi-tree subdivisioning of the picture area, or a combination thereof or any other sort of blockation. Just as an aside, it is noted that blocks 80, 82 and 84 are not restricted to being of quadratic, rectangular or any other shape.
[0073] Fig. 3 further illustrates that the combination of the prediction signal 26 and the prediction residual signal 24”” directly results in the reconstructed signal 12’. However, it should be noted that more than one prediction signal 26 may be combined with the prediction residual signal 24”” to result into picture 12’ in accordance with alternative embodiments.
[0074] In Fig. 3, the transform blocks 84 shall have the following significance. Transformer 28 and inverse transformer 54 perform their transformations in units of these transform blocks 84. For instance, many codecs use some sort of DST or DCT for all transform blocks 84. Some codecs allow for skipping the transformation so that, for some of the transform blocks 84, the prediction residual signal is coded in the spatial domain directly. However, in accordance with embodiments described below, encoder 10 and decoder 20 are configured in such a manner that they support several transforms. For example, the transforms supported by encoder 10 and decoder 20 could comprise: o DCT-II (or DCT-III), where DCT stands for Discrete Cosine Transform o DST-IV, where DST stands for Discrete Sine Transform o DCT-IV o DST-VII o Identity Transformation (IT) Naturally, while transformer 28 would support all of the forward transform versions of these transforms, the decoder 20 or inverse transformer 54 would support the corresponding backward or inverse versions thereof: o Inverse DCT-II (or inverse DCT-III) o Inverse DST-IV o Inverse DCT-IV o Inverse DST-VII o Identity Transformation (IT)
[0075] The subsequent description provides more details on which transforms could be supported by encoder 10 and decoder 20. In any case, it should be noted that the set of supported transforms may comprise merely one transform such as one spectral-to-spatial or spatial- to-spectral transform.
[0076] As already outlined above, Fig. 1 , Fig. 2 and Fig. 3 have been presented as an example where the inventive concept described further below may be implemented in order to form specific examples for encoders and decoders according to the present application. Insofar, the encoder and decoder of Fig. 1 , and Fig. 2, respectively, may represent possible implementations of the encoders and decoders described herein below. Fig. 1 , and Fig. 2 are, however, only examples. An encoder according to embodiments of the present application may, however, perform block-based encoding of a picture 12 using the concept outlined in more detail below and being different from the encoder of Fig. 1 such as, for instance, in that same is no video encoder, but a still picture encoder, in that same does not support inter-prediction, or in that the sub-division into blocks 80 is performed in a manner different than exemplified in Fig. 3. Likewise, decoders according to embodiments of the present application may perform block-based decoding of picture 12’ from data stream 14 using the coding concept further outlined below, but may differ, for instance, from the decoder 20 of Fig. 2 in that same is no video decoder, but a still picture decoder, in that same does not support intra-prediction, or in that same sub-divides picture 12’ into blocks in a manner different than described with respect to Fig. 3 and / or in that same does not derive the prediction residual from the data stream 14 in transform domain, but in spatial domain, for instance.
[0077] Fig. 4 illustrates an apparatus 20 for decoding a video 1 T from a data stream 14, also referred to as decoder 20, according to an embodiment of the invention. The video 1 T may comprise a sequence of pictures, e.g., picture 12’ and picture 12*’ illustrated in Fig. 4, wherein the apostrophe is used to indicate that the pictures reconstructed by decoder 20 may deviate from the original pictures, e.g., in terms of coding loss, e.g., quantization loss. In Fig. 4, picture 12’ represents a picture, which is currently decoded by decoder 20. For example, decoder 20 may decode the pictures of the video according to a coding order 21. Decoder 20 reconstructs picture 12’ in units of portions, into which the picture is subdivided, e.g., as described with respect to the portioning into prediction block with respect to Fig. 3 above.
[0078] In the following, a reconstruction process for reconstructing a predetermined portion 16 of picture 12’ performed by decoder 20 is described. The predetermined portion 16 may be the currently reconstructed portion of picture 16. In other words, the predetermined picture may be a picture of the sequence, for which decoder 20 performs the decoding as described with respect to picture 12’.
[0079] Decoder 20 of Fig. 4 comprises a decoding module 60, which obtains a prediction residual 61 of the predetermined portion 16. For example, decoding module 60 decodes a residual signal of the portion 16 from the data stream, e.g., by entropy decoding 50 and inverse transform 54 as described with respect to Fig. 2. In other words, decoding module 60 may optionally comprise entropy decoder 50, inverse quantizer 52 and inverse transformer 54, and prediction residual 61 may optionally correspond to prediction residual 24”” of Fig 2. In examples, the prediction residual 61 may be zero. In such cases, for example, decoding module 60 may derive from the data stream 14 that the prediction residual for the portion 16 is zero, e.g., by decoding a respective indication, e.g., in form of a syntax element and set the prediction residual zero.
[0080] Decoding module 60 further obtains prediction parameters 63 for the predetermined portion, e.g., by deriving the prediction parameters 63 from information signaled in the data stream.
[0081] For example, the prediction parameters 63 may indicate how to derive the prediction signal 74 and the extension prediction signal 72 from the previously reconstructed portion of the video. In embodiments in which the prediction is a temporal prediction, the prediction parameters may comprise, for example, an indication of a reference picture comprising the previously reconstructed portion based on which the prediction is performed, e.g., the previously reconstructed picture 12*’, and one or more motion parameters, e.g., motion vectors, which indicate a motion of the content of the picture 12’ with respect to the previously reconstructed portion based on which the prediction is performed.
[0082] Decoder 20 further comprises a prediction module 64, which uses the prediction parameters 63 of portion 16 to derive a prediction signal 74 for the portion 16 and an extension prediction signal 72 for the extension region 70 based on a previously reconstructed portion of the video. For example, prediction module 70 may use a previously reconstructed portion of the video for predicting the predetermined portion 16 and the extension region 70.
[0083] The previously reconstructed portion of the video, which forms the basis for deriving the prediction signal 74 and the extension prediction signal 72, may, e.g., depending on a coding mode or prediction mode of the portion 16, be a part of picture 12’ or may be part of one or more previously decoded pictures, such as picture 12*’ shown in Fig. 4. In other words, the prediction performed to derive prediction signal 74 and extension prediction signal 72 may be a temporal prediction (“inter-prediction”), according to which one or more pictures, which belong to different time stamps than picture 12’, but which have been decoded previous to picture 12’ (note that the decoding order may deviate from the temporal order), e.g., previously decoded portion 18’ of Fig. 4, are used for the prediction. In other examples, the prediction may be an intra-prediction, according to a previously reconstructed portion of picture 12’, e.g., portion 18 of Fig. 4, is used for the prediction. Also combinations are possible, in which the prediction is derived based on temporal prediction and intraprediction.
[0084] For example, decoder 20 decodes the portions, into which picture 12’ is subdivided, sequentially, according to a coding order of the portions within the picture. For example, the coding is performed row per row from top to bottom, and within one row, portion-wise from left to right, as in the illustrative example of Fig. 4, where the dotted region 18 of picture 12’ is reconstructed previous to the predetermined portion 16. Of course, other schemes are possible, e.g. from bottom-right to top-left.
[0085] The extension region 70 is located in the previously reconstructed portion 18 of picture 12’, i.e. , is part of the currently decoded picture to which the predetermined portion belongs.
[0086] For example, extension region 70 is an extension of the predetermined portion 16. The extension region 70 may be a simply connected region of the previously reconstructed portion 18 of picture 12’, or may, as exemplarily depicted in Fig. 4, be composed of multiple regions. The extension region 70 may directly adjoin the predetermined portion 16, but in other examples, it does not directly adjoin the previously reconstructed portion 18 but is separated therefrom.
[0087] For example, the extension region may spatially neighbor the predetermined portion in that the extension region adjoins the predetermined portion. In other examples, “neighboring” does not mean that the extension region necessarily adjoins the predetermined portion but the extension region may be spatially separated from the predetermined portion. For example, the extension region may comprise one or more rows I columns of samples which may adjoin the predetermined portion or which may be separated by one or more rows / columns of samples from the predetermined portion.
[0088] Decoder 20 further comprises a difference signal module 66, which derives a difference signal 78 based on the extension prediction signal 72 and based on reconstructed samples 77 of the extension region. Decoder 20 further comprises a refinement module 68, which derives a refined prediction signal 76 based on the prediction signal 74 for the predetermined portion and the difference signal 78.
[0089] Decoder 20 reconstructs, in operation 62 illustrated in Fig. 4, also referred to as reconstructor 62, the predetermined portion 16 based on the refined prediction signal 76 and the prediction residual 61 , e.g., by combining the refined prediction signal with a residual signal of the predetermined portion, e.g., to obtain reconstructed samples of the portion. For example, operation 62 may correspond to combiner 56 as described with respect to Fig. 2.
[0090] Decoder 20 of Fig. 4 may optionally implement some or all features of decoder 20 described with respect to Fig. 2, but decoder 20 of Fig. 4 may also differ from decoder 20 of Fig. 2. For example, modules 64, 66, and 68 of Fig. 4 may be part of a prediction entity 59 of decoder 20. Decoder 20 of Fig. 4 may correspond to decoder 20 of Fig. 2 in that prediction entity 59 of Fig. 4 may take the place of prediction module 58 of Fig. 2 may correspond to wherein prediction module 64 may optionally operate as described with respect to prediction module 58 of Fig. 2. It is noted, however, that the reconstruction process comprising the refinement of the prediction signal described with respect to Fig. 4 is not necessarily applied to all portions of the picture, but may be selectively applied to specific portions of the picture or video. There may also be pictures of the video, to which the prediction refinement is not applied at all.
[0091] Fig. 5 illustrates a cut-out of picture 12’ comprising the predetermined portion 16 according to an example. As illustrated in Fig. 5, picture 12’ may comprise an array of sample positions arranged in rows and columns of the array. For example, for reconstructing portion 16, decoder 20 may derive, for each sample position of the predetermined portion 16, a reconstructed sample value. Accordingly, for the previously reconstructed portion 18 of picture 12’, a reconstructed sample value may be available for each sample position at the time of reconstructing the predetermined portion 16.
[0092] For example, the prediction signal 74 may comprise a predicted sample value for each sample position of the predetermined portion. Similarly, the extension prediction signal 72 may comprise a predicted sample value for each sample position of the extension region 70.
[0093] According to an embodiment, difference signal module 66 derives for sample position 81 , which is part of the extension region 70, a weighted difference between the reconstructed sample value of the sample position 81 and the predicted sample value of the extension prediction signal 72 for the extension region, e.g., to obtain a difference sample value of the difference signal described with respect to Fig. 6.
[0094] In other words, according to an embodiment, the decoder 20 is configured for deriving the difference signal by deriving, for a sample position 81 of the extension region (e.g., for each sample position of the extension region), a weighted difference between a reconstructed sample value of the respective sample position (e.g., a sample value of the reconstructed portion of the picture at the respective sample position) and a predicted sample value of the extension prediction signal for the extension region (e.g., to obtain a difference sample value of the difference signal).
[0095] According to a further embodiment, difference signal module 66 may include neighboring sample position of the sample position 81 in the determination of the weighted difference at the sample position 81. For example, difference signal module 66 may derive a reconstruction value obtained from one or more reconstructed sample values positioned within a first region 82 around the sample position 81 . Additionally or alternatively, difference signal module 66 may derive a prediction value obtained from one or more predicted sample values of the extension prediction signal for the extension region being positioned within a second region. The second region may be equivalent to the first region 81 , or may be different. Difference signal module 66 may for example, average the sample values within the first and second regions, respectively, to obtain the reconstruction value and the prediction value for sample position 81 , respectively. Difference signal module 66 may then form a weighted difference of the reconstruction value and the prediction value for sample position 81.
[0096] In other words, according to an embodiment, decoder 20 is configured for deriving the difference signal by deriving, for a sample position 81 of the extension region (e.g., for each sample position of the extension region), a weighted difference between a reconstruction value obtained from one or more reconstructed sample values positioned within a region 82 (e.g., a first region) around the respective sample position 81 (e.g., by averaging the one or more reconstructed sample values of the region) and a prediction value obtained from one or more predicted sample values of the extension prediction signal for the extension region, the one or more predicted sample values being positioned within a region 82 (e.g., a second region, or the first region) around the respective sample position 81 (e.g., by averaging the one or more predicted sample values of the region).
[0097] According to an embodiment, the weighted difference is an equal-weighted difference.
[0098] According to another embodiment, weights of the weighted difference depend on the sample position, for which the weighted difference is derived.
[0099] According to an embodiment, decoder 20 is configured for deriving the refined prediction signal by forming, for a predetermined sample position 17 (or one of the sample positions, e.g., predetermined in terms of a currently considered one) of the predetermined portion 16 (e.g., for each sample position of the predetermined portion 16), a combination (e.g. a sum, e.g., a weighted sum or an equal weighted sum) of a sample value of the prediction signal for the predetermined sample position 17 and the output of a function of the difference signal.
[0100] According to an embodiment, the function depends on one or more of the predetermined sample position 17, a shape of the predetermined portion 16, a size of the predetermined portion 16, a coding mode for the predetermined portion 16, and a prediction mode for the predetermined portion 16 (e.g., a manner according to which the prediction signal is derived).
[0101] In other words, for deriving the refined prediction signal 76 at a predetermined sample position 17, refinement module 68 may form a combination of the predicted sample value of the prediction signal 74 at the predetermined sample position 17 and the output of a function, which receives the difference signal 78 as an input, at the predetermined sample position 17. In other words, the function may depend on the sample position. For example, the function may define respective contributions of sample values of the difference signal at different sample positions to the refined prediction signal. For example, in one embodiment, sample positions 81 and 83 may contribute to the refined prediction signal 76 for sample position 17, e.g., in form of difference signal values of the difference signal determined based on the predicted sample values of the extension prediction signal and the reconstructed sample values of the sample positions 81 and 83. However, as already described, further and / or other sample positions may contribute to the refined prediction signal for the predetermined sample position 17.
[0102] For example, the difference signal may be a signal comprising difference signal values associated with sample positions of the extension region 70.
[0103] According to an embodiment, decoder 20 is configured for deriving the refined prediction signal by forming, for a predetermined sample position 17 of the predetermined portion 16 (e.g., for each sample position of the predetermined portion 16), a combination of a sample value of the prediction signal for the predetermined sample position 17 and one or more difference signal values of the difference signal (E.g., a weighted sum of a sample value of the prediction signal for the predetermined sample position 17 and one or more difference signal values of the difference signal. E.g., the weighted sum comprises, or consists of, a sample value of the prediction signal for the predetermined sample position 17 and one or more difference signal values of the difference signal).
[0104] For example, decoder 20 is configured for for deriving the refined prediction signal by forming, for a predetermined sample position 17 of the predetermined portion 16 (e.g., for each sample position of the predetermined portion 16), a combination of a sample value of the prediction signal for the predetermined sample position 17 and a weighted sum of one or more difference signal values of the difference signal wherein the weighted sum comprises all difference signal values of the difference signal (e.g., difference signal values for all sample positions of the extension region).
[0105] Fig. 6 illustrates an embodiment of difference signal module 66 and refinement module 68. According to this embodiment, difference signal module 66 derives difference signal 78, which comprises for a sample position of the extension region, e.g., sample position 81 , a difference signal value. Difference signal module 66 may determine the difference signal value based on a weighted difference between one or more reconstructed samples of the extension region on the one hand and one or more predicted sample values of the extension prediction signal for the extension region on the other hand, e.g., as already described with respect to Fig. 5.
[0106] According to this embodiment, refinement module 68 comprises a refinement signal provider 91 to derive a refinement signal 79. According to this embodiment, the difference signal 78 is provided to the refinement signal provider 92, which determines, for a sample position of the predetermined portion 16, e.g., for the sample position 17, a refinement signal value of the refinement signal 79.
[0107] Thus, in examples, the difference signal 78 may have the dimensionality or size of the extension region, while the refinement signal 79 may have the dimensionality or size of the predetermined portion 16.
[0108] For example, refinement signal provider 92 may determine the refinement signal value of the refinement signal 79 for sample position 17 based on one or more difference signal values of the difference signal 78.
[0109] For example, refinement signal provider 92 may determine the refinement signal value of the refinement signal 79 for sample position 17 by combining, e.g. by forming a weighted sum (e.g., referred to as first weighted sum) of a plurality of difference signal values.
[0110] According to an embodiment, refinement signal provider 92 derives respective weights for the difference signal values of the first weighted sum in dependence on a sample position of the respective difference signal value and / or a shape of the predetermined portion and / or a size of the predetermined portion and / or an indication derived from the data stream. According to an embodiment, decoding module 60 derives a syntax element from the data stream, which indicates respective weights for the difference signal values of the first weighted sum. For example, the syntax element indicates one out of a set of sets of weights to be applied for the one or more difference signal values, or the syntax element indicates one out of a set of functions for deriving the weights to be applied for the one or more difference signal values, (E.g., the weights may depend on further parameters, such as position, block size, block shape).
[0111] For example, the syntax element is signaled in the data stream individually for the predetermined portion. For example, the apparatus derives the syntax element individually for portions, to which the prediction refinement is applied.
[0112] According to an embodiment, refinement signal provider 92 derives respective weights for the difference signal values of the first weighted sum in dependence on a distance of respective sample positions of the difference signal values from a reference position (e.g., a corner, e.g., a top left corner) of the predetermined portion.
[0113] According to an embodiment, refinement signal provider 92 derives respective weights for the difference signal values of the first weighted sum in a manner that a smaller or equal weight (e.g., in terms of absolute value is assigned to a difference signal value positioned at a higher distance from the reference position of the predetermined portion.
[0114] According to an embodiment, refinement signal provider 92 derives a weight for a difference signal value of the first weighted sum, which difference signal value has a sample position in a horizontally adjacent portion of the predetermined portion, in a manner that a smaller or equal weight (e.g., in terms of absolute value) is assigned to a difference signal value positioned at a higher distance in vertical direction from the reference position of the predetermined portion. Additionally or alternatively, refinement signal provider 92 derives a weight for a difference signal value of the first weighted sum, which difference signal value has a sample position in a vertically adjacent portion of the predetermined portion, in a manner that a smaller or equal weight (e.g., in terms of absolute value) is assigned to a difference signal value positioned at a higher distance in horizontal direction from the reference position of the predetermined portion. For example, weights may be chosen according to Fig. 8, where the abscissa indicates the distance from the reference position, e.g., the top left corner, in vertical or horizontal direction, respectively.
[0115] According to an embodiment, refinement signal provider 92 derives a weight for a difference signal value of the first weighted sum, which difference signal value has a sample position in a horizontally adjacent portion of the predetermined portion, using a first function, which depends on a distance in vertical direction of the sample position of the refinement sample value from the reference position of the predetermined portion in a manner that a smaller or equal weight (e.g., in terms of absolute value) is assigned to a difference signal value positioned at a higher distance in vertical direction from the reference position. Similarly, refinement signal provider 92 may derive a weight for a difference signal value of the first weighted sum, which difference signal value has a sample position in a vertically adjacent portion of the predetermined portion, using a second function, which depends on a distance in horizontal direction of the sample position of the refinement sample value from the reference position of the predetermined portion in a manner that a smaller or equal weight (e.g., in terms of absolute value) is assigned to a difference signal value positioned at a higher distance in horizontal direction from the reference position.
[0116] According to an embodiment, the first function equals the second function. Alternatively, the first function differs from the second function.
[0117] According to an embodiment, the first function and the second function depend on a size and / or a shape of the predetermined portion and / or a coding mode for the predetermined portion 16, and / or a prediction mode for the predetermined portion 16.
[0118] According to one example, the plurality of difference signal values comprises those having equal coordinates in the sample array as the predetermined sample position 17 in either the vertical or the horizontal direction such as sample positions 81 and 83, respectively, in Fig. 5. In other examples, further sample positions may contribute to the refinement signal value of sample position 17, e.g., in case that the extension region comprises further rows and / or columns, such that further sample positions with equal vertical or horizontal coordinate are available. In such case, the contributions of the different difference signal values may, for example, be weighted differently, e.g., depending on their position. In other examples, the plurality of difference signal value for the refinement signal value at sample position 17 may be defined differently. According to an embodiment, the difference signal refiner 66 derives difference signal 78 by deriving, for a sample position, such as position 81 , or position 83, of the extension region (e.g., for each sample position of the extension region), a difference signal value of the difference signal based on a weighted difference between one or more reconstructed samples of the extension region on the one hand and one or more predicted sample values of the extension prediction signal for the extension region on the other hand.
[0119] According to an embodiment, refinement module 68 may further comprise a prediction signal refiner 93, which derives a sample value of the refined prediction signal 76 at the predetermined sample position 17 based on a combination, e.g. a sum, e.g., a weighted sum or an equal weighted sum, of a predicted sample value of the prediction signal for the predetermined sample position and the refinement sample value of the refinement signal for the predetermined sample position.
[0120] Thus, according to an embodiment, refinement module 68 comprises refinement signal provider 92, which may derive, for a predetermined sample position 17 of the predetermined portion 16 (e.g., for each sample position of the predetermined portion 16), a refinement sample value of a refinement signal 79 based on (e.g., based on, e.g. by forming a weighted sum of) one or more difference signal values of the difference signal 78, and further, refinement module 68 comprises prediction signal refiner 93, which may derive a sample value of the refined prediction signal 76 at the predetermined sample position 17 based on a combination (e.g. a sum, e.g., a weighted sum or an equal weighted sum) of a predicted sample value of the prediction signal 74 for the predetermined sample position 17 and the refinement sample value for the predetermined sample position 17.
[0121] According to an embodiment, the refinement signal provider 92 derives the refinement sample value of the refinement signal by forming a first weighted sum of a plurality of difference signal values of the difference signal.
[0122] For example, weights of the first weighted sum depend on one or more of the predetermined sample position 17, a size of the predetermined portion 16, a shape of the predetermined portion 16, a coding mode of the predetermined portion 16, a prediction mode of the predetermined portion 16, and a dedicated syntax element derived from the data stream which indicates the weights of the weighted sum (e.g., a syntax element dedicated for the predetermined portion 16). According to an embodiment, prediction signal refiner 93 derives the refined prediction signal by combining the refinement signal and the prediction signal in a position-wise manner (e.g., in a manner according to which each predicted sample value of the prediction signal is combined with a respective collocated (e.g., located at the same sample position within the predetermined portion 16) refinement sample value of the refinement signal).
[0123] According to an embodiment, prediction signal refiner 93 derives the sample value of the refined prediction signal at the predetermined sample position 17 by forming a second weighted sum of the predicted sample value and the refinement sample value for the predetermined sample position 17.
[0124] For example, weights of the second weighted sum depend on one or more of the predetermined sample position 17, a size of the predetermined portion 16, a shape of the predetermined portion 16, a coding mode of the predetermined portion 16, a prediction mode of the predetermined portion 16, and a dedicated syntax element derived from the data stream which indicates the weights of the weighted sum (e.g., a syntax element dedicated for the predetermined portion 16).
[0125] For example, the one or more difference signal values for deriving the refinement signal value for the predetermined sample position 17 comprise (or consist of) one or more difference sample values which are equally positioned with respect to a horizontal direction as the predetermined sample position 17 of the predetermined portion 16, such as, for example, position 81 in Fig. 5 and / or one or more positions to the left of position 81 , and / or one or more difference sample values which are equally positioned with respect to a vertical direction as the predetermined sample position 17 of the predetermined portion 16, such as, for example, position 83 in Fig. 5 and / or one or more positions at the top of position 83.
[0126] According to an embodiment, the difference signal 78 is filtered and the refined prediction signal is derived using the filtered difference signal, e.g., instead of the difference signal 78.
[0127] For example, the filtering comprises a smoothing and / or a low-pass-filtering.
[0128] According to an embodiment, the decoder 20, e.g., the refinement module 68, sub-samples the difference signal and derives the refined prediction signal based on the sub-sampled difference signal. In the following, further optional features and details are described, which may be combined with any of the embodiments described above with respect to Fig. 1 to Fig. 6.
[0129] According to an embodiment, the video is a color video, and the picture comprises multiple channels, e.g., color channels, e.g., Y, Cb, Cr. In other words, the picture 12’, or each of the pictures of the sequence, may comprise multiple channels, or components, e.g., a luma and two chroma channels, or three color channels. For example, the picture comprises, for each of the channels, a respective sample array, which may be of equal size or of different sizes.
[0130] For example, decoder 20 may reconstruct each of the channels using predictive residual coding, e.g., at least for a portion of each of the channels.
[0131] According to an embodiment, decoder 20 reconstructs the channels of picture 12’ by, in reconstructing each of the channels, reconstructing a predetermined portion, e.g., a currently reconstructed portion, of the respective channel as follows. Decoding module 60 may obtain a prediction residual of the predetermined portion of the channel, e.g., by deriving the prediction residual from the data stream. Decoding module 60 may further obtain prediction parameters for the predetermined portion of the channel, by deriving the prediction parameters from the data stream. For example, the prediction parameters are specific to the channel. Prediction module 64 may then use the prediction parameters of the predetermined portion of the channel for deriving, based on a previously reconstructed portion of the channel, e.g., a previously reconstructed picture of the channel or a previously reconstructed portion of the current picture of the channel, a prediction signal for the predetermined portion of the channel. Reconstructor 62 may then reconstruct the predetermined portion of the channel based on the prediction signal and the prediction residual of the channel, e.g., to obtain reconstructed samples of the portion.
[0132] For example, decoder 20 may perform, for one or more or all of the multiple channels, the refinement of the prediction signal of a predetermined portion of the respective channel as described above with respect to the predetermined portion 16. In other words, the predetermined portion 16 may be a portion of a channel of a picture of the color video.
[0133] Accordingly, according to an embodiment, decoder 20 may refine, for one or more or all of the multiple channels, the prediction signal of the predetermined portion of the respective channel as follows: Prediction module 64 may use the prediction parameters obtained for the predetermined portion of the channel for deriving, based on the previously reconstructed portion of the channel, an extension prediction signal for an extension region of the predetermined portion of the channel which extension region spatially neighbors the predetermined portion of the channel, the extension region being part of a previously reconstructed portion of the channel. Difference signal module 66 may derive a difference signal (or correction signal) for the predetermined portion of the channel based on the extension prediction signal of the predetermined portion of the channel and based on reconstructed samples of the extension region of the predetermined portion of the channel. Refinement module 68 may refine the prediction signal of the predetermined portion of the channel using the difference signal of the predetermined portion of the channel. Reconstructor 62 may combine the refined prediction signal (e.g., instead of the prediction signal) for the predetermined portion of the channel with the residual signal of the predetermined portion 16 of the channel.
[0134] Any of the features and details described above with respect to the reconstruction of the predetermined portion 16 may optionally apply to the reconstruction of the predetermined portion of the channel.
[0135] According to an embodiment, decoder 30 performs the refinement of the respective prediction signals of the one or more or all channels independently from each other, e.g., independently from the refinement of the prediction signals of further channels.
[0136] Alternatively, decoder 20 may perform the refinement of the prediction signal of a predetermined one of the channels by deriving the difference signal 78 for the predetermined portion of the predetermined channel further based on the extension prediction signal derived for one or more further ones of the channels.
[0137] In other words, decoder may derive an extension prediction signal as described with respect to the predetermined portion 16 for each of respective portions of at least two of the channels. For example, the respective portions are co-located. That is, for example, the portions have corresponding or equivalent positions within respective sample arrays representing the channels.
[0138] According to an embodiment, the difference signal module 68 derives the difference signal for the predetermined portion of the predetermined channel further based on a difference signal and / or a prediction signal and / or a residual signal and / or a reconstructed signal derived for the one or more further ones of the channels.
[0139] In the following, further optional details and features of the extension region 70 are described, which may be implemented in combination with any of the previously described embodiments.
[0140] According to an embodiment, the extension region 70 comprises, or consists of, one or more of the following: one or more, e.g., one or two, columns of sample positions of a horizontally adjacent portion of the predetermined portion 16, the one or more columns being horizontally adjacent to the predetermined portion 16; see, for example, the portion of extension region 70 of Fig. 5, which is located to the left of portion 16. one or more, e.g., one or two, rows of sample positions of a vertically adjacent portion of the predetermined portion 16, the one or more rows being horizontally adjacent to the predetermined portion 16; see, for example, the portion of extension region 70 of Fig. 5, which is located to the top of portion 16. one or more samples of a horizontally and vertically adjacent portion of the predetermined portion 16; e.g., in Fig. 5, the extension region may additionally include the position, which shares the horizontal coordinate with position 81 , and which shares the vertical coordinate of position 83.
[0141] In the following, further optional details and features regarding the prediction performed by prediction module 64 to derive the prediction signal and the extension prediction signal are described, which may be implemented in combination with any of the previously described embodiments.
[0142] According to an embodiment, the previously reconstructed portion 18’ of the video based on which the prediction signal is derived, comprises, or is part of, one or more previously reconstructed pictures, e.g., as exemplarily indicated in Fig. 4.
[0143] In other words, prediction module 64 may derive the prediction signal 74, and optionally the extension prediction signal 72, for the predetermined portion 16 using temporal prediction.
[0144] In the following, further optional details and features regarding an activation or deactivation of the refinement of the prediction signal are described, which may be implemented in combination with any of the previously described embodiments, e.g., as illustrated as optional feature in form of switch 85 in Fig. 7 described later. It is noted, however, that the switching is independent of the specific implementation shown in Fig. 7, and may optionally be implemented in any of the embodiments.
[0145] For example, decoder 20 may decide between activating and deactivating the abovedescribed refinement of the prediction signal 74 obtained for a portion of the picture. For example, decoder 20 may switch between (i) reconstructing the portion using prediction signal 74 and (ii) deriving the refined prediction signal 76 for the portion and reconstructing the portion using the refined prediction signal. In other words, in (i), the prediction signal 74 may be input to reconstructor 62, while in (ii) the refined prediction signal 76 may be input to reconstructor 62 to be combined with the prediction residuum.
[0146] The decision may be performed on a per portion basis, e.g., individually for each of a set of portions of the picture.
[0147] In more generic terms, according to an embodiment, decoder 20 may activate or deactivate, e.g., individually, for each of portions of the picture, a refinement of a prediction signal of the respective portion. For each of portions, for which the refinement of the prediction signal is activated, reconstructor 62 may reconstruct the respective portion by combining a refined prediction signal of the respective portion with a prediction residual of the respective portion, and for each of portions, for which the refinement of the prediction signal is deactivated, reconstructor 62 may reconstruct the respective portion by combining the prediction signal of the respective portion with a prediction residual of the respective portion.
[0148] For example, decoder 20 may, for each of portions, for which the refinement of the prediction signal is activated, derive a refined prediction signal 74 as described above regarding the predetermined portion.
[0149] Several options may exist for decoder 20 to perform the decision whether to activate of deactivate the prediction refinement. For example, the decision may be signaled in the data stream 14. Alternatively, the decision may be predefined, e.g., in dependence on one or more criterions that relate to properties of the portion and / or a current coding situation, such as size and / or position of the portion within the picture, a prediction mode, a coding profile currently used. In an even further alternative, decoder 20 may infer the decision, e.g., from a reference portion of the currently coded portion. Thus, in more generic terms, decoder 20 may perform the decision according to one or more of the following: for each of a first set of portions of the picture, (e.g., individually) activating or deactivating the refinement of the prediction signal of the respective portion in dependence on an indication (e.g., specific to the respective portion) derived from the data stream, and / or for each of a second set of portions of the picture, activating the refinement of the prediction signal of the respective portion, and / or for each of a third set of portions of the picture, deactivating the refinement of the prediction signal of the respective portion, and / or for each of a fourth set of portions of the picture, inferring whether to activate or deactivate the refinement of the prediction signal of the respective portion, e.g., in dependence on a reference portion, e.g., a merge candidate, assigned to the respective portion.
[0150] For example, decoder 20 may assign a portion of the picture to one of the first set of portions, the second set of portions, the third set of portions, or the fourth set of portions based on a prediction mode assigned to the respective portion.
[0151] According to an embodiment, decoder 20 may perform the decision whether to activate or deactivate the prediction refinement for a portion in dependence on one or both of: an indication derived from the data stream (e.g., a flag, e.g., referred to as refine flag) which indicates whether or not to activate or deactivate the derivation of the refinement of the prediction signal for the respective portion, and / or a prediction mode of the respective portion (e.g., the prediction mode may be a mode according to which the prediction signal may be obtained) (e.g., the prediction modes may include a temporal prediction mode, e.g., referred to as inter-prediction mode, e.g., a regular inter-prediction mode, an affine inter-prediction mode, and a geometric mode, a merge mode, e.g., a regular merge mode, or a combined inter / intra prediction (CUP) mode, and or one or more intra-prediction modes).
[0152] For example, if the prediction mode of the respective portion is a first temporal prediction mode (e.g., a regular inter-prediction mode), decoder 20 may derive an indication from the data stream which indicates whether or not to activate or deactivate the derivation of the refinement of the prediction signal for the respective portion. Additionally or alternatively, if the prediction mode of the respective portion is a merge mode (e.g., a mode according to which prediction parameters for the respective block are inherited from a previously reconstructed portion, e.g. a specified previously reconstructed portion (e.g., referred to as specified merging candidate)), decoder 20 may derive an indication which indicates whether or not to activate or deactivate the derivation of the refinement of the prediction signal for the respective portion from a respective indication for a previously reconstructed portion (e.g. the specified previously reconstructed portion).
[0153] For example, if the prediction mode of the respective portion is not the first temporal prediction mode and not the merge mode, decoder 20 may deactivate the refinement of the prediction signal of the respective portion.
[0154] Fig. 7 illustrates a prediction refinement method according to an embodiment. The embodiment of Fig. 7 may optionally be combined with any of the features and details described above with respect to Fig. 1 to Fig. 6. Any of the features and details described with respect to Fig. 7 may optionally be implemented in combination with any of the embodiments described previously with respect to Fig. 1 to Fig. 6. An optional correspondence between the elements of Fig. 7 and the feature described above is given by the reference signs in Fig. 7.
[0155] In the following, one exemplary implementation of the prediction refinement method is described:
[0156] • The inter prediction signal of the current block is extended by one sample row to the top and one sample column to the left, using the same motion parameters (e.g., motion vectors and reference indices) as used for generating the inter prediction signal of the current block.
[0157] • These additionally generated prediction samples correspond to sample locations of adjacent, neighboring blocks for which there are already reconstructed samples available.
[0158] • For each spatial position in the additional row and column, the difference between the reconstructed sample value of the neighboring block and the sample value of the extended prediction signal is computed.
[0159] • Now, there is a top row and a left column, both consisting of the samples values of the difference signals. • A position-dependent weighting is applied separably to the sample values in the difference signals for the top row and the left column, resulting in a refinement offset signal. This refinement offset signal is then added to the prediction signal of the current block, resulting in the refined prediction signal.
[0160] • The position-dependent weighting factors can be, e.g., obtained from Fig. 8. For a prediction block of size l / l / x / - / , the following applies:
[0161] For each sample position 0 < x < W , 0 < y < H, the refinement offset is obtained as weightw[x] ■ left_difference_signal[y] + weightH[y] ■ top_difference_signal[x]
[0162] Here, weightwand weightHcorrespond to the respective curves of Fig. 8, depending on the values of l / l / and H.
[0163] • As illustrated in Fig. 7, it is indicated by a flag (refine flag) whether this refinement method is to be applied to a given block. For regular inter prediction mode (i.e. , not for merge mode and not for affine inter prediction), this flag is transmitted in the bitstream. For regular merge mode (i.e., not for geometric, affine, or combined inter / intra prediction [CUP]), the value of this flag is inherited from the specified merging candidate. Otherwise, the value of this flag is inferred to be equal to 0.
[0164] As mentioned, the foregoing specific implementation is exemplary.
[0165] More generally, embodiments of the prediction refinement method can be described as follows:
[0166] • A prediction signal is spatially enlarged / extended to also cover picture samples which have already been processed, and for which, consequently, their recon- structed sample values are already available. For example, the motion-compensated prediction signal is extended to the top / left by one sample row / column, by using the same motion parameters as used for generating the prediction signal of the block, but applying it to the enlarged spatial region.
[0167] • Now, for each spatial position in the enlarged region, two sample values are available: The reconstructed sample values r (of a preceding block) and the enlarged prediction samples p (corresponding to the current block). Out of these two sample values, a correction signal is derived. In its simplest form, the correction signal may be the difference d = r - p.
[0168] • The correction signal is blended into the current block, thus refining the prediction signal of the current block: - In its simplest form, the blending is an addition of a refinement signal to the prediction signal of the current block.
[0169] - The refinement signal may be obtained, e.g. as a position-dependent weighting of the top / left correction signal, or more generally, as the result of applying for each position inside the current block a function, which may depend on the position and the block shape, to the correction signal which yields the value of the refinement signal at that position.
[0170] - This position-dependent weighting will be typically of such a kind that positions which are closer to the top and left border have a higher weight than positions far apart (e.g., in the bottom / right region). One particular example for such a kind of weighting is the one used by PDPC [4] in the VVC stan- dard [1], possibly with minor modifications, e.g. in the way how the weights are derived depending on the block shape.
[0171] • Note that in contrast to established methods, such as CUP (combined inter / intra prediction) [2] in VVC, not the neighboring reconstructed signal itself, but a refinement signal derived thereof is blended into the current block.
[0172] In the following, variants of some of the features are described, which may be implemented individually or in any combination.
[0173] • The application of the refinement process is either explicitly signalled per block via a new syntax element or the application condition may be (implicitly) derived, e.g. from the prediction mode.
[0174] • A combination is also possible: For certain prediction modes (e.g., “regular” in- ter prediction, sometimes also called AMVP mode [2]), it is explicitly signalled, whereas for other modes (e.g., affine or geometric prediction modes) it is always disabled, for a third class of modes (e.g., LIC mode [5]) it is always applied, and for a fourth class of modes (e.g., merge modes [2]) it is implicitly derived (e.g., from the respective merging candidate) whether it is to be applied or not.
[0175] • In its simplest form, the correction signal is derived as d = r - p. However, a more general form is also possible, like: d = wr■ r-wp■ p. Here, for the so-called simplest case, wr= wp= 1. In this more general form, it is also possible to use a different d, i.e. different combinations of wrand wp, for each position within the current block. In this case, consequently, there would be a different correction signal for each combination of wrand wp. • In its simplest form, the prediction signal is enlarged by one sample row / column, but it is also possible to use a larger enlarged region and derive the correc- tion / refinement signal thereof. In particular, something akin to MRL (multiple reference line intra prediction) [4] as in VVC may be used here. Alternatively, a correction signal with one row / column may be achieved by averaging the values of two or more rows / columns of rand p, respectively.
[0176] • In a more general form, the position dependent weighting consists of a matrix of weights with size (num. samples d)x(num. samples block), i.e. the refinement value for each position within the current block is a weighted sum of all values of the correction signal d.
[0177] • The refinement may be applied to all the color channels (e.g., Y, Cb, Cr) or only a subset thereof.
[0178] • For different blocks, different kinds of position-dependent weightings may be used. The selection among these different weightings may be either implicitly (e.g., depending on the block shape or prediction mode) or explicitly (e.g., by a new syntax element that indicates the used weighting).
[0179] • Prior to generating the refinement signal out of the correction signal, a filtering (e.g., smoothing or low-pass filtering) or subsampling (e.g. for larger blocks) may be applied to the correction signal.
[0180] • The top-left corner point / area may also be included in the correction signal. In the example described above, the prediction signal is only extended to the top and to the left, leaving this area out.
[0181] • The prediction signal which is extended may comprise uni-prediction, bi-prediction, and multi-hypothesis prediction, as well as any subset thereof.
[0182] • In the simplest case, the refined prediction signal is obtained as a sample-wise addition of the refinement signal and the initial (unrefined) prediction signal. However, variants thereof are also possible. The refined prediction signal may also be a weighted sum of the refinement signal and the initial (unrefined) prediction signal.
[0183] The weights may be spatially varying over the prediction block and / or may be depending on, e.g., the block shape, the prediction mode. In a more general case, the refined prediction signal is the output of applying a (possibly spatially varying) function to the refinement signal and the initial (unrefined) prediction signal.
[0184] Fig. 8 shows a diagram, which indicates, for different block width / heights, position dependent weights. For example, the position on the abscissa may refer to the horizontal or vertical coordinate of the sample position in the block, counted from top left corner in horizontal direction or vertical direction, respectively.
[0185] Fig. 9 illustrates an apparatus 10 for encoding a video 11 into a data stream 14, also referred to as encoder 10, according to an embodiment of the invention. Encoder 10 may be the counterpart of decoder 20 of Fig. 4, e.g., encoder 10 encodes the video 11 into the data stream 14 to be decoded by decoder 20. Accordingly, encoder 10 encodes the video in units of portions, using the same partitioning as decoder 20.
[0186] Encoder 10 encodes the predetermined portion 16 of the picture 12 by the deriving prediction parameters 63 for the predetermined portion 16. For example, encoder 10 encodes the prediction parameters into the data stream 14. Encoder 10 performs a prediction of the predetermined portion by means of prediction entity 59, which may operate equivalently as the prediction entity 59 of decoder 20. E.g., the relationship between encoder 10 of Fig. 9 and decoder 20 of Fig. 4 may be as described with respect to encoder 10 and decoder 20 of Fig. 1 and Fig. 2. Encoder 10 uses the refined prediction signal 76 provided by prediction entity 59 for deriving a prediction residual 57, see operation 97 of Fig. 9, which may operate like subtractor 22 of Fig. 1 . Predication entity 59 may operate as described with respect to prediction entity 59 of Fig. 4. In the optional block 99 illustrated in Fig. 9 the prediction residual 57 may be encoded into data stream 14, e.g., by transforming, quantizing, and entropy encoding the prediction residual. Block 99 may further provide a previously encoded portion of the video for prediction entity 59, e.g., as will be described in the following with respect to Fig. 10.
[0187] For example, encoder 10 may derive the prediction parameters 63 by selecting the prediction parameters from a plurality of coding option, e.g., by evaluating the plurality of coding options, e.g., with respect to a rate-distortion relation.
[0188] Fig. 10 illustrates another embodiment of encoder 10, which may be based on encoder 10 of Fig. 9. In other words, Fig. 10 may provide further optional details of encoder 10 of Fig. 9.
[0189] For example, encoder 10 comprises an encoding module 95, which encodes the prediction residual into the data stream 14, e.g. by block-bases transform coding and quantization, e.g. as described with respect to transformer 28 and quantizer 23 of Fig. 1. In examples, encoder comprises an entropy encoder as described with respect to Fig. 1 to encode the prediction residual into the data stream, e.g. as described with respect to entropy encoder 34 of Fig. 1.
[0190] Encoder 10 may further comprise decoding module 60’ to reconstruct the encoded prediction residual of a previously encoded portion of the video, such as portions 18 or 18’ described with respect to Fig. 4, to obtain a reconstructed prediction residual 61 of a previously encoded portion of the video. Combiner 62 may combine the reconstructed prediction residual with a prediction signal of this portion to reconstruct the previously encoded portion, which is to be provided to prediction entity 59 for performing the prediction of the predetermined portion 16. For example, decoding module 60’ may perform the inverse operation of encoding module 95, e.g. dequantization and inverse transform, e.g. as described with respect to Fig. 1 .
[0191] It is clear that whenever, in the description of decoder 20, the derivation of an information or indication from data stream 14 is described, a corresponding embodiment of encoder 10 inserts this information or indication into data stream 14.
[0192] It is noted that the block diagrams of Fig. 4 and Fig. 9 may alternatively be considered as illustrations of respective methods, in which each of the blocks represents a step of the respective method.
[0193] Thus, what is further disclosed with respect to Fig. 4 is a method for decoding a video 11 from a data stream 14, the method comprising reconstructing a picture 12’ of the video in units of portions, into which the picture is subdivided, wherein the method comprises a step of reconstructing a predetermined portion 16 of the picture, wherein the step of reconstructing the predetermined portion comprises the following steps: obtaining a prediction residual 61 of the predetermined portion; obtaining prediction parameters 63 for the predetermined portion 16, and using the prediction parameters for deriving 64, based on a previously reconstructed portion 18, 18’ of the video, a prediction signal 74 for the predetermined portion 16 and an extension prediction signal 72 for an extension region 70 which spatially neighbors the predetermined portion 16, the extension region being part of a previously reconstructed portion 18 of the picture; deriving 66 a difference signal 78 based on the extension prediction signal 72 and based on reconstructed samples 77 of the extension region; deriving 68 a refined prediction signal 76 for the predetermined portion 16 based on the prediction signal 74 for the predetermined portion 16 and the difference signal 78; and reconstructing 62 the predetermined portion 16 based on the refined prediction signal 76 and the prediction residual 61.
[0194] What is further disclosed with respect to Fig. 9 is a method for encoding a video 11 into a data stream 14, wherein the method comprises encoding a picture 12 of the video in units of portions, into which the picture is subdivided, wherein the method comprises a step of encoding a predetermined portion 16 of the picture. The step of encoding the predetermined portion 16 comprises the following steps: deriving 60 prediction parameters 63 for the predetermined portion 16, and using the prediction parameters for deriving 64, based on a previously encoded portion 18, 18’ of the video, a prediction signal 74 for the predetermined portion 16 and an extension prediction signal 72 for an extension region 70 which spatially neighbors the predetermined portion 16, the extension region being part of a previously encoded portion 18 of the picture; deriving 66 a difference signal 78 based on the extension prediction signal 72 and based on previously encoded and reconstructed samples 77 of the extension region; deriving 68 a refined prediction signal 76 for the predetermined portion 16 based on the prediction signal 74 for the predetermined portion 16 and the difference signal 78; and using the refined prediction signal to derive a prediction residual 61 of the predetermined portion 16.
[0195] Although some aspects have been described as features in the context of an apparatus it is clear that such a description may also be regarded as a description of corresponding features of a method. Although some aspects have been described as features in the context of a method, it is clear that such a description may also be regarded as a description of corresponding features concerning the functionality of an apparatus.
[0196] Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, one or more of the most important method steps may be executed by such an apparatus.
[0197] The inventive encoded image signal can be stored on a digital storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet. In other words, further embodiments provide a video data stream product including the video data stream according to any of the herein described embodiments, e.g. a digital storage medium having stored thereon the video data stream. Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software or at least partially in hardware or at least partially in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
[0198] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0199] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier.
[0200] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0201] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0202] A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non- transitory.
[0203] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet. A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0204] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0205] A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
[0206] In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.
[0207] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0208] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0209] In the foregoing Detailed Description, it can be seen that various features are grouped together in examples for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed examples require more features than are expressly recited in each claim. Rather, as the following claims reflect, subject matter may lie in less than all features of a single disclosed example. Thus the following claims are hereby incorporated into the Detailed Description, where each claim may stand on its own as a separate example. While each claim may stand on its own as a separate example, it is to be noted that, although a dependent claim may refer in the claims to a specific combination with one or more other claims, other examples may also include a combination of the dependent claim with the subject matter of each other dependent claim or a combination of each feature with other dependent or independent claims. Such combinations are proposed herein unless it is stated that a specific combination is not intended. Furthermore, it is intended to include also features of a claim to any other independent claim even if this claim is not directly made dependent to the independent claim.
[0210] The above-described embodiments are merely illustrative for the principles of the present disclosure. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the pending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.
[0211] References
[0212] [1] B. Bross, Y.-K. Wang, Y. Ye, S. Liu, J. Chen, G. J. Sullivan, and J.-R. Ohm, “Overview of the Versatile Video Coding (VVC) standard and its applications,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31 , no. 10, pp. 3736-3764, 2021.
[0213] [2] W.-J. Chien, L. Zhang, M. Winken, X. Li, R.-L. Liao, H. Gao, C.-W. Hsu, H. Liu, and C.-C. Chen, “Motion vector coding and block merging in the Versatile Video Coding standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31 , no. 10, pp. 3848-3861 , 2021.
[0214] [3] H. Yang, H. Chen, J. Chen, S. Esenlik, S. Sethuraman, X. Xiu, E. Alshina, and J. Luo, “Subblock-based motion derivation and inter prediction refinement in the Versatile Video Coding standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31 , no. 10, pp. 3862-3877, 2021.
[0215] [4] J. Pfaff, A. Filippov, S. Liu, X. Zhao, J. Chen, S. De-Lux'an-Hern'andez, T. Wiegand, V. Rufitskiy, A. K. Ramasubramonian, and G. Van der Auwera, “Intra prediction and mode coding in VVC,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31 , no. 10, pp. 3834-3847, 2021.
[0216] [5] H. Liu, Y. Chen, J. Chen, L. Zhang, and M. Karczewicz, “Local illumination compensation,” document VCEG-AZ06, ITU-T Q.6 / SG 16 (VCEG), 2015.
Claims
Claims1. Apparatus (20) for decoding a video (1 T) from a data stream (14), configured for reconstructing a picture (12’) of the video in units of portions, into which the picture is subdivided, wherein the apparatus is configured for reconstructing a predetermined portion (16) of the picture by obtaining a prediction residual (61) of the predetermined portion (16), obtaining prediction parameters (63) for the predetermined portion (16), and using the prediction parameters for deriving (64), based on a previously reconstructed portion (18, 18’) of the video, a prediction signal (74) for the predetermined portion (16) and an extension prediction signal (72) for an extension region (70) which spatially neighbors the predetermined portion (16), the extension region being part of a previously reconstructed portion (18) of the picture, deriving (66) a difference signal (78) based on the extension prediction signal (72) and based on reconstructed samples (77) of the extension region, deriving (68) a refined prediction signal (76) for the predetermined portion (16) based on the prediction signal (74) for the predetermined portion (16) and the difference signal (78), and reconstructing (62) the predetermined portion (16) based on the refined prediction signal (76) and the prediction residual (61).
2. Apparatus according to claim 1 , configured for deriving the difference signal by deriving, for a sample position (81) of the extension region], a weighted difference between a reconstructed sample value of the respective sample position and a predicted sample value of the extension prediction signal for the extension region.
3. Apparatus according to claim 1 or 2, configured for deriving the difference signal by deriving, for a sample position (81) of the extension region, a weighted difference between a reconstruction value obtained from one or more reconstructed sample values positioned within a region (82) around the respective sample position (81)and a prediction value obtained from one or more predicted sample values of the extension prediction signal for the extension region, the one or more predicted sample values being positioned within a region (82) around the respective sample position (81).
4. Apparatus according to claim 2 or 3, wherein the weighted difference is an equal- weighted difference.
5. Apparatus according to claim 2 or 3, wherein weights of the weighted difference depend on the sample position, for which the weighted difference is derived.
6. Apparatus according to any of the claims 1 to 5, configured for deriving the refined prediction signal by forming, for a predetermined sample position (17) of the predetermined portion (16), a combination of a sample value of the prediction signal for the predetermined sample position (17) and the output of a function of the difference signal.
7. Apparatus according to claim 6, wherein the function depends on one or more of the predetermined sample position (17), a shape of the predetermined portion (16), a size of the predetermined portion (16), a coding mode for the predetermined portion (16), and a prediction mode for the predetermined portion (16).
8. Apparatus according to any of the claims 1 to 7, configured for deriving the refined prediction signal by forming, for a predetermined sample position (17) of the predetermined portion (16), a combination of a sample value of the prediction signal for the predetermined sample position (17) and one or more difference signal values of the difference signal.
9. Apparatus according to claim 8, configured for deriving the refined prediction signal by forming, for a predetermined sample position (17) of the predetermined portion (16), a combination of a sample value of the prediction signal for the predetermined sample position (17) and a weighted sum of one or more difference signal values of the difference signal wherein the weighted sum comprises all difference signal values of the difference signal.
10. Apparatus according to claim 8 or 9, configured for deriving weights according to which each of the one or more difference signal values of the combination contribute to the refined prediction signal at the predetermined sample position (17) in dependence on a sample position of the respective difference signal value and / or a shape of the predetermined portion (16) and / or a size of the predetermined portion(16) and / or an indication derived from the data stream.
11. Apparatus according to any of claims 8 to 10, configured for deriving a syntax element from the data stream, which indicates weights according to which each of the one or more difference signal values of the combination contribute to the refined prediction signal at the predetermined sample position (17).
12. Apparatus according to claim 11 , wherein the syntax element is signaled in the data stream individually for the predetermined portion (16).
13. Apparatus according to any of claims 8 to 12, configured for deriving weights, according to which each of the one or more difference signal values of the combination contribute to the refined prediction signal at the predetermined sample position (17), in dependence on a distance of respective sample positions of the difference signal values from a reference position of the predetermined portion (16).
14. Apparatus according to claim 13, configured for deriving weights, according to which each of the one or more difference signal values of the combination contribute to the refined prediction signal at the predetermined sample position (17), in a manner that a smaller or equal weight is assigned to a difference signal value positioned at a higher distance from the reference position of the predetermined portion (16).
15. Apparatus according to claim 13 or 14, configured for deriving a weight according to which a difference signal value of the combination contributes to the refined prediction signal at the predetermined sample position(17), which difference signal value has a sample position in a horizontally adjacent portion of the predetermined portion (16), in a manner that a smaller or equal weight is assigned to a difference signal value positioned at a higher distance in vertical direction from the reference position of the predetermined portion (16), and / orderiving a weight according to which a difference signal value of the combination contributes to the refined prediction signal at the predetermined sample position (17), which difference signal value has a sample position in a vertically adjacent portion of the predetermined portion (16), in a manner that a smaller or equal weight is assigned to a difference signal value positioned at a higher distance in horizontal direction from the reference position of the predetermined portion (16).
16. Apparatus according to claim 13 or 14, configured for deriving a weight according to which a difference signal value of the combination contributes to the refined prediction signal at the predetermined sample position (17), which difference signal value has a sample position in a horizontally adjacent portion of the predetermined portion (16), using a first function, which depends on a distance in vertical direction of the sample position of the difference signal value from the reference position of the predetermined portion (16) in a manner that a smaller or equal weight is assigned to a difference signal value positioned at a higher distance in vertical direction from the reference position, and deriving a weight according to which a difference signal value of the combination contributes to the refined prediction signal at the predetermined sample position (17), which difference signal value has a sample position in a vertically adjacent portion of the predetermined portion (16), using a second function, which depends on a distance in horizontal direction of the sample position of the difference signal value from the reference position of the predetermined portion (16) in a manner that a smaller or equal weight is assigned to a difference signal value positioned at a higher distance in horizontal direction from the reference position, wherein the first function equals the second function, or wherein the first function differs from the second function.
17. Apparatus according to claim 16, wherein the first function and the second function depend on one or more of a size, a shape of the predetermined portion (16), a coding mode for the predetermined portion (16), and a prediction mode for the predetermined portion (16).
18. Apparatus according to any of the claims 1 to 17, configured for deriving the refined prediction signal by deriving (66) the difference signal (78) by deriving, for a sample position (81 , 83) of the extension region, a difference signal value of the difference signal based on a weighted difference between one or more reconstructed samples of the extension region on the one hand and one or more predicted sample values of the extension prediction signal for the extension region on the other hand.
19. Apparatus according to claim 18, configured for deriving the refined prediction signal by deriving (92), for a predetermined sample position (17) of the predetermined portion (16), a refinement sample value of a refinement signal (79) based on one or more difference signal values of the difference signal (78), and deriving (93) a sample value of the refined prediction signal (76) at the predetermined sample position (17) based on a combination of a predicted sample value of the prediction signal (74) for the predetermined sample position (17) and the refinement sample value for the predetermined sample position (17).
20. Apparatus according to claim 19, configured for deriving the refinement sample value of the refinement signal by forming a first weighted sum of a plurality of difference signal values of the difference signal, wherein weights of the first weighted sum depend on one or more of the predetermined sample position (17), a size of the predetermined portion (16), a shape of the predetermined portion (16), a coding mode of the predetermined portion (16), a prediction mode of the predetermined portion (16), and a dedicated syntax element derived from the data stream which indicates the weights of the weighted sum.
21. Apparatus according to claim 19 or 20, configured for deriving the refined prediction signal by combining the refinement signal and the prediction signal in a positionwise manner.
22. Apparatus according to any of claims 19 to 21 , configured for deriving the sample value of the refined prediction signal at the predetermined sample position (17) by forming a second weighted sum of the predicted sample value and the refinement sample value for the predetermined sample position (17).
23. Apparatus according to claim 22, wherein weights of the second weighted sum depend on one or more of the predetermined sample position (17), a size of the predetermined portion (16), a shape of the predetermined portion (16), a coding mode of the predetermined portion (16), a prediction mode of the predetermined portion (16), and a dedicated syntax element derived from the data stream which indicates the weights of the weighted sum.
24. Apparatus according to any of claims 19 to 23, wherein the one or more difference signal values for deriving the refinement signal value for the predetermined sample position (17) comprise one or more difference sample values which are equally positioned with respect to a horizontal direction as the predetermined sample position (17) of the predetermined portion (16), and / or one or more difference sample values which are equally positioned with respect to a vertical direction as the predetermined sample position (17) of the predetermined portion (16).
25. Apparatus according to any of the claims 1 to 24, configured for filtering the difference signal and deriving the refined prediction signal based on the filtered difference signal.
26. Apparatus according to claim 25, wherein the filtering comprises a smoothing and / or a low-pass-filtering.
27. Apparatus according to any of the claims 1 to 26, configured for sub-sampling the difference signal and deriving the refined prediction signal based on the subsampled difference signal.
28. Apparatus according to any of the claims 1 to 27, wherein the video is a color video, wherein the picture comprises multiple channels, and wherein the apparatus is configured for reconstructing each of the channels by reconstructing a predetermined portion (16) of the respective channel by obtaining a prediction residual of the predetermined portion (16) of the channel, obtaining prediction parameters for the predetermined portion (16) of the channel, and using the prediction parameters of the predetermined portion (16) of the channel for deriving, based on a previously reconstructed portion of the channel, a prediction signal for the predetermined portion (16) of the channel, and reconstructing the predetermined portion (16) of the channel based on the prediction signal and the prediction residual of the channel, wherein the apparatus is configured for performing, for one or more or all of the multiple channels, a refinement of the prediction signal of the predetermined portion (16) of the respective channel by using the prediction parameters obtained for the predetermined portion (16) of the channel for deriving, based on the previously reconstructed portion of the channel, an extension prediction signal for an extension region of the predetermined portion (16) of the channel which extension region spatially neighbors the predetermined portion (16) of the channel, the extension region being part of a previously reconstructed portion of the channel, deriving a difference signal for the predetermined portion (16) of the channel based on the extension prediction signal of the predetermined portion (16) of the channel and based on reconstructed samples of the extension region of the predetermined portion (16) of the channel, and refining the prediction signal of the predetermined portion (16) of the channel using the difference signal of the predetermined portion (16) of the channel.
29. Apparatus according to claim 28, configured to perform the refinement of the respective prediction signals of the one or more or all channels independently from each other.
30. Apparatus according to claim 28, configured for performing the refinement of the prediction signal of a predetermined one of the channels by deriving the difference signal for the predetermined portion (16) of the predetermined channel further based on the extension prediction signal derived for one or more further ones of the channels31. Apparatus according to claim 30, configured for deriving the difference signal for the predetermined portion (16) of the predetermined channel further based on a difference signal and / or a prediction signal and / or a residual signal and / or a reconstructed signal derived for the one or more further ones of the channels.
32. Apparatus according to any of the claims 1 to 31 , wherein the extension region comprises one or more columns of sample positions of a horizontally adjacent portion of the predetermined portion (16), the one or more columns being horizontally adjacent to the predetermined portion (16), and / or one or more rows of sample positions of a vertically adjacent portion of the predetermined portion (16), the one or more rows being horizontally adjacent to the predetermined portion (16), and / or one or more samples of a horizontally and vertically adjacent portion of the predetermined portion (16).
33. Apparatus according to any of the claims 1 to 32, wherein the previously reconstructed portion of the video, comprises one or more previously reconstructed pictures.
34. Apparatus according to any of the claims 1 to 33, configured for deriving the prediction signal for the predetermined portion (16) using temporal prediction.
35. Apparatus according to any of the claims 1 to 34, configured for activating or deactivating, for each of portions of the picture, a refinement of a prediction signal of the respective portion, and for each of portions, for which the refinement of the prediction signal is activated, reconstructing the respective portion by combining a refined prediction signal of the respective portion with a prediction residual of the respective portion, and for each of portions, for which the refinement of the prediction signal is deactivated, reconstructing the respective portion by combining the prediction signal of the respective portion with a prediction residual of the respective portion.
36. Apparatus according to claim 35, configured for for each of portions, for which the refinement of the prediction signal is activated, deriving an extension prediction signal for an extension region which spatially neighbors the respective portion, the extension region being part of a previously reconstructed portion of the picture, deriving a difference signal based on the extension prediction signal and based on reconstructed samples of the extension region, and deriving the refined prediction signal for the respective portion based on the prediction signal for the portion and the difference signal.
37. Apparatus according to claim 35 or 36, configured for for each of a first set of portions of the picture, activating or deactivating the refinement of the prediction signal of the respective portion in dependence on an indication derived from the data stream, and / or for each of a second set of portions of the picture, activating the refinement of the prediction signal of the respective portion, and / orfor each of a third set of portions of the picture, deactivating the refinement of the prediction signal of the respective portion, and / or for each of a fourth set of portions of the picture, inferring whether to activate or deactivate the refinement of the prediction signal of the respective portion.
38. Apparatus according to claim 37, configured for assigning a portion of the picture to one of the first set of portions, the second set of portions, the third set of portions, or the fourth set of portions based on a prediction mode assigned to the respective portion.
39. Apparatus according to claim 35 or 36, configured for activating or deactivating, for each of portions of the picture the refinement of the prediction signal of the respective portion in dependence on an indication derived from the data stream which indicates whether or not to activate or deactivate the derivation of the refinement of the prediction signal for the respective portion, and / or a prediction mode of the respective portion.
40. Apparatus according to claim 39, configured for if the prediction mode of the respective portion is a first temporal prediction mode, derive an indication from the data stream which indicates whether or not to activate or deactivate the derivation of the refinement of the prediction signal for the respective portion.
41. Apparatus according to claim 39 or 40, configured for if the prediction mode of the respective portion is a merge mode, deriving an indication which indicates whether or not to activate or deactivate the derivation of the refinement of the prediction signal for the respective portion from a respective indication for a previously reconstructed portion.
42. Apparatus according to claim 40 or 41 , configured for if the prediction mode of the respective portion is not the first temporal prediction mode and not the merge mode, deactivate the refinement of the prediction signal of the respective portion.
43. Apparatus (10) for encoding a video (11) into a data stream (14), configured for encoding a picture (12) of the video in units of portions, into which the picture is subdivided, wherein the apparatus is configured for encoding a predetermined portion (16) of the picture by deriving prediction parameters (63) for the predetermined portion (16), and using the prediction parameters for deriving (64), based on a previously encoded portion (18, 18’) of the video, a prediction signal (74) for the predetermined portion (16) and an extension prediction signal (72) for an extension region (70) which spatially neighbors the predetermined portion (16) (16), the extension region being part of a previously encoded portion (18) of the picture, deriving (66) a difference signal (78) based on the extension prediction signal (72) and based on previously encoded and reconstructed samples (77) of the extension region, deriving (68) a refined prediction signal (76) for the predetermined portion (16) based on the prediction signal (74) for the predetermined portion (16) and the difference signal (78), and using the refined prediction signal to derive a prediction residual (61) of the predetermined portion (16).
44. Apparatus according to claim 43, configured for deriving the difference signal by deriving, for a sample position (81) of the extension region], a weighted difference between a reconstructed sample value of the respective sample position and a predicted sample value of the extension prediction signal for the extension region.
45. Apparatus according to claim 43 or 44, configured for deriving the difference signal by deriving, for a sample position (81) of the extension region, a weighted difference between a reconstruction value obtained from one or more reconstructed sample values positioned within a region (82) around the respective sample position (81) and a prediction value obtained from one or more predicted sample values of the extension prediction signal for the extension region, the one or more predicted sample values being positioned within a region (82) around the respective sample position (81).
46. Apparatus according to claim 44 or 45, wherein the weighted difference is an equal- weighted difference.
47. Apparatus according to claim 44 or 45, wherein weights of the weighted difference depend on the sample position, for which the weighted difference is derived.
48. Apparatus according to any of the claims 43 to 47, configured for deriving the refined prediction signal by forming, for a predetermined sample position (17) of the predetermined portion (16), a combination of a sample value of the prediction signal for the predetermined sample position (17) and the output of a function of the difference signal.
49. Apparatus according to claim 48, wherein the function depends on one or more of the predetermined sample position (17), a shape of the predetermined portion (16) a size of the predetermined portion (16), a coding mode for the predetermined portion (16), and a prediction mode for the predetermined portion (16).
50. Apparatus according to any of the claims 43 to 49, configured for deriving the refined prediction signal by forming, for a predetermined sample position (17) of the predetermined portion (16), a combination of a sample value of the prediction signal for the predetermined sample position (17) and one or more difference signal values of the difference signal.
51. Apparatus according to claim 50, wherein the weighted sum comprises all difference signal values of the difference signal.
52. Apparatus according to claim 50 or 51 , configured for deriving weights according to which each of the one or more difference signal values of the combination contribute to the refined prediction signal at the predetermined sample position (17) in dependence on a sample position of the respective difference signal value and / or a shape of the predetermined portion (16) and / or a size of the predetermined portion(16).
53. Apparatus according to any of claims 50 to 52, configured for encoding a syntax element into the data stream, which indicates weights according to which each of the one or more difference signal values of the combination contribute to the refined prediction signal at the predetermined sample position (17).
54. Apparatus according to claim 53, wherein the syntax element is signaled in the data stream individually for the predetermined portion (16).
55. Apparatus according to any of claims 50 to 54, configured for deriving weights, according to which each of the one or more difference signal values of the combination contribute to the refined prediction signal at the predetermined sample position (17), in dependence on a distance of respective sample positions of the difference signal values from a reference position of the predetermined portion (16).
56. Apparatus according to claim 55, configured for deriving weights, according to which each of the one or more difference signal values of the combination contribute to the refined prediction signal at the predetermined sample position (17), in a manner that a smaller or equal weight is assigned to a difference signal value positioned at a higher distance from the reference position of the predetermined portion (16).
57. Apparatus according to claim 55 or 56, configured for deriving a weight according to which a difference signal value of the combination contributes to the refined prediction signal at the predetermined sample position(17), which difference signal value has a sample position in a horizontally adjacent portion of the predetermined portion (16), in a manner that a smaller or equal weight is assigned to a difference signal value positioned at a higher distance in vertical direction from the reference position of the predetermined portion (16), and / orderiving a weight according to which a difference signal value of the combination contributes to the refined prediction signal at the predetermined sample position (17), which difference signal value has a sample position in a vertically adjacent portion of the predetermined portion (16), in a manner that a smaller or equal weight is assigned to a difference signal value positioned at a higher distance in horizontal direction from the reference position of the predetermined portion (16).
58. Apparatus according to claim 55 or 56, configured for deriving a weight according to which a difference signal value of the combination contributes to the refined prediction signal at the predetermined sample position (17), which difference signal value has a sample position in a horizontally adjacent portion of the predetermined portion (16), using a first function, which depends on a distance in vertical direction of the sample position of the difference signal value from the reference position of the predetermined portion (16) in a manner that a smaller or equal weight is assigned to a difference signal value positioned at a higher distance in vertical direction from the reference position, and deriving a weight according to which a difference signal value of the combination contributes to the refined prediction signal at the predetermined sample position (17), which difference signal value has a sample position in a vertically adjacent portion of the predetermined portion (16), using a second function, which depends on a distance in horizontal direction of the sample position of the difference signal value from the reference position of the predetermined portion (16) in a manner that a smaller or equal weight is assigned to a difference signal value positioned at a higher distance in horizontal direction from the reference position, wherein the first function equals the second function, or wherein the first function differs from the second function.
59. Apparatus according to claim 58, wherein the first function and the second function depend on one or more of a size, a shape of the predetermined portion (16), a coding mode for the predetermined portion (16), and a prediction mode for the predetermined portion (16).
60. Apparatus according to any of the claims 43 to 59, configured for deriving the refined prediction signal by deriving the difference signal by deriving, for a sample position of the extension region, a difference signal value of the difference signal based on a weighted difference between one or more reconstructed samples of the extension region on the one hand and one or more predicted sample values of the extension prediction signal for the extension region on the other hand.
61. Apparatus according to claim 60, configured for deriving the refined prediction signal by deriving, for a predetermined sample position (17) of the predetermined portion (16), a refinement sample value of a refinement signal based on one or more difference signal values of the difference signal, and deriving a sample value of the refined prediction signal at the predetermined sample position (17) based on a combination of a predicted sample value of the prediction signal for the predetermined sample position (17) and the refinement sample value for the predetermined sample position (17).
62. Apparatus according to claim 61 , configured for deriving, the refinement sample value of the refinement signal by forming a first weighted sum of a plurality of difference signal values of the difference signal, wherein weights of the first weighted sum depend on one or more of the predetermined sample position (17), a size of the predetermined portion (16), a shape of the predetermined portion (16), a coding mode of the predetermined portion (16), a prediction mode of the predetermined portion (16).
63. Apparatus according to claim 61 or 62, configured for deriving the refined prediction signal by combining the refinement signal and the prediction signal in a positionwise manner.
64. Apparatus according to any of claims 61 to 63, configured for deriving the sample value of the refined prediction signal at the predetermined sample by forming asecond weighted sum of the predicted sample value and the refinement sample value for the predetermined sample position (17).
65. Apparatus according to claim 64, wherein weights of the second weighted sum depend on one or more of the predetermined sample position (17), a size of the predetermined sample portion, a shape of the predetermined portion (16), a coding mode of the predetermined portion (16), a prediction mode of the predetermined portion (16).
66. Apparatus according to any of claims 61 to 65, wherein the one or more difference signal values for deriving the refinement signal value for the predetermined sample position (17) comprise one or more difference sample values which are equally positioned with respect to a horizontal direction as the predetermined sample position (17) of the predetermined portion (16), and / or one or more difference sample values which are equally positioned with respect to a vertical direction as the predetermined sample position (17) of the predetermined portion (16).
67. Apparatus according to any of the claims 43 to 66, configured for filtering the difference signal and deriving the refined prediction signal based on the filtered difference signal.
68. Apparatus according to claim 67, wherein the filtering comprises a smoothing and / or a low-pass-filtering.
69. Apparatus according to any of the claims 43 to 68, configured for sub-sampling the difference signal and deriving the refined prediction signal based on the subsampled difference signal.
70. Apparatus according to any of the claims 43 to 69, wherein the video is a color video, wherein the picture comprises multiple channels, and wherein the apparatus is configured for encoding each of the channels by encoding a predetermined portion (16) of the respective channel by obtaining prediction parameters for the predetermined portion (16) of the channel, and using the prediction parameters of the predetermined portionchannel, a prediction signal for the predetermined portion (16) of the channel, and using the prediction signal to derive a prediction residual (61) of the predetermined portion (16), wherein the apparatus is configured for performing, for one or more or all of the multiple channels, a refinement of the prediction signal of the predetermined portion (16) of the respective channel by using the prediction parameters obtained for the predetermined portion (16) of the channel for deriving, based on the previously encoded portion of the channel, an extension prediction signal for an extension region of the predetermined portion (16) of the channel which extension region spatially neighbors the predetermined portion (16) of the channel, the extension region being part of a previously encoded portion of the channel, deriving a difference signal for the predetermined portion (16) of the channel based on the extension prediction signal of the predetermined portion (16) of the channel and based on reconstructed samples of the extension region of the predetermined portion (16) of the channel, and refining the prediction signal of the predetermined portion (16) of the channel using the difference signal of the predetermined portion (16) of the channel.
71. Apparatus according to claim 70, configured to perform the refinement of the respective prediction signals of the one or more or all channels independently from each other.
72. Apparatus according to claim 70, configured for performing the refinement of the prediction signal of a predetermined one of the channels by deriving the difference signal for the predetermined portion (16) of the predetermined channel further based on the extension prediction signal derived for one or more further ones of the channels73. Apparatus according to claim 72, configured for deriving the difference signal for the predetermined portion (16) of the predetermined channel further based on a difference signal and / or a prediction signal and / or a residual signal and / or a reconstructed signal derived for the one or more further ones of the channels.
74. Apparatus according to any of the claims 43 to 73, wherein the extension region comprises one or more columns of sample positions of a horizontally adjacent portion of the predetermined portion (16), the one or more columns being horizontally adjacent to the predetermined portion (16), and / or one or more rows of sample positions of a vertically adjacent portion of the predetermined portion (16), the one or more rows being horizontally adjacent to the predetermined portion (16), and / or one or more samples of a horizontally and vertically adjacent portion of the predetermined portion (16).
75. Apparatus according to any of the claims 43 to 74, wherein the previously encoded portion of the video, comprises one or more previously encoded pictures.
76. Apparatus according to any of the claims 43 to 75, configured for deriving the prediction signal for the predetermined portion (16) using temporal prediction.
77. Apparatus according to any of the claims 43 to 76, configured for activating or deactivating, for each of portions of the picture, a refinement of a prediction signal of the respective portion, and for each of portions, for which the refinement of the prediction signal is activated, encoding the respective portion by using a refined prediction signal of the respective portion to derive a prediction residual of the respective portion, andfor each of portions, for which the refinement of the prediction signal is deactivated, encoding the respective portion by using the prediction signal of the respective portion to derive a prediction residual of the respective portion.
78. Apparatus according to claim 77, configured for for each of portions, for which the refinement of the prediction signal is activated, deriving an extension prediction signal for an extension region which spatially neighbors the respective portion, the extension region being part of a previously encoded portion of the picture, deriving a difference signal based on the extension prediction signal and based on reconstructed samples of the extension region, and deriving the refined prediction signal for the respective portion based on the prediction signal for the portion and the difference signal.
79. Apparatus according to claim 77 or 78, configured for for each of a first set of portions of the picture, encoding an indication into the data stream, which indicates whether to activate or deactivate the refinement of the prediction signal of the respective portion, and / or for each of a second set of portions of the picture, activating the refinement of the prediction signal of the respective portion, and / or for each of a third set of portions of the picture, deactivating the refinement of the prediction signal of the respective portion, and / or for each of a fourth set of portions of the picture, inferring whether to activate or deactivate the refinement of the prediction signal of the respective portion.
80. Apparatus according to claim 79, configured for assigning a portion of the picture to one of the first set of portions, the second set of portions, the third set of portions, orthe fourth set of portions based on a prediction mode assigned to the respective portion.
81. Apparatus according to claim 77 or 78, configured for activating or deactivating, for each of portions of the picture the refinement of the prediction signal of the respective portion in dependence on a prediction mode of the respective portion.
82. Apparatus according to claim 81 , configured for if the prediction mode of the respective portion is a first temporal prediction mode, encode an indication into the data stream which indicates whether or not to activate or deactivate the derivation of the refinement of the prediction signal for the respective portion.
83. Apparatus according to claim 81 or 82, configured for if the prediction mode of the respective portion is a merge mode, deriving an indication which indicates whether or not to activate or deactivate the derivation of the refinement of the prediction signal for the respective portion from a respective indication for a previously encoded portion.
84. Apparatus according to claim 82 or 83, configured for if the prediction mode of the respective portion is not the first temporal prediction mode and not the merge mode, deactivate the refinement of the prediction signal of the respective portion.
85. Method for decoding a video (11) from a data stream (14), the method comprising reconstructing a picture (12’) of the video in units of portions, into which the picture is subdivided, wherein the method comprises reconstructing a predetermined portion (16) of the picture by obtaining a prediction residual (61) of the predetermined portion, obtaining prediction parameters (63) for the predetermined portion (16), and using the prediction parameters for deriving (64), based on a previouslyreconstructed portion (18, 18’) of the video, a prediction signal (74) for the predetermined portion (16) and an extension prediction signal (72) for an extension region (70) which spatially neighbors the predetermined portion (16), the extension region being part of a previously reconstructed portion (18) of the picture, deriving (66) a difference signal (78) based on the extension prediction signal (72) and based on reconstructed samples (77) of the extension region, deriving (68) a refined prediction signal (76) for the predetermined portion (16) based on the prediction signal (74) for the predetermined portion (16) and the difference signal (78), and reconstructing (62) the predetermined portion (16) based on the refined prediction signal (76) and the prediction residual (61).
86. Method for encoding a video (11) into a data stream (14), wherein the method comprises encoding a picture (12) of the video in units of portions, into which the picture is subdivided, wherein the method comprises encoding a predetermined portion (16) of the picture by deriving prediction parameters (63) for the predetermined portion (16), and using the prediction parameters for deriving (64), based on a previously encoded portion (18, 18’) of the video, a prediction signal (74) for the predetermined portion (16) and an extension prediction signal (72) for an extension region (70) which spatially neighbors the predetermined portion (16), the extension region being part of a previously encoded portion (18) of the picture, deriving (66) a difference signal (78) based on the extension prediction signal (72) and based on previously encoded and reconstructed samples (77) of the extension region, deriving (68) a refined prediction signal (76) for the predetermined portion (16) based on the prediction signal (74) for the predetermined portion (16) and the difference signal (78), andusing the refined prediction signal to derive a prediction residual (61) of the predetermined portion (16).
87. A data stream comprising a video, the video being encoded into the data stream using the method of claim 86.
88. A computer program for implementing the method of claim 85 or 86 when being executed on a computer or signal processor.