Block-based prediction encoding and decoding of images

By performing spectral decomposition, noise reduction, and synthesis on the neighborhood of the prediction block, an improved prediction padding version is generated and transmitted as a signal notification in the data stream. This solves the problem of low coding efficiency in video codecs and achieves more efficient data compression and synchronization.

CN115190301BActive Publication Date: 2026-02-13FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210663788.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-01-05
Filing Date
2017-12-20
Publication Date
2026-02-13
Estimated Expiration
2037-12-20

AI Technical Summary

Technical Problem

Existing video codecs have room for improvement in coding efficiency in block-based predictive coding, especially in terms of the amount of data required to maintain predictive synchronization between the encoder and decoder.

Method used

By utilizing previously encoded or reconstructed versions of the neighborhood of a predetermined block, spectral decomposition, noise reduction, and spectral synthesis are performed to generate an improved version of the prediction padding. A signal is transmitted in the data stream to notify the selection of the appropriate prediction padding version. Combining spatial prediction and entropy coding, coding efficiency is optimized.

Benefits of technology

It improves coding efficiency, reduces prediction residual errors, enhances synchronization between the encoder and decoder, and improves data compression performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115190301B_ABST
    Figure CN115190301B_ABST
Patent Text Reader

Abstract

The previous encoded or reconstructed version of the neighborhood of the predetermined block to be predicted is used in order to enable more efficient predictive encoding of the predicted block. In particular, a spectral decomposition of a region consisting of a first version of the predicted padding of the predetermined block and of the neighborhood yields a first spectrum, the first spectrum is subjected to noise reduction, and the second spectrum resulting therefrom can be subjected to spectral synthesis, yielding a modified version of the region, which comprises a second version of the predicted padding of the predetermined block. The second version of the predicted padding of the predetermined block tends to improve the encoding efficiency thanks to the use of the processed (i.e. encoded / reconstructed) neighborhood of the predetermined block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application “Block-based prediction encoding and decoding of images” with application number 201780087979.2 and filing date 20 December 2017. TECHNICAL FIELD

[0002] The present application relates to block-based prediction encoding and decoding of images, for example suitable for hybrid video codecs. BACKGROUND

[0003] Many video codecs and still image codecs nowadays use block-based prediction encoding to compress data representing image content. The better the prediction, the less data is needed to encode the prediction residual. The overall benefit of using prediction depends on the amount of data needed to keep the prediction synchronized between the encoder and the decoder, i.e. the data needed for prediction parameterization. SUMMARY

[0004] It is an object of the present application to provide a concept for block-based prediction encoding / decoding of images that enables improved coding efficiency.

[0005] This object is achieved by the subject matter of the independent claims of the present application.

[0006] A basic finding of the present application is that a previous encoded or reconstructed version of a neighborhood of a predetermined block to be predicted can be exploited in order to enable more efficient prediction encoding of the predicted block. In particular, a spectral decomposition of a region consisting of a first version of the prediction padding of the predetermined block and the neighborhood yields a first spectrum, the first spectrum is subjected to denoising, and a second spectrum resulting therefrom can be subjected to spectral synthesis, yielding a modified version of the region comprising a second version of the prediction padding of the predetermined block. The second version of the prediction padding of the predetermined block tends to improve coding efficiency due to the exploitation of the processed, i.e. encoded / reconstructed, neighborhood of the predetermined block.

[0007] According to embodiments of the present application, a first signaling can be spent in a data stream in order to select between using the first version of the prediction padding and the second version of the prediction padding. Although this first signaling requires an additional amount of data, the ability to select between the first version of the prediction padding and the second version of the prediction padding can improve coding efficiency. The first signaling can be conveyed within the data stream at sub-image granularity, such that the selection between the first version and the second version can be made at sub-image granularity.

[0008] Likewise, additionally or alternatively, according to another embodiment of the present application, a second signalization can be spent in the data stream, the second signalization being used for setting the size of the neighborhood used for extending the predetermined block and forming the region with respect to which the spectral decomposition, the noise reduction and the spectral synthesis are performed. The second signal can likewise be conveyed within the data stream in a manner that varies in sub-picture granularity.

[0009] Even further, additionally or alternatively, a further signalization, a third signalization, can be conveyed within the data stream, the third signalization signaling the strength of the noise reduction, e.g. by indicating a threshold to be applied to the first spectrum resulting from the spectral decomposition. The third signal can likewise be conveyed within the data stream in a manner that varies in sub-picture granularity.

[0010] The second and / or third signalization can be encoded into the data stream using spatial prediction and / or using entropy coding using a spatial context, i.e. using an estimation of the probability distribution for the possible signalization values that depends on the spatial neighborhood of the region for which the respective signalization is contained in the data stream. BRIEF DESCRIPTION OF DRAWINGS

[0011] Advantageous implementations of the present application are subject of dependent claims. Preferred embodiments of the present application are described below with reference to the accompanying drawings, in which:

[0012] Figure 1 A block diagram of an encoding apparatus according to embodiments of the present application is shown;

[0013] Figure 2 A schematic diagram showing an image containing a block to be predicted using an illustration on the right hand side for the block currently to be predicted is shown according to embodiments, wherein about how the block currently to be predicted is extended in order to result in a region, which is then the starting point for implementing an alternative version of the prediction padding for the block;

[0014] Figure 3 A schematic diagram of a noise reduction according to embodiments is shown, wherein in particular two alternative ways of performing the noise reduction using a threshold are shown;

[0015] Figure 4 A schematic diagram showing a selection among possible strengths of noise reduction according to embodiments is shown; and

[0016] Figure 5 A block diagram of a decoding apparatus according to embodiments is shown, which is adapted to the apparatus of Figure 1 DETAILED DESCRIPTION

[0017] Figure 1 An apparatus 10 for encoding an image 12 into a data stream 14 on a block prediction basis is shown. Figure 1 ​The apparatus comprises a prediction provider 16, a spectral decomposer 18, a noise reducer 20, a spectral synthesizer 22 and an encoding stage 24. In the manner described in more detail below, these components 16 to 24 are connected in series in the order in which they are mentioned to the prediction loop of the encoder 10. For illustration purposes, Figure 1 It is indicated that the encoding stage 24 can internally comprise an adder 26, a transformer 28 and a quantization stage 30, which are connected in series in the order in which they are mentioned in the aforementioned prediction loop. In particular, the non-inverting input of the adder 26 is connected, directly or indirectly via a selector (as further outlined below), to the output of the spectral synthesizer 22, while the non-inverting input of the adder 26 receives the signal to be encoded (i.e. the image 12). As further shown in Figure 1 The encoding stage 24 can further comprise, as further shown in, an entropy encoder 32 connected between the output of the quantizer 30 and the output of the apparatus 10 at which the encoded data stream 14 representing the image 12 is output. As further shown in Figure 1 The apparatus 10 can comprise, as further shown in, a reconstruction stage 34 connected along the aforementioned prediction loop between the encoding stage 24 and the prediction provider 16, which provides the prediction provider 16 with previously encoded portions (i.e. portions of the image 12 or a video to which the image 12 belongs) that have been previously encoded by the encoder 10 as well as with a version of these portions that can be reconstructed at the decoder side, in particular even taking into account the encoding losses introduced by the quantization within the quantization stage 30. As shown in Figure 1 The reconstruction stage 34 can comprise, as shown in, a dequantizer 36, an inverse transformer 38 and an adder 40, which are connected in series to the aforementioned prediction loop in the order in which they are mentioned, with the input of the dequantizer being connected to the output of the quantizer. In particular, the output of the adder 40 is connected to the input of the prediction provider 16 and, in addition, there is a further input of the spectral decomposer 18 in addition to the input of the spectral decomposer 18 being connected to the output of the prediction provider 16, as set out in more detail below. While the first input of the adder 40 is connected to the output of the inverse transformer 38, the further input of the adder 40 receives the final prediction signal directly or optionally indirectly via the output of the spectral synthesizer 22. As shown in Figure 1 Optionally, the encoder 10 comprises, as shown in, a selector 42 configured to select between applying the prediction signal output by the spectral synthesizer 22 to the respective input of the adder 40 or to the output of the prediction provider 16.

[0018] Having illustrated the internal structure of the encoder 10, it should be noted that the implementation of the encoder 10 can be done in software, firmware or hardware or any combination thereof. Thus, Figure 1Any of the blocks or modules illustrated in the figures can correspond to a portion of a computer program running on a computer, a portion of firmware (e.g., a field-programmable array), or a portion of an electronic circuit (e.g., an application-specific IC).

[0019] Figure 1 The apparatus 10 is configured to encode the image 12 into the data stream 14 using block-based prediction. Thus, on the basis of blocks, the prediction provider 16 and the subsequent modules 18, 20 and 22 operate. Figure 2 The blocks 46 in the image 12 are prediction blocks. However, the apparatus 10 can also operate on the basis of blocks for other tasks. For example, the residual coding performed by the encoding stage 24 can also be performed on the basis of blocks. However, the prediction blocks on which the prediction provider 16 operates can be different from the residual blocks on which the encoding stage 24 operates. That is, the image 12 can be subdivided into prediction blocks differently than it is subdivided into residual blocks. For example, but not exclusively, the subdivision into residual blocks can represent an extension of the subdivision into prediction blocks such that each residual block is a part of a corresponding prediction block or coincides with a certain prediction block but does not cover onto neighboring prediction blocks. Furthermore, the prediction provider 16 can use different coding modes in order to perform its prediction and the switching between these modes can occur in blocks which can be referred to as coding blocks, which can also be different from the prediction blocks and / or the residual blocks. For example, the subdivision of the image 12 into coding blocks can be such that each prediction block covers only one corresponding coding block but can be smaller than the corresponding coding block. The coding modes just mentioned can include spatial prediction modes and temporal prediction modes.

[0020] In order to further illustrate the functioning or operating mode of the apparatus 10, reference is made to Figure 2 which shows that the image 12 can be an image belonging to a video, i.e. one of the images 44 in a temporal sequence, but it should be noted that this is merely illustrative and the apparatus 10 can also be applicable to still images 12. Figure 2Specifically, the prediction block 46 within the image 12 is indicated. The prediction block should be the block for which the prediction provider 16 currently performs a prediction. For the prediction block 46, the prediction provider 16 uses the previously encoded part of the image 12 and / or the video 44, or, alternatively, the part that the decoder can already reconstruct from the data stream 14 when trying to perform the same prediction for the block 46. To this end, the prediction provider 16 uses the reconstructable version, i.e. the version that can also be reconstructed at the decoder side. Different coding modes can be used. For example, the prediction provider 16 can predict the block 46 by temporal prediction, e.g. motion-compensated prediction, based on the reference image 48. Alternatively, the prediction provider 16 can use spatial prediction for predicting the block 46. For example, the prediction provider 16 can extrapolate a previously encoded neighborhood of the block 46 along a certain extrapolation direction into the interior of the block 46. In case of motion-compensated prediction, a motion vector can be signaled to the decoder as a prediction parameter for the block 46 within the data stream 14. Likewise, in case of spatial prediction as a prediction parameter, an extrapolation direction for the block 46 can be signaled to the decoder within the data stream 14.

[0021] That is, the prediction provider 16 outputs a prediction padding of the predetermined block 46. The prediction padding is at 48 in Figure 1 It is actually a first version 48 of the prediction padding, since this version will be "improved" by a subsequent series of components 18, 20 and 22, as further explained below. That is, the prediction provider 16 predicts a prediction sample value for each sample 50 within the block 46, which prediction padding represents the first version 48.

[0022] As shown in Figure 1 The block 46 can be rectangular or even square, as shown. However, this is merely an example and should not be regarded as a limiting alternative embodiment of the present application.

[0023] The spectral decomposer is configured to spectrally decompose a region 52 consisting of the first version 48 of the prediction padding for the block 46 and an extension thereof, i.e. a previously encoded version of the neighborhood 54 of the block 46. That is, in a geometric manner, the spectral decomposer 20 performs a spectral decomposition on the region 52, which comprises, in addition to the block 46, the neighborhood 54 of the block 46, wherein the part of the region 52 corresponding to the block 46 is filled with the first version 48 of the prediction padding of the block 46 and the neighborhood 54 is filled with sample values that can be reconstructed at the decoding side from the data stream 14. The spectral decomposer 18 receives the prediction padding 48 from the prediction provider 16 and the reconstructed sample values for the neighborhood 54 from the reconstruction stage 34.

[0024] For example, Figure 2shows the encoding of image 12 into data stream 14 by means of device 10 according to a certain encoding / decoding order 58. For example, this encoding order may traverse image 12 from the upper left corner to the lower right corner. The traversal may run row by row as shown in Figure 2 , but may alternatively run column by column or diagonally. However, all these examples are merely illustrative and should not be considered restrictive. Due to this general propagation of the encoding / decoding order from the upper left to the lower right of image 12, for most prediction blocks 46, the neighborhood of this block 46 (which is at the top of block 46 or adjacent to the top edge 461 and at the left of block 46 or adjacent to the left edge 464) is encoded into data stream 14 and is reconstructed from data stream 14 at the decoding side when performing encoding / reconstruction on block 46. Thus, in the example of Figure 2 , neighborhood 54 represents the spatial extension of block 46 outside edges 461 and 464 of block 46, thus describing an L-shaped region which, together with block 46, forms a rectangular region 52, the lower side and the right side of rectangular region 52 coinciding or being collinear with the left side 462 and the bottom side 463 of block 46.

[0025] That is to say, spectral decomposer 18 performs spectral decomposition on the sample array corresponding to region 52, where the samples corresponding to neighborhood 54 are sample values that can be reconstructed from data stream 14 using the prediction residuals encoded into data stream 14 through encoding stage 24, while the samples within block 46 in region 52 are the sample values of prediction fill 48 of prediction provider 16. The spectral decomposition (i.e., its transform type) that spectral decomposer 18 performs on this region 52 can be DCT, DST or wavelet transform. Optionally but not exclusively, the transform T2 used by spectral decomposer 18 can be of the same type as the transform T1 used by transformer 28 to transform the prediction residuals into the spectral domain (as the output of subtractor 28). If they are of the same type, spectral decomposer 18 and transformer 28 can share a certain circuit and / or computer code responsible or designed to perform this type of transform. However, alternatively, the transforms performed by spectral decomposer 18 and transformer 28 can be different.

[0026] Therefore, the output of spectral decomposer 18 is a first spectrum 60. Spectrum 60 can be an array of spectral coefficients. For example, the number of spectral coefficients can be equal to the number of samples within region 52. As long as it involves the spatial frequency along the horizontal axis x, the spatial frequency to which the spectral component belongs can increase column by column from left to right, and as long as it involves the spatial frequency within region 52 along the y-axis, it increases from top to bottom. However, it should be noted that as an alternative to the above example, T2 can be "overcomplete", such that the number of transform coefficients produced by T2 can even be greater than the number of samples within region 52.

[0027] The noise reducer 20 then performs noise reduction on the spectrum 60 to obtain a second, or noise-reduced, spectrum 62. Examples will be provided below on how the noise reduction can be performed by the noise reducer 20. In particular, the noise reducer 20 can involve thresholding of the spectral coefficients. Spectral coefficients below a certain threshold can be set to zero, or can be shifted towards zero by an amount equal to the threshold. However, all these examples are merely illustrative, and there are many alternatives with respect to performing noise reduction on the spectrum 60 to yield the spectrum 62.

[0028] The spectrum decomposer 22 then performs an inverse process of the spectrum decomposition performed by the spectrum decomposer 18. That is, the spectrum synthesizer 22 uses an inverse transform compared to the spectrum decomposer 18. As a result of the spectrum synthesis, which can alternatively be referred to as synthesis, the spectrum synthesizer 22 outputs a second version of the prediction padding for the block 46 (indicated by the hatching at 64 in Figure 1 It should be understood that the spectrum synthesis by the synthesizer 22 yields a modified version of the entire region 52. However, as indicated by the different hatching for the block 46 on the one hand and the neighborhood 54 on the other hand, only the portion corresponding to the block 46 is of interest, and forms the second version of the prediction padding of the block 46. It is indicated by the cross-hatching in Figure 1 The spectrum synthesis in the neighborhood 54 in the region 52 is shown using simple hatching in Figure 1 may not even be computed by the spectrum synthesizer 22. It is briefly noted here that the “inverse” has been satisfied when the spectrum decomposition T2 -1 is a left inverse of T2 (i.e. T2 -1 • T2 = 1). That is, a two-sided inverse is not required. For example, in addition to the above transformation examples, T2may be a shearlet transform or a contourlet transform. It is advantageous if all basis functions of T2extend over the entire region 52, or if all basis functions cover at least a large portion of the region 52, regardless of the transformation type used for T2.

[0029] It is worth noting that due to the fact that the spectrum 64 has already been obtained by transforming, noise-reducing and re-transforming the region 52, which also encompasses the already encoded, as well as the reconstructable version involved on the decoding side, the second version of the prediction padding 64 is very likely to yield a lower prediction error, and thus can represent an improved predictor for finally encoding the block 46 into the data stream 14 by the encoding stage 24 (i.e. for performing residual encoding).

[0030] As mentioned above, the selector 42 can optionally be present in the encoder 10. If not, the second version 64 of the prediction padded block 46 inevitably represents the final prediction of the block 46 into the inverting input of the subtractor 26, and thus the subtractor 26 computes the prediction residual or prediction error by subtracting the final prediction from the actual content of the image 12 within the block 46. Then, the encoding stage 24 transforms this prediction residual into the spectral domain, where the quantizer 30 performs quantization on the individual spectral coefficients representing this prediction residual. In addition, the entropy encoder 32 entropy encodes these quantized coefficient levels into the data stream 14. As mentioned above, due to the encoding / decoding order 58, the spectral coefficients of the prediction residual within the neighborhood 54 are already present within the data stream 14 before the prediction of the block 46. The quantizer 36 and the inverse transformer 28 recover the prediction residual of the block 46 in a version that can also be reconfigurable at the decoding side, and the adder 40 adds this prediction residual to the final prediction, resulting in a reconstructed version of the encoded portion, which, as mentioned above, also includes the neighborhood 54, i.e. a reconstructed version 56 of the neighborhood 54, which is used to fill a portion of the region 52, which is then subjected to spectral decomposition via the spectral decomposer 18.

[0031] However, if the optional selector 42 is present, the selector 42 can perform a selection between the first version 48 and the second version 64 of the prediction padded block 46, and use either of these two versions as the final prediction into the respective inputs of the subtractor 26 and the adder, respectively.

[0032] The way in which the second or refined version 64 of the prediction padding for the block 46 is derived from the blocks 18 to 22 and the optional selector 42 can be parameterizable for the encoder 10. That is, the encoder 10 can parameterize this way with regard to one or more of the following options, wherein the parameterization is signaled to the decoder by a respective signalization. For example, the encoder 10 can decide to select the version 48 or 64 and signal the result of the selection by a signalization 70 in the data stream. Also, the granularity at which the selection 70 is performed can be sub-picture granularity and can be done, for example, in regions or blocks into which the picture 12 is subdivided. In particular, the encoder 10 can perform the selection individually for each prediction block (e.g., the block 46) and signal the selection for each such prediction block by a signalization 70 in the data stream 14. A simple flag can be signaled in the data stream 14 for each block (e.g., the block 46). The signalization 70 in the data stream 14 can be encoded using spatial prediction. For example, the flag can be spatially predicted for a neighboring block in the neighboring blocks 46 based on the signalization 70 contained in the data stream 14. Additionally or alternatively, the signalization 70 can be encoded into the data stream using context adaptive entropy coding. The context used for entropy coding the signalization 70 for a certain block 46 into the data stream 14 can be determined from properties contained in the data stream 14 for neighboring blocks 46 (e.g., the signalization 70 signaled in the data stream 14 for such neighboring blocks).

[0033] Additionally or alternatively, another parameterization option for the encoder 10 can be the size of the region 52 or, alternatively, the size of the neighborhood 54. For example, the encoder 10 can set the position of a corner in the region 52, which is opposite to a corner 74 of the block 46 that is co-located with a corresponding corner of the block 46. The signalization 72 can indicate the position of this corner 76 or the size of the region 52, respectively, by an index into a list of available corner positions or sizes, respectively. The corner position can be indicated relative to the top-left corner of the block 46, i.e., as a vector relative to a corner in the block 46 that is opposite to the corner shared between the region 52 and the block 46. The setting of the size of the region 52 can also be done by the device 10 at sub-picture granularity (e.g., regions or blocks into which the picture 12 is subdivided, wherein these regions or blocks can coincide with the prediction blocks, i.e., the encoder 10 can perform the setting of the size of the region 52 for each block 46 individually). The signalization 72 can be encoded into the data stream 14 using prediction coding as explained with regard to the signalization 70 and / or using context adaptive entropy coding with a spatial context similar to the signalization 70.

[0034] As an alternative or additional scheme to signal notifications 70 and 72, device 10 can also be configured to determine the intensity of noise reduction performed by noise reduction unit 20. For example, via signal notification 78 ( Figure 3 Device 10 can signal the determination or selection of intensity. For example, signal notification 78 can indicate a threshold κ, which is also mentioned in the more mathematically presented example described below. Figure 3 The noise reduction unit 20 is shown to use the threshold κ to set all spectral components or coefficients in the spectrum 60 above κ to zero to generate the spectrum 62, or to cut and fold the spectrum 60 below the threshold κ, or to shift a portion of the spectrum 60 above the threshold κ to 0 so that it starts from 0 (e.g., Figure 3 (As shown). The same applies to signal notification 78 as noted above with respect to signal notifications 70 and 72. That is, device 10 can set the noise reduction intensity or threshold κ either globally or at the sub-image granularity. In the latter case, encoder 10 can optionally perform the settings individually for each block 46. (Refer to reference...) Figure 4 In the specific embodiment shown, encoder 10 can select either a noise reduction intensity or a threshold κ from a set of possible values ​​for κ, which itself is selected from multiple sets 80. Selection within set 80 can be based on a quantization parameter Q, which quantizer 30 performs quantization based on, and dequantizer 36 performs dequantization on the predicted residual signal based on that quantization parameter. Then, signal notification 78 signals the actual noise reduction intensity or threshold κ selected from the possible values ​​of κ in the selected set 80. It should be understood that the quantization parameter Q can be signaled in data stream 14 at a different granularity than the signal notification 78 signaling within data stream 14. For example, the quantization parameter Q can be signaled in data stream 14 on a slice or image basis, while signal notification 78 can (as described above) signal notification for each block 46 in data stream 14. Similar to the above, signal notification 78 can be transmitted within the data stream using predictive coding and / or context-adaptive entropy coding using spatial context.

[0035] Figure 5 It shows the method for adapting to Figure 1 The device's data stream 14 is used to predict and decode image 12 (a reconstructed version of image 12) based on block-by-block prediction. To a large extent, Figure 5 The internal structure of the decoder 100 is consistent with that of the encoder 10, as long as they relate to those encoding parameters (which are ultimately determined by...). Figure 1 The task of the device 10 (selected) is thus. Therefore, Figure 5 It shows Figure 5The apparatus 100 comprises a prediction loop in which the components 40, 16, 18, 20, 22 and optionally the signal 42 are connected in series in the manner shown and described above with reference to Figure 1 When the reconstructed portion of the signal to be reconstructed (i.e. the image 12) is produced at the output of the adder 40, this output represents the output of the decoder 100. Optionally, an image improvement module (e.g. a post filter) can be located in front of the output.

[0036] It should be considered that whenever the apparatus 10 has the freedom to select a certain coding parameter, the apparatus 10 selects this coding parameter to maximize, for example, a certain optimization criterion (e.g., a rate / distortion cost measure). The signaling in the data stream 14 is then used to keep the predictions performed by the encoder 10 and the decoder 100 synchronized. Corresponding modules or components of the decoder 100 can be controlled by the encoder 10 including the respective signaling in the data stream 14 and signaling the selected coding parameters. For example, the prediction provider 16 of the decoder 100 is controlled via the coding parameters in the data stream 14. For example, these coding parameters indicate a prediction mode and prediction parameters for the indicated prediction mode. The coding parameters are selected by the apparatus 10. Examples for the prediction parameters 102 have been mentioned above. The same holds for each of the signalings 70, 72 and 78, all of which are optional (i.e., none, one, two or all of them can be present). At the encoding side of the apparatus 10, the respective signalings are selected to optimize certain criteria and the selected parameters are indicated by the respective signalings. The signalings 70, 72 and 78 steer the following: the selector 42 (optional) as to the selection between the prediction padded versions; the spectral decomposer 18 as to the size of the region 52 (e.g., via a relative vector indicating the top-left corner of the region 52); the noise reducer 20 as to the strength of the noise reduction (e.g., via an indication of the threshold to be used). The loop just outlined is continuously fed with new residual data via the other input of the adder 40 (i.e., the input not connected to the selector 42), wherein the adder 40 in the reconstructor 34, the prediction provider 16 following it, the spectral decomposer 18, the noise reducer 20, the spectral composer 22 and the optional selector 42 are connected in series in said loop. Specifically, the entropy decoder 132 performs the inverse of the entropy encoder 32, i.e., the same entropy to decode the residual signal (i.e., the coefficient levels) from the data stream 14 in the spectral domain in the same way as applied to the blocks 46 along the above-mentioned encoding / decoding order 58. The entropy decoder 132 forwards these coefficient levels to the reconstruction stage 34, which dequantizes the coefficient levels in the dequantizer 36 and transforms the dequantized coefficient levels to the spectral domain by the inverse transformer 38, thereby obtaining the residual signal to be added to the final prediction signal (either the second prediction padded version 64 or the first version 48).

[0037] In summary, the decoder 100 can have access to the same information base for performing the prediction by the prediction provider 16 and has already reconstructed the samples within the neighborhood 54 of the current prediction block 46 using the prediction signals obtained from the data stream 14 via the block sequences 32, 36 and 38. The signalizations 70, 78 and 72, if present, allow for synchronization between the encoder and the decoder 100. As described above, the decoder 100 can be configured to change the corresponding parameters, i.e. the selection of the selector 42, the size of the region 52 at the spectral decomposer 18 and / or the noise reduction strength in the noise reducer 20, at sub-picture granularity, which can differ between these parameters as already set out above. The decoder 100 changes these parameters at this granularity, because the signalizations 70, 72 and / or 78 are signaled in the data stream 14 at this granularity. As described above, the apparatus 100 can decode any of the signalizations 70, 72 and 78 from the data stream 14 using spatial decoding. Additionally or alternatively, context adaptive entropy decoding using spatial contexts can be utilized. Furthermore, with respect to the signalization 78, i.e. the signalization controlling the noise reduction 20, the apparatus 100 can be configured to: as described above with respect to the encoder 10, determine the quantization parameter Q from the data stream 14 for the region in which the prediction block 46 currently is located; and then determine the noise reduction strength to actually be used for noise reducing the block 46 based on the signalization 78, which selects one noise reduction strength out of a pre-selected set of possible noise reduction strengths. For example, each set 80 can comprise eight possible noise reduction strengths. Since there is more than one set 80 selection, the total number of possible noise reduction strengths covered by all sets 80 can be eight times the number of sets 80. However, the sets 80 can overlap, i.e. some possible noise reduction strengths can be members of more than one set 80. Naturally, here only eight are used as an example and the number of possible noise reduction strengths per set 80 can differ from 8 and can even vary between sets 80. Figure 4 As outlined, based on the quantization parameter Q selection of one out of several subsets of possible noise reduction strengths, the apparatus 100 determines the quantization parameter Q from the data stream 14 for the region in which the prediction block 46 currently is located; then determines the noise reduction strength to actually be used for noise reducing the block 46 based on the signalization 78, which selects one noise reduction strength out of a pre-selected set of possible noise reduction strengths. For example, each set 80 can comprise eight possible noise reduction strengths. Since there is more than one set 80 selection, the total number of possible noise reduction strengths covered by all sets 80 can be eight times the number of sets 80. However, the sets 80 can overlap, i.e. some possible noise reduction strengths can be members of more than one set 80. Naturally, here only eight are used as an example and the number of possible noise reduction strengths per set 80 can differ from 8 and can even vary between sets 80.

[0038] Various variations can be made with respect to the above-described embodiments. For example, the encoding stage 24 and the reconstruction stage 34 need not be based on a transform. That is, the prediction residuals can be encoded in the data stream 14 in a manner different from using a spectral domain. Furthermore, the concept can work losslessly. As described before with respect to the relationship between the decomposer 18 and the transformer 38, the inverse transform of the inverse transformer 38 can be of the same type as the transform performed by the combiner 22.

[0039] The above idea can be implemented in such a way that a non-linear transform domain based prediction depending on the initial predictor and on surrounding reconstructed samples is generated. The above idea can be used to generate a prediction signal in video coding. In other words, the basic principle of the idea can be described as follows. In a first step, an image or video decoder generates a starting prediction signal, e.g. by motion compensation or intra or spatial image prediction, as in some underlying image or video compression standard. In a second step, the decoder proceeds as follows. First, it defines an extended signal, which consists of a combination of the prediction signal and the already reconstructed signal. Then, the decoder applies a linear analysis transform to the extended prediction signal. Next, the decoder applies a non-linear thresholding to the transformed extended prediction signal. In a last step, the decoder applies a linear synthesis transform to the result of the previous step and replaces the starting prediction signal by the result of the synthesis transform (restricted to the domain of the prediction signal).

[0040] In the next section, we will give a more mathematical presentation of the idea as an implementation example.

[0041] As an example, we consider a hybrid video coding standard, in which the content of a video frame on a block B Figure 2 )

[0042] B cmp := {(x, y) ∈ Z 2 : k 1,cmp ≤ x < k 2,cmp : l 1,cmp ≤ y < l 2,cmp},

[0043] where k 1,cmp , k 2,cmp , l 1,cmp , l 1,cmp ∈ Z have k 1,cmp < k 2,cmp and l 1,cmp < l 2,cmp is to be generated by the decoder. The latter content is given by the following function for each color component cmp: .

[0044] We assume that the hybrid video coding standard operates by predictively coding the current block B cmp .

[0045] This means that, as part of the standard, for each component cmp, the decoder constructs a prediction signal (P Figure 2 48) in a way that is uniquely determined by the already decoded bitstream. This predicted signal can be generated, for example, through intra-image prediction or through motion-compensated prediction. We are applying for copyright to the following extensions to this predictive coding, thereby enabling the generation of new predictive signals. cmp Apply for copyright.

[0046] 1. Change the standard so that, based on the bitstream, the decoder can determine for each component (cmp) that exactly one of the following two options is true:

[0047] (i) Option 1: Do not apply the new prediction method.

[0048] (ii) Option 2: Apply new prediction methods.

[0049] 2. Change the standard so that if the decoder determines in step 1 that option 2 of cmp is true for a given component, the decoder can determine the following data based on the bitstream:

[0050] (i) Unique integer

[0051] This makes it possible to achieve the following on block or L-shaped neighborhood 54 ( Figure 2 )

[0052] ,

[0053] The image has already been constructed It can be used in decoders. Define the size of the neighborhood 54, and it can be notified by 72 signals.

[0054] (ii) Unique integer The unique analytical transformation T (which is a linear mapping) is analyzed.

[0055]

[0056] And the unique composite transformation S (which is a linear mapping).

[0057]

[0058] S is Figure 1 and Figure 5 60 in. For example, T can be a discrete cosine or discrete sine transform, in which case, And S = T -1 .

[0059] However, in this case, an overcomplete transformation can also be used.

[0060]

[0061] (iii) Unique threshold and unique threshold processing operator

[0062] , it is either given as a hard threshold processing operator with threshold , where is defined by

[0063]

[0064] or as a soft threshold processing operator with threshold , where is defined by

[0065]

[0066] Here . The threshold processing can be done as part of the denoising 20 in Figure 1 and Figure 5 and is signaled by 78.

[0067] For example, the underlying codec can be changed such that the introduced parameters in the last three items can be determined by the decoder from the current bitstream, as described below. The decoder determines the index from the current bitstream and the decoder uses this index to determine the integer , the transform T and S and the threshold k and the specific threshold processing operator from a pre-defined look-up table. The latter look-up table can depend on some additional parameters already available to the decoder, e.g. the quantization parameter. That is, in addition or as an alternative to 70, 72 and 78, there can be a signaling for changing the transforms used by the synthesizer 22 and the decomposer 18 and their inverses. The signaling and the changing can be done globally for the picture or sub-picture granularly using spatial prediction and / or using spatial context entropy coding / decoding.

[0068] 3. Change the criterion such that if the decoder determines in step 1 that option two is true for a given component cmp, the decoder extends the block B cmp to a larger block:

[0069] B cmp,ext := {(x, y) e Z 2 : k’ 1,cmp ≤ x < k 2,cmp : l’ 1,cmp ≤ y < l 2,cmp},

[0070] where k’ 1,cmp , l’ 1,cmp are as in step 2 and the extended prediction signal is defined by :

[0071] ,

[0072] where B rec,cmp and rec cmp As in step 2. The signal pred cmp,ext can be canonically seen as a vector in.

[0073] Next, the decoder defines a new prediction signal (pred

[0074] pred Figure 1 and Figure 5 in 64

[0075] ,

[0076] where the analysis transform T, the threshold , the thresholding operator and the synthesis transform S are as in step 2.

[0077] 4. Change the normal codec so that if the decoder determines in step 1 that option two is true for a given component cmp, the decoder defines a new prediction signal (pred is as in the previous steps. The signalization 70 can be used as shown in Figure 1 and Figure 5 .

[0078] As mentioned above, the codec need not be a video codec.

[0079] While some aspects have been described in the context of an apparatus, it is clear that other aspects of the disclosure also represent a description, albeit an implicit one, of corresponding methods, wherein blocks or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps can be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or electronic circuit. In some embodiments, one or more of the most important method steps can be executed by such an apparatus.

[0080] The encoded image (or video) of the present application can be stored on a digital storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium, e.g. the Internet.

[0081] Depending on certain implementation requirements, embodiments of the application can be implemented in hardware or in software. The implementation can be carried out using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate with a programmable computer system such that one of the methods described herein is performed. Therefore, the digital storage medium can be computer readable.

[0082] Some embodiments according to the application comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0083] Generally, embodiments of the present application can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code can for example be stored on a machine readable carrier.

[0084] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0085] In other words, an embodiment of the inventive methods is, therefore, a computer program for performing one of the methods described herein, when the computer program runs on a computer.

[0086] A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitionary.

[0087] A still further embodiment of the inventive methods is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals can for example be configured to be transferred via a data communication connection, for example via the Internet.

[0088] A further embodiment comprises processing means, for example a computer, or a programmable logic device, configured to or adapted for performing one of the methods described herein.

[0089] One embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0090] Another embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.

[0091] In some embodiments, a programmable logic device (for example a field programmable gate array) can be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.

[0092] The apparatuses described herein can be implemented using a hardware apparatus, or using a computer, or using a combination of hardware and computer.

[0093] The apparatuses described herein, or any components of the apparatuses described herein, can be implemented at least partially in hardware and / or in software.

[0094] The methods described herein can be performed using a hardware apparatus, or using a computer, or using a combination of hardware and computer.

[0095] The methods described herein, or any components of the apparatuses described herein, can be performed at least partially in hardware and / or in software.

[0096] The above-described embodiments are merely illustrative for the principles of the present invention. It should be understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the appended patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.

Claims

1. A method for decoding an image from a data stream, comprising: deriving, from the data stream, a quantization parameter associated with a block of the image; applying a transform to a region of the image, wherein the region comprises: a first set of samples included in the block and a second set of samples that is contiguous to the block, wherein the first set of samples represents a prediction signal for the block and the second set of samples represents a reconstructed portion of the image that is contiguous to the block; determining a threshold based at least in part on the quantization parameter and a lookup table; filtering the transformed region using the threshold; applying an inverse of the transform to the filtered region to obtain a modified version of the region; and reconstructing the block using samples included in the modified version of the region that correspond to respective positions of the first set of samples.

2. The method of claim 1, wherein, Filtering the transformed region comprises: comparing a spectral coefficient of the transformed region to the determined threshold; and modifying the spectral coefficient when the spectral coefficient is less than the determined threshold.

3. The method of claim 1, further comprising: identifying, based on the quantization parameter, a set of thresholds from a plurality of sets of thresholds for noise reduction; and selecting the threshold from the identified set of thresholds based on a signaling in the data stream.

4. The method of claim 3, wherein, Each set of thresholds comprises 8 thresholds.

5. The method of claim 1, wherein the reconstructed portion of the image is L-shaped and comprises a lower portion and a right portion; the lower portion is collinear with a left side of the block along a vertical axis; and the right portion is collinear with a top side of the block along a horizontal axis.

6. The method of claim 1, further comprising: determining whether to reconstruct the block using a prediction residual and the first set of samples included in the block; in response to a first determination, reconstructing the block using a prediction residual and the first set of samples included in the block; and in response to a second determination, reconstructing the block using a prediction residual and a portion of the samples included in the modified version of the region.

7. An electronic device for decoding an image from a data stream, the electronic device comprising: a processor configured to: derive, from the data stream, a quantization parameter associated with a block of the image; apply a transform to a region of the image, the region comprising: a first set of samples included in the block and a second set of samples that is contiguous to the block, wherein the first set of samples represents a prediction signal for the block and the second set of samples represents a reconstructed portion of the image that is contiguous to the block; determine a threshold based at least in part on the quantization parameter and a lookup table; filter the transformed region using the threshold; apply an inverse of the transform to the filtered region to obtain a modified version of the region; and reconstruct the block using samples included in the modified version of the region that correspond to respective positions of the first set of samples.

8. The electronic device of claim 7, wherein, To filter the transformed region, the processor is further configured to: compare a spectral coefficient of the transformed region to the determined threshold; and modify the spectral coefficient when the spectral coefficient is less than the determined threshold. ​ 9. The electronic device of claim 7, wherein, The processor is further configured to: identify, based on the quantization parameter, a set of thresholds for noise reduction from a plurality of sets of thresholds; and select the thresholds from the identified set of thresholds based on a signaling in the data stream.

10. The electronic device of claim 9, wherein, Each set of thresholds includes 8 thresholds.

11. The electronic device of claim 7, wherein the reconstructed portion of the image is L-shaped and includes a lower portion and a right portion; the lower portion is collinear with a left side of the block along a vertical axis; and the right portion is collinear with a top side of the block along a horizontal axis.

12. The electronic device of claim 7, wherein, The processor is further configured to: determine whether to use a prediction residual and a first set of samples included in the block to reconstruct the block; in response to a first determination, use the prediction residual and the first set of samples included in the block to reconstruct the block; and in response to a second determination, use the prediction residual and a portion of samples included in a modified version of the region to reconstruct the block.

13. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor of an electronic device, cause the electronic device to: derive, from a data stream, a quantization parameter associated with a block of an image; apply a transform to a region of the image, wherein the region includes: a first set of samples included in the block, and a second set of samples that are contiguous to the block, wherein the first set of samples represents a prediction signal for the block, and the second set of samples represents a reconstructed portion of the image that is contiguous to the block; determine a threshold based at least in part on the quantization parameter and a lookup table; filter the transformed region using the threshold; apply an inverse of the transform to the filtered region to obtain a modified version of the region; and reconstruct the block using samples included in the modified version of the region that correspond to respective positions of the first set of samples.

14. The non-transitory computer readable medium of claim 13, wherein, The instructions that, when executed, cause the at least one processor to filter the transformed region further include instructions that, when executed, cause the at least one processor to: compare a spectral coefficient of the transformed region to the determined threshold; and modify the spectral coefficient when the spectral coefficient is less than the determined threshold.

15. The non-transitory computer-readable medium of claim 13, further comprising instructions that, when executed, cause the at least one processor to: identify, based on the quantization parameter, a set of thresholds for noise reduction from a plurality of sets of thresholds, wherein each set of thresholds includes 8 thresholds; and select the thresholds from the identified set of thresholds based on a signaling in the data stream.

16. The non-transitory computer-readable medium of claim 13, wherein the reconstructed portion of the image is L-shaped and includes a lower portion and a right portion; the lower portion is collinear with a left side of the block along a vertical axis; and the right portion is collinear with a top side of the block along a horizontal axis.

17. The non-transitory computer-readable medium of claim 13, further comprising instructions that, when executed, cause the at least one processor to: determine whether to use a prediction residual and a first set of samples included in the block to reconstruct the block; in response to a first determination, reconstructing the block using the prediction residual and a first set of samples included in the block; and in response to a second determination, reconstructing the block using the prediction residual and a portion of samples included in the modified version of the region.

Citation Information

Patent Citations

  • Method for decoding moving images, method for encoding moving images, moving image decoder, moving image encoder, program and integrated circuit

    CN102227910A

  • Scalable video coding using base-layer hints for enhancement layer motion parameters

    US20160014430A1