Cross-component prediction

The multi-component image coding method reconstructs second component signals from spatially corresponding first component signals with adaptive weights and domain switching, addressing inefficiencies in existing color space transformations to improve coding efficiency and reduce redundancy.

JP2026035646APending Publication Date: 2026-03-04DOLBY VIDEO COMPRESSION LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing image and video compression technologies face inefficiencies due to suboptimal color space transformations that introduce correlation between color components, leading to higher bit rates and unnatural signals, particularly in R'G'B' color synthesis, and fail to completely remove inter-component redundancy despite decorrelation efforts.

Method used

A multi-component image coding approach that reconstructs a second component signal from spatially corresponding portions of the first component signal and a correction signal, adaptively adjusting weights and domains to reduce inter-component redundancy, allowing for flexible switching of inter-component prediction at various granularities.

Benefits of technology

This method enhances coding efficiency by effectively removing residual redundancy, achieving higher bit rates over a wider range of multi-component image content with reduced signaling overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035646000001_ABST
    Figure 2026035646000001_ABST
Patent Text Reader

Abstract

A method and apparatus for generating a reconstructed component signal is provided. The decoder (200) increases coding efficiency over a wider range of multi-component image content by reconstructing a second component signal (256'2) for a second component of a multi-component image (202) from a spatially corresponding portion of a reconstructed first component signal (2561) and a correction signal (2562) derived from a data stream (104) for the second component (208). By including a spatially corresponding portion of the reconstructed first component signal in the reconstruction of the second component signal, any remaining inter-component redundancies / correlations are also present, either despite the a priori performed component space transformation or because they were introduced by the a priori performed component space transformation.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to inter-component prediction in multi-component image coding, for example between luma and chroma. [Background technology]

[0002] In image and video signal processing, color information is typically represented in a three-component color space, such as R'G'B' or Y'CbCr. The first component, Y', in the case of Y'CbCr, is often called luma, and the remaining two components, the Cb and Cr components or planes in the case of Y'CbCr, are called chroma. The advantage of the Y'CbCr color space over the R'G'B' color space is primarily the residual property of the chroma components, i.e., they contain less energy or amplitude compared to chroma signals in absolute color spaces such as R'G'B'. For Y'CbCr in particular, the luma component represents the grayscale information of the image or video, and the chroma component Cb represents the difference relative to the blue primary color, and the chroma component Cr represents the difference relative to the red primary color.

[0003] In the application space of image and video compression and processing, Y'CbCr signals are preferred over R'G'B' due to the fact that color space conversion from R'G'B' to Y'CbCr reduces or eliminates correlation between different color components or planes. In addition to decorrelation, less information must be transmitted, and therefore color conversion also serves as a compression approach. Such preprocessing, such as decorrelation or reduction, enables higher compression efficiency while maintaining or increasing complexity by a meaningful amount, for example. Hybrid video compression schemes are often designed for Y'CbCr input because correlation between different color components is eliminated or reduced; furthermore, the design of the hybrid compression scheme only needs to consider separate processing of the different components. However, conversion from R'G'B' to Y'CbCr and vice versa is not lossless; therefore, information, i.e., sample values ​​available in the original color space, may be lost after such color conversion. This problem can be avoided by using a color space that includes a lossless transformation from and to the original color space, such as the Y'CoCg color space when using R'G'B' input. Nevertheless, fixed color space transformations can yield suboptimal results depending on the application. For image and video compression, fixed color transformations are often suboptimal due to higher bit rates and unnatural signals with high or no correlation between color planes. In the second case, a fixed transformation introduces correlation between different signals, and in the first case, a fixed transformation may not remove all correlation between different signals. Furthermore, due to the global application of the transformation, correlation may not be completely removed from different components or planes, even locally or globally. Another problem introduced by color space transformations lies in the architecture of the image or video encoder. Typically, the optimization process attempts to reduce a cost function, which is often a distance metric defined over the input color space.In the case of a transformed input signal, it may be difficult to achieve optimal results for the original input signal due to additional processing steps. As a result, the optimization process may result in the minimum cost for the transformed signal but not for the original input signal. Although the transformation is often linear, the cost calculation in the optimization process often includes signaling overhead, and the cost for the final decision is then calculated using the Lagrangian formula. The latter may result in different cost values ​​and different optimization decisions. Color transformation aspects are particularly important in the area of ​​color representation, as modern image and video displays typically use R'G'B' color synthesis for content representation. Generally speaking, transformations are applied when correlations within or between signals should be removed or reduced. As a result, color space transformations are a special case of a more general transformation approach.

[0004] It would therefore be desirable to have a multi-component image coding concept at hand that is even more efficient, i.e., that achieves higher bit rates over a wider range of multi-component image content. Summary of the Invention [Problem to be solved by the invention]

[0005] It is therefore an object of the present invention to provide such a more efficient multi-component image coding concept. [Means for solving the problem]

[0006] This object is achieved by the subject matter of the pending independent claims. The invention is based on the discovery that reconstructing a second component signal associated with a second component of a multi-component image from spatially corresponding portions of the reconstructed first component signal and a correction signal derived from the data stream for the second component promises increased coding efficiency over a wider range of multi-component image content. By including spatially corresponding portions of the reconstructed first component signal in the reconstruction of the second component signal, any remaining inter-component redundancy / correlation, possibly still present despite the deductively performed component space transformation, or present because it was introduced by such a deductively performed component space transformation, can be readily removed, for example, by inter-component redundancy / correlation reduction of the second component signal.

[0007] According to an embodiment of the present application, the multi-component image codec is interpreted as a block-based hybrid video codec that operates in units of code blocks, prediction blocks, residual blocks, and transform blocks, and inter-component dependencies are switched on and off at the granularity of residual blocks and / or transform blocks by respective signaling in the data stream. The additional overhead for signaling is overcompensated by the coding efficiency gain, as the amount of inter-component redundancy can vary within an image. According to an embodiment of the present application, the first component signal is a prediction residual of temporal, spatial, or inter-view prediction of a first component of the multi-component image, and the second component signal is a prediction residual of temporal, spatial, or inter-view prediction of a second component of the multi-component image. With this approach, the exploited inter-component dependencies focus on the remaining inter-component redundancies, as inter-component prediction tends to exhibit smoother spatial behavior.

[0008] According to an embodiment, a first weight, denoted α below, by which a spatially corresponding portion of a reconstructed first component signal influences the reconstruction of a second component signal, is adaptively set at sub-picture granularity. This approach allows intra-picture variations in inter-component redundancy to be more closely followed. According to an embodiment, a mixture of a high-level syntax element structure and a sub-picture granularity first weight syntax element is used to signal the first weight at sub-picture granularity, where the high-level syntax element structure defines a mapping from a set of possible bin sequence regions of a given binarization of the first weight syntax element to a joint region of possible values ​​for the first weight. This approach keeps the overhead for side information for controlling the first weight low. Adaptation can be performed forward-adaptively. A syntax element can be used for each block, e.g., a residual or transform block, with a limited number of signalable states that symmetrically index one of multiple weight values ​​for α, which is symmetrically distributed around zero. In one embodiment, the number of signalable states is not equal to the number of weight values, including zero, and signaling zero is used to signal non-use of inter-component prediction, so that an extra flag is obsolete. Furthermore, the magnitude is signaled before the conditionally signaled code, and the magnitude is mapped to the number of weight values, and further, if the magnitude is zero, the code is not signaled, so that the signaling cost is further reduced.

[0009] According to an embodiment, a second weight, by which the correction signal influences the reconstruction of the second component signal, is set at sub-picture granularity in addition to or instead of adaptively setting the first weight. This measure can further increase the adaptability of inter-component redundancy reduction. In other words, according to an embodiment, when reconstructing the second component signal, the weight of a weighted sum of spatially corresponding parts of the correction signal and the reconstructed first component signal can be set at sub-picture granularity. The weighted sum can be used as a scalar argument of a scalar function that is constant, at least for each image. The weight can be set in a backward direction based on a local neighborhood. The weight can be corrected in a forward-driven manner.

[0010] According to an embodiment, the domain in which the reconstruction of the second component signal from the spatially corresponding portion of the reconstructed first component signal using the correction signal is performed is the spatial domain. Alternatively, the spectral domain is used. Alternatively, the domain used is changed between the spatial and spectral domains. The switching is performed at sub-picture granularity. It has been found that the ability to switch the domain in which the combination of the reconstructed first component signal and the correction signal occurs at sub-picture granularity increases coding efficiency. The switching can be performed in a backward-adaptive or forward-adaptive manner.

[0011] According to an embodiment, syntax elements in the data stream are used to enable changing the roles of first and second component signals within a component of a multi-component image, where the additional overhead for signaling the syntax elements is low compared to the possible gains in coding efficiency.

[0012] According to an embodiment, the reconstruction of the second component signal can be switched, at sub-picture granularity, between a reconstruction based only on the reconstructed first component signal and a reconstruction based on the reconstructed first component signal and further reconstructed component signals of further components of the multi-component image. With relatively low additional effort, this possibility increases the flexibility in removing residual redundancy between components of the multi-component image.

[0013] Similarly, according to an embodiment, a first syntax element in the data stream is used to enable or disable, globally or at an increased scope level, the reconstruction of a second component signal based on a reconstructed first component signal. When enabled, a sub-picture level syntax element in the data stream is used to adapt the reconstruction of the second component signal based on the reconstructed first component signal at sub-picture granularity. With this approach, spending side information for sub-picture level syntax elements is only necessary in applications or multi-component image content where enablement results in coding efficiency gains.

[0014] Instead, the switching between enablement and disablement is performed locally in the opposite direction. In this case, the first syntax element does not even need to be present in the data stream. According to an embodiment, for example, the local switching is performed locally depending on whether the first and second component signals are prediction residuals of spatial prediction in intra-prediction mode of spatial prediction that match or do not deviate by more than a predetermined amount. By this measure, the local switching between enablement and disablement does not consume bitrate.

[0015] According to an embodiment, a second syntax element in the data stream is used to switch between adaptive reconstruction of the second component signal based on a first component signal reconstructed at sub-picture granularity using sub-picture level syntax elements in the data stream in a forward-adaptive manner and non-adaptively performing reconstruction of the second component signal based on the reconstructed first component signal, where the signaling overhead for the second syntax element is low compared to the possibility of avoiding overhead for transmitting sub-picture level syntax elements for multi-component image content where non-adaptively performing reconstruction is already sufficiently efficient.

[0016] According to an embodiment, the concept of inter-component redundancy reduction is transferred to a three chroma component image. According to an embodiment, a luma and two chroma components are used. The luma component may be selected as the first component.

[0017] According to an embodiment, sub-picture level syntax elements for accommodating the reconstruction of a second component signal from a reconstructed first component signal are coded in a data stream using Golomb-Rice coding. Bins of the Golomb-Rice code may be subjected to binary arithmetic coding. Different contexts may be used for different bin positions of the Golomb-Rice code.

[0018] According to an embodiment, the reconstruction of the second component signal from the reconstructed first component signal includes spatially rescaling and / or bit-depth fine mapping to spatially corresponding portions of the reconstructed first component signal. The adaptation of the spatial rescaling and / or bit-depth fine mapping may be performed in a backward and / or forward adaptive manner. The adaptation of the spatial rescaling may include selecting a spatial filter. The adaptation of the bit-depth fine mapping may include selecting a mapping function.

[0019] According to an embodiment, the reconstruction of the second component signal from the reconstructed first component signal is performed indirectly via a spatially low-pass filtered version of the reconstructed first component signal.

[0020] Advantageous developments of the embodiments of the present application are the subject of the dependent claims.

[0021] Preferred embodiments of the present application are described below with reference to the figures. [Brief explanation of the drawings]

[0022] [Figure 1] FIG. 1 shows a block diagram of an encoder configured to encode a multi-component image according to an embodiment. [Figure 2] FIG. 2 shows a block diagram of a decoder compatible with the encoder of FIG. 1 according to an embodiment. [Figure 3] FIG. 3 shows a schematic representation of an image and its subdivision / division into various blocks of different types according to an embodiment. [Figure 4] FIG. 4 illustrates schematically a reconstruction module incorporated in the encoder and decoder of FIGS. 1 and 2 according to an embodiment of the present application. [Figure 5] FIG. 5 shows schematically an image with two components, a current inter-component predicted block and its spatially corresponding part in the first (base) component. [Figure 6] FIG. 6 shows a schematic image with three components to illustrate an embodiment in which one component always acts as the base (first) component according to the embodiment. [Figure 7] FIG. 7 shows schematically an image with three components, each of which can alternately act as a first (base) component according to an embodiment. [Figure 8] FIG. 8 shows schematically the possibility of adaptively varying the domain for inter-component prediction. [Figure 9a] FIG. 9a illustrates schematically a possibility for signaling inter-component prediction parameters. [Figure 9b] Figure 9b shows the embodiment of Figure 9a in more detail. [Figure 9c] FIG. 9c shows, for illustrative purposes, a schematic illustration of the application of the inter-component prediction parameters to which FIGS. 9a and 9b relate. [Figure 10] FIG. 10 illustrates schematically forward-adaptive inter-component prediction according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0023] The following carryover description begins with a detailed embodiment description of an encoder and a detailed embodiment description of a decoder that matches the encoder, followed by a generalized embodiment.

[0024] Figure 1 illustrates an encoder 100 configured to encode a multi-component image 102 into a data stream 104. While Figure 1 illustratively illustrates the case of a multi-component image 102 including three components 106, 108, and 110, the following carryover discussion will make clear that advantageous embodiments of the present application can be readily derived from the discussion of Figure 1 by migrating the present three-component image example to an example in which only two components are present, such as components 106 and 108. Accordingly, some of the elements of encoder 100 illustrated in Figure 1 are indicated using dotted lines and, to the extent that they are optional features with respect to more generalized embodiments as described further below.

[0025] In principle, the encoder of Figure 1 and all of the embodiments described below can operate with any type of "component." In practice, multi-component image 102 is an image of a scene that samples the scene with spatially varying components 106, 108, and 110 under other matching conditions, e.g., with respect to timestamps, viewing conditions, etc. However, for ease of understanding of the embodiments described further below, components 106, 108, and 110 are all assumed to be texture-related and, further, color components, e.g., luma and chroma components, or color components of any other color space.

[0026] 1 is substantially configured to apply block-based hybrid video coding to each of components 106, 108, and 110 of image 102 separately, but using inter-component prediction to reduce inter-component redundancy. Inter-component prediction relates to various aspects as outlined in further detail below.

[0027] Thus, the encoder 100 includes, for each component 106, 108, and 110, a prediction residual former 112, exemplarily implemented here as a series of subtractors, a transformer 114, and a quantizer 116, connected in series, in the order of their listing, between the input at which the respective component 106, 108, and 110 arrives and a respective input of a data stream former 118, which is configured to multiplex quantization coefficients and other coding parameters, described in more detail below, into the data stream 104. The non-inverting input of the prediction residual former 112 is arranged to receive the respective component of the image 102, while its inverting (subtractive) input receives a prediction signal 120 from a predictor 122, the input of which is connected to the output of the quantizer 116 via a reconstruction path including a series of an inverse quantizer 124, a re-transformer 126, and a prediction / residual combiner 128, exemplarily implemented here as an adder. The output of the prediction / residual recombiner 126 is connected to the input of the predictor 122, and furthermore, a first input is connected to the output of the reconverter 126, a further input of which receives the prediction signal 120 output by the predictor 122.

[0028] Elements 112, 114, 116, 124, 126, 128, 122, and 120 exist in encoder 100 in parallel for each of components 106-110 to form three separate block-based hybrid video coding paths 1301, 1302, and 1303. Indices 1, 2, and 3 are used to distinguish between different components of image 102, with index 1 associated with component 106, index 2 associated with component 108, and index 3 associated with component 110.

[0029] In the example shown in Figure 1, component 106 represents a kind of "base component," as will become more apparent from the carryover description below. However, it should be noted that the roles of components 106-110 may be equalized, in that the configuration described and shown in Figure 1 has been chosen solely for easier understanding of the principles of embodiments of the present application, as will be described in more detail below. In any case, indices are also used for elements 112, 114, 116, 124, 126, 128, 122, and 120 of paths 1301-1303 described thus far to distinguish these elements.

[0030] 1, components 108 and 110 represent "dependent components," and for this reason, in addition to those already described for encoding paths 1301-1303, further elements are connected between these already described elements. In particular, the spatial domain inter-component prediction residual former 132 is connected by its non-inverting input and its output between the prediction residual former 112, on the one hand, and the input of the transformer 114, on the other hand; similarly, the spatial domain inter-component prediction and residual recombiner 134 is connected between the retransformer 126, on the one hand, and the prediction / residual recombiner 128, on the other hand. Figure 1 also shows the spectral domain inter-component prediction residual former 136, which is located by its non-inverting input and its output between the transformer 114 and the quantizer 116, together with a corresponding spectral domain inter-component prediction / residual recombiner 138, which is connected between the inverse quantizer 124, on the one hand, and the retransformer 126, on the other hand. The pair of spatial domain inter-component prediction former 132 and corresponding recombiner 134 is supplied at its further input by a respective spatial domain inter-component residual predictor 140, and similarly, a further input of the spectral domain inter-view prediction former 136, together with its associated recombiner 138, is supplied with a spectral domain residual prediction signal 142 provided by a spectral domain inter-component residual predictor 144. At their input interfaces, the predictors 140 and 144 are connected to any internal node of the respective previous coding path. For example, the predictor 1402 is connected to the output of the prediction residual former 1121, and the predictor 1442 has an input connected to the output of the inverse quantizer 1241. Similarly, predictors 1403 and 1443 have inputs connected to the outputs of prediction residual former 1122 and inverse quantizer 1242, respectively, but in addition, a further input of predictor 1403 is shown connected to the output of prediction residual former 1121, and a further input of predictor 1443 is shown connected to the output of inverse quantizer 1241.Predictors 140 and 144 are shown to optionally generate prediction parameters for implementing forward-adaptive control of the inter-component prediction performed thereby, the parameters output by predictor 140 being indicated by 146 and the parameters output by predictor 144 being indicated by reference numeral 148. The prediction parameters 146 and 148 of predictors 140 and 144, as well as the residual signal 150 output by quantizer 116, are received by data stream former 118, which, for example, encodes all of this data in a lossless manner into data stream 104.

[0031] Before describing the mode of operation of the encoder of Fig. 1 in more detail below, reference is made to Fig. 2, which shows a corresponding decoder. Substantially, decoder 200 corresponds, component by component, to the portion of encoder 100 extending from inverse quantizer 124 to predictor 122, and accordingly, the same reference numerals, increased by 100, are used for these corresponding elements. More precisely, decoder 200 of Fig. 2 is configured to receive data stream 104 at the input of a data stream extractor 218 of decoder 200, which is configured to extract from data stream 104 residual signals 1501-1503 and all associated coding parameters, including prediction parameters 146 and 148. Corresponding to encoder 100 of Fig. 1, decoder 200 of Fig. 2 is configured into parallel decoding paths 230, one path per component of the multi-component image, the reconstruction of which is indicated using reference numeral 202, and the reconstructed components of which are output at the outputs of the respective decoding paths as 206, 208, and 210. Due to quantization, the reconstruction deviates from the original image 102 .

[0032] As already mentioned above, the decoder paths 2301 to 2303 of the decoder 200 substantially correspond to those paths of the encoding paths 1301 to 1303 of the encoder 100, including the elements 124, 138, 126, 134, 128 and 122 thereof. That is, the decoding paths 2302 and 2303 of the "dependent components" 208 and 210 comprise a series connection of an inverse quantizer 224, a spectral domain inter-component prediction / residual recombiner 238, an inverse transformer 226, a spatial domain inter-component prediction / residual recombiner 234 and a prediction / residual recombiner 228, connected, in that order, between the respective output of the data stream extractor 218 on the one hand and the respective output for outputting the respective components 208 and 210, and between the data stream extractor 218 on the one hand and the output of the decoder 200 for outputting the multi-component image 202 on the other hand, with the predictor 222 connected in a feedback loop from the prediction / residual recombiner 228 back to its further input. Further inputs of 238 and 234 are provided by respective spatial and spectral domain inter-component predictors 240 and 244. The decoding path 2301 of the "base component" 206 differs in that elements 238, 234, 244, and 240 are absent. Predictor 2443 has its input connected to the output of inverse quantizer 2242 and another output connected to the output of inverse quantizer 2241, and predictor 2403 has a first input connected to the output of inverse transformer 2262 and another input connected to the output of inverse transformer 2261. Predictor 2442 has an input connected to the output of inverse quantizer 2241, and further, predictor 2402 has its input connected to the output of inverse transformer 2261.

[0033] It should be noted that although the predictors 1403 and 1402 of Figure 1 are shown connected to the as-yet-quantized residual signals of the underlying components in the spatial domain merely for the sake of easier explanation, as will become clear from Figure 2, in order to avoid prediction conflicts between the encoder and decoder, the predictors 1402 and 1403 of Figure 1 may alternatively and advantageously be connected to the outputs of the inverse transformers 1262 and 1261, respectively, instead of the unquantized versions of the spectral domain residual signals of the underlying components.

[0034] After describing the structure of both the encoder 100 and the decoder 200, their modes of operation are described below. In particular, as already mentioned above, the encoder 100 and the decoder 200 are configured to use hybrid coding to encode / decode each component of the multi-component image. In effect, each component of the multi-component image 102 represents a sample array or image, each spatially sampling the same scene with respect to a different color component. The spatial resolution of the components 106, 108, and 110, i.e., the spectral resolution at which the scene is sampled for each component, can differ from component to component.

[0035] As mentioned above, the components of the multi-component image 102 are subjected to hybrid encoding / decoding separately. "Separately" does not necessarily mean that the encoding / decoding of the components is performed completely independently of each other. First, inter-component prediction eliminates redundancy between components, and some coding parameters may be commonly selected for the components. In FIG. 3, for example, it is assumed that the predictors 122 and 222 select the same subdivision of the image 102 into coding blocks. However, individual subdivision of the image 102 into coding blocks by the predictors 122 and 222 is also possible. The subdivision of the image 102 into coding blocks may be fixed or signaled within the data stream 104. In the latter case, the subdivision information may be part of the prediction parameters output by the predictor 122 to the data stream former 118, further indicated by reference numeral 154, and the data stream extractor 218 extracts these prediction parameters, including the subdivision information, and outputs them to the predictor 222.

[0036] 3 exemplarily illustrates that the subdivision of an image 102 into coding blocks or code blocks may be performed according to a two-stage process, whereby the multi-component image 102 is first regularly subdivided into tree root blocks, the outline of which is shown in FIG. 3 using double lines 300, followed by the application of a recursive multi-tree subdivision to subdivide each tree root block 302 into code blocks 304, the outline of which is shown in FIG. 3 using simple solid lines 306. In this manner, the code blocks 304 represent leaf blocks of the recursive multi-tree subdivision of each tree root block 302. The above-mentioned subdivision information, possibly included within the data stream 104, for the components jointly or individually may include multi-tree subdivision information for each tree root block 302 for signaling the subdivision of each tree root block 302 to the code blocks 304, and may optionally further include subdivision information controlling and signaling the subdivision of the image 102 into a regular array of tree root blocks 302 in rows and columns.

[0037] For example, predictors 122 and 222 change between multiple prediction modes supported by encoder 100 and decoder 200 on a code block 304 basis. For example, predictors 1221-1223 individually select a prediction mode for the code block and indicate their selection to predictor 2223 via prediction parameters 1541-1543. Available prediction modes may include temporal and spatial prediction modes. Other prediction modes may be supported, such as inter-view prediction modes. Using recursive multi-tree subdivision, such as dual-tree subdivision, code blocks may be further subdivided into prediction blocks 308, the outline 310 of which is shown in FIG. 3 using dotted lines. Respective recursive subdivision information may be included in prediction parameters 1541-1543 for each code block. In an alternative embodiment, this selection of prediction mode is performed at the granularity of the prediction block. Similarly, each code block 304 may be further subdivided into residual blocks 312, for example using a recursive multi-tree subdivision, such as dual-tree subdivision, the outline of which may be shown in FIG. 3 using dashed-dotted lines 314. Thus, each code block 304 is divided in parallel into a prediction block 308 and a residual block 312. The residual blocks 312 may optionally be further subdivided into transform blocks 316, the outline of which is shown in FIG. 3 using dashed-dotted lines 318. Alternatively, the transform and residual blocks form the same entity, i.e., the residual blocks are transform blocks and vice versa. In other words, the residual blocks may coincide with the transform blocks according to alternative embodiments. Residual block and / or transform block-related subdivision information may or may not be included in the prediction parameters 154.

[0038] Depending on the prediction mode associated with each codeblock or prediction block, each prediction block 308 has a respective prediction parameter associated therewith that is appropriately selected by the predictors 1221-1223 and thus inserted into the parameter information 1541-1543 and further used by the predictors 2221-2223 to control prediction within each prediction block 308. For example, a prediction block 308 having a temporal prediction mode associated therewith may have a motion vector for motion-compensated prediction associated therewith and, optionally, a reference image index indicating a reference image from which the prediction of each prediction block 308 is derived / copied by the displacement indicated by the motion vector. A prediction block 308 having a spatial prediction mode may have a spatial prediction direction associated therewith that is included in the prediction information 1511-1543, the latter indicating the direction in which the already reconstructed surroundings of the respective prediction block are spatially extrapolated to the respective prediction block.

[0039] Thus, using the prediction mode and prediction parameters, the predictors 122 and 222 derive prediction signals 1201-1203 for each of the components, which are then corrected using residual signals 1561-1563 for each of the components. These residual signals are coded using transform coding. That is, the transformers 1141-1143 perform a transform, i.e., a spectral decomposition, such as a DCT or DST, on each transform block 116 individually, and the inverse transformer 226 inverts it individually for the transform blocks 308, i.e., performs an IDCT or IDST, for example. That is, as far as the encoder 100 is concerned, the transformer 114 performs a transform on the not-yet-quantized residual signal as formed by the prediction residual former 112. The inverse transformers 126 and 226 reverse the spectral decomposition based on the quantized residual signal 150 of each component, which is inserted into the data stream by the data stream former 118 in a lossless manner, for example using Huffman or arithmetic coding, and which is further extracted therefrom by the data stream extractor 218, for example using Huffman or arithmetic decoding.

[0040] However, to achieve a lower data rate for encoding the residual signals 2561-2563, with which the prediction signals 2201-2203 are corrected in the prediction / residual recombiners 2281-2283, the encoders 100 and 200 support inter-component prediction for encoding the component residual signals. As described in more detail below, according to embodiments of the present application, the inter-component prediction for encoding the residual signals can be switched on / off and / or forward- and / or backward-adaptively adjusted at the granularity of residual blocks and / or transform blocks. When switched off, the inter-component prediction signals output by the predictors 140, 144, 240, and 244 are zero, and further, the residual signals 2561-2563 of all components are derived solely from the quantized transform coefficients included in their respective residual signals 1501 and 1503. However, when switched on, inter-component redundancy / correlation is removed as far as the dependent components are concerned, i.e., the residual signals 2562 and 2563 are encoded / decoded using inter-component prediction. The base (first) component, which acts as the inter-component prediction source, is left unchanged as far as the inter-component prediction is concerned. How this is done is outlined below.

[0041] For the time being, the description of the inter-component redundancy reduction achieved by predictors 140, 144, 240, and 244 will focus on the inter-component redundancy reduction between components 106, 108, 206, and 208, respectively. Thereafter, for ease of understanding, the description will be extended to the case of three components shown in Figures 1 and 2.

[0042] As will become apparent from the carryover description below, the embodiments outlined below utilize, for example, residual characterization of chroma components, particularly in cases where absolute chroma components or planes serve as inputs. For example, in the three-component case of FIGS. 1 and 2 representing components of a luma / chroma color space, component 106 / 206 may be, for example, a luma component, while components 108 / 208 and 110 / 210 are chroma components, such as blue- and red-related chroma components. By keeping distortion calculations defined across the original input space, the embodiments outlined below enable higher fidelity. In other words, the encoder of FIG. 1 together with the decoder of FIG. 2 enables it to input the original multi-component image, for example, without any color space conversion having been performed beforehand. The coding efficiency achievable by the embodiments of FIGS. 1 and 2 is independent of the input color space; i.e., the embodiments outlined below for performing inter-component redundancy elimination act as an additional decorrelation step for y'CbCr inputs 106-110 and as a color transform in the case of R'G'B' inputs 106-110. Additionally, the embodiments outlined below apply redundancy reduction locally between components or planes; i.e., the encoder 100 adaptively determines for each region of the image / picture 102 or video whether inter-plane / component prediction involving the residual signal is applied. However, it should also be noted that embodiments of the present invention are not limited to color components or planes; rather, the techniques described herein can be applied to general planes / components, as well as cases where the resulting residual of two or more planes consists of correlations. For example, the embodiments described herein can be applied to planes from different layers in a scalable video compression scheme. Upon prediction, the processing order of components or planes can be controlled locally, e.g., by syntax elements signaling the component order in the bitstream. This alternative is also described below.

[0043] 1 and 2, the case of a hybrid block-based image and video compression scheme is shown, in which the prediction performed by the predictor 122 / 222 is performed on a log basis, and furthermore, transform coding is applied to the prediction error, i.e. the output of the prediction residue former 112, called the residue, to which, as far as the encoding side is concerned, the transformer 114, the quantizer 116 and the data stream inserter 118 contribute, and as far as the decoding side is concerned, the inverse transformer 226, the inverse quantizer 224 and the data stream extractor 218 contribute.

[0044] It should be noted that the term block shall be understood as generally referring to a rectangular shape as described below, ie a block may have a rectangular shape.

[0045] In the embodiment described next, the encoder determines the application of claim / inter-component prediction for each prediction block 308 or transform block 304 .

[0046] However, as an intermediate remark, it is presented here that the embodiments of the present invention are not limited to the cases outlined with reference to Figures 1 and 2, where inter-plane / component prediction is applied to prediction residuals of spatial, temporal and / or inter-view predicted signals. Theoretically, the embodiments shown here can be similarly transferred to the case where inter-plane / component prediction is performed directly on component samples. In other words, when prediction itself is skipped, the original samples of the input signal should be called or can be treated as residuals for the rest of the description.

[0047] For each residual block 312 (or rectangular shape), according to an embodiment, a syntax element is transmitted in the data stream 104, and the syntax element further indicates whether inter-plane / inter-component prediction by the predictors 140, 144, 240, 244 should be used. In the case of a video compression scheme such as H.265 / EVC, also shown in FIG. 3, the residual is further divided into smaller transform blocks or shapes 216, and the encoder 100, according to an embodiment, can transmit the just-mentioned syntax element specifying the use or non-use of inter-plane / inter-component prediction for each transform block 316 or group of transform blocks. Note that the grouping, i.e., signaling level, can be adaptively selected, and further, such grouping decision can be transmitted as a further syntax element in the bits / data stream 104, for example, in the header of the data stream. At the highest signaling level, the use or non-use of inter-plane / inter-component prediction can be signaled for a coding unit or coding block 304, a group of code blocks 304, or even for an entire image or frame.

[0048] Alternatively, reference is made to Figure 4. Figure 4 shows a reconstruction module 400, which is divided into several parts: a spatial domain inter-component predictor 1402 and associated spatial domain inter-component prediction / residual recombiner 1342, which receives x in the form of a data stream extracted, dequantized and inverse transformed residual signal 1561, receives y in the form of a data stream extracted, dequantized and inverse transformed residual signal 1562, and outputs z as a residual signal 156'2 that replaces the residual signal 1562; a spectral domain inter-component predictor 1442 and associated spectral domain inter-component prediction / residual recombiner 1382, which receives x in the form of a data stream extracted and dequantized residual signal 1701, receives y in the form of a data stream extracted and dequantized residual signal 1702, and outputs z as a residual signal 170′2 that replaces the residual signal 1702; a spatial domain inter-component predictor 1403 and associated spatial domain inter-component prediction / residual recombiner 1343, which receives x in the form of a data stream extracted, dequantized and inverse transformed residual signal 1561 or a data stream extracted, dequantized and inverse transformed residual signal 1562 (or any combination thereof or both), receives y in the form of a data stream extracted, dequantized and inverse transformed residual signal 1563, and outputs z as a residual signal 156'3 that replaces the residual signal 1563; a spectral domain inter-component predictor 1443 and associated spectral domain inter-component prediction / residual recombiner 1383, which receives x in the form of a data stream extracted and dequantized residual signal 1701 or a data stream extracted and dequantized residual signal 1702 (or any combination thereof or both), receives y in the form of a data stream extracted and dequantized residual signal 1703, and outputs z as a residual signal 170'3 replacing the residual signal 1703; a spatial domain inter-component predictor 2402 and an associated spatial domain inter-component prediction / residual recombiner 2342, which receives x in the form of a data stream extracted, dequantized and inverse transformed residual signal 2561, receives y in the form of a data stream extracted, dequantized and inverse transformed residual signal 2562, and outputs z as a residual signal 256'2 that replaces the residual signal 2562; a spectral domain inter-component predictor 2442 and associated spectral domain inter-component prediction / residual recombiner 2382, which receives x in the form of a data stream extracted and dequantized residual signal 2701, receives y in the form of a data stream extracted and dequantized residual signal 2702, and outputs z as the residual signal 170'2 replacing the residual signal 2702; a spatial domain inter-component predictor 2403 and associated spatial domain inter-component prediction / residual recombiner 2343, which receives x in the form of a data stream extracted, dequantized and inverse transformed residual signal 2561 or a data stream extracted, dequantized and inverse transformed residual signal 2562 (or any combination thereof or both), receives y in the form of a data stream extracted, dequantized and inverse transformed residual signal 2563, and outputs z as a residual signal 256'3 replacing the residual signal 2563; a spectral domain inter-component predictor 2443 and associated spectral domain inter-component prediction / residual recombiner 2383, which receives x in the form of a data stream extracted and dequantized residual signal 2701 or a data stream extracted and dequantized residual signal 2702 (or any combination thereof or both), receives y in the form of a data stream extracted and dequantized residual signal 2703, and outputs z as a residual signal 270'3 replacing the residual signal 2703; is embodied in the embodiment shown in FIGS.

[0049] In all of these cases in Figures 1 and 2, inter-plane / component prediction is performed as described in more detail subsequently, and furthermore, in any of these cases, the embodiments of Figures 1 and 2 may be supplemented by the more generalized inter-plane / component prediction module 400 of Figure 4. Note that only some of the cases may actually be used.

[0050] As shown in Figure 4, the inter-plane / component prediction module 400 or reconstruction module 400 has two inputs and one output, and optionally uses prediction parameters signaling between the encoder and decoder sides. The first input 402 is shown as "x" and represents the received reconstructed first component signal based on which the inter-plane / component redundancy reduction for the second component is performed by the reconstruction module 400. This reconstructed first component signal 402 may be a transmitted residual signal extracted from the data stream for the portion of the first component that is co-located in the spatial or spectral domain with the portion of the second component that is currently the subject of inter-plane / component redundancy reduction, as shown in Figures 1 and 2.

[0051] Another input signal 404 of the reconstruction module 400, denoted as "y", further represents the transmitted residual signal of the portion of the second component currently being subjected to inter-plane / inter-component redundancy reduction by the module 400 in the same domain as the signal 402, i.e., the spectral or spatial domain. The reconstruction module 400 also reconstructs a second component signal 406 in the same domain, denoted as "z" in FIG. 4, which component signal 406 represents the main output of the reconstruction module 400 and which also participates at least in reconstructing the dependent components of the multi-component image by replacing x. "At least" means that the component signal 406 output by the reconstruction module 400 may represent a residual prediction and must accordingly still be combined with the prediction signal 154i of each dependent component i, as shown in FIGS. 1 and 2.

[0052] Like other modules in the encoder and decoder, the reconstruction module 400 operates on a block basis. This can manifest itself, for example, in block-wise adaptation of the inter-component redundancy reduction inversion performed by the reconstruction module 400. "Block-wise adaptation" can optionally include explicit signaling of the prediction parameters 146 / 148 in the data stream. However, backward adaptive setting of the parameters for controlling the inter-component redundancy reduction is also possible. That is, with reference to FIG. 4, in the case of a reconstruction module 400 incorporated in a decoder, the prediction parameters input the regulation module and therefore represent its further input, while in the case of a reconstruction module 400 incorporated in an encoder, the prediction parameters are determined internally in the manner exemplified below, including, for example, the solution of the LSE optimization problem.

[0053] TIFF2026035646000002.tif47170

[0054] 5 shows a currently reconstructed portion or block 440 of the second component 108. FIG. 5 also shows a spatially corresponding portion / block 442 of the first component 106, i.e., the portion spatially co-located in the image 10. The input signals 402 and 404 that module 400 receives for components 106 and 108 with respect to the co-located blocks 440 and 442 may represent the residual signals as transmitted for components 106 and 108 in the data stream, as outlined with respect to FIGS. 1 and 2. With respect to the second component 108, module 400 calculates z. For each block, such as block 440, the parameters α, β, and γ, or simply a subset thereof, are adapted in a manner further exemplified below.

[0055] In particular, for a given block, such as block 440, whether inter-component prediction is performed may be signaled in the data stream by a syntax element. The parameters α, β, and γ in the case of inter-component prediction being switched on merely represent possible examples. For block 440 in which inter-component prediction is applied, the prediction mode may be signaled in the data stream, the prediction source may be signaled in the data stream, the prediction domain may be signaled in the data stream, and further, parameters related to the above-mentioned parameters may be signaled in the data stream. The meanings of "prediction mode," "prediction source," "prediction domain," and "related parameters" will become clear from the carryover explanation below. In the example described so far, inter-component prediction operates on the residual signal. That is, x and y are prediction residuals as transmitted in the data stream, and furthermore, both represent prediction residuals of hybrid prediction. As mentioned above, x and y may be prediction residuals in the spatial domain or in the frequency domain in the exemplary case of using transform coding as outlined above. Applying prediction at the encoder or decoder stage has several advantages. First, additional memory is usually unnecessary. Second, inter-component prediction can be performed locally, i.e., without introducing an additional intermediate step after the analysis process from the decoder's perspective. To distinguish the prediction domain, an additional syntax element can be transmitted in the bitstream. That is, the latter additional syntax element can indicate whether the inter-component prediction domain can be the spatial domain or the spectral domain. In the first case, x, y, and z are in the spatial domain, and in the latter case, x, y, and z are in the spectral domain. It should be noted that from the decoder's perspective, the residuals reconstructed from the bitstream may be different from those generated in the encoder before the quantization step. However, it is preferable to use the already quantized and reconstructed residual as a prediction source in an encoder implementing an embodiment of the present application.Furthermore, in the case of skipping the transform stage, the inter-component prediction in the spatial and frequency domains is exactly the same. For such a configuration, the signaling of the prediction domain, i.e., spatial or frequency domain, can be skipped.

[0056] TIFF2026035646000003.tif38169

[0057] To keep the processing chain as simple as possible, an example configuration may keep the processing of the first component, such as the luma component, unchanged and further use the luma-reconstructed residual signal as a predictor for the component residual signal. Note that this is a possible prediction source configuration, and furthermore, such a simple prediction source simplifies the general transform approach; all three components or planes of the input signal are needed to generate the transform samples.

[0058] Another possible configuration is to perform prediction source adaptation, i.e., signal that all available residual signals or each reconstructed residual component is used for prediction. Alternatively, the processing order can be changed locally; for example, the second component is reconstructed first and then used as a prediction source for the remaining components. Such a configuration benefits from the fact that the delta between two components using bijective (or approximately invertible) predictors is the same, but for opposite codes, the absolute cost for coding the prediction source is different. Furthermore, combining several prediction sources is possible. In this case, the combining weights can be transmitted in the bitstream or estimated in a backward-driving manner using available or respective coded neighborhood statistics.

[0059] The specification of model parameters can be performed with backward driving, consisting of backward driving estimation and forward signaling, or can be fully forward signaled in the data stream. An example configuration is to use a fixed set of model parameters known to both the encoder and decoder and further signal a set index to the decoder for each block or shape. Another configuration is to use a dynamic set or list, where the order of predictors is changed after a certain number of blocks or shapes. Such an approach enables more local adaptation to the source signal. More detailed examples on prediction modes, prediction sources, prediction regions, and parameter signaling are carried forward below.

[0060] Regarding prediction modes, the following may be noted:

[0061] The prediction mode can be an affine, non-linear, or more complex function realized by approaches such as spline or support vector regression.

[0062] Note that color space conversion is mostly linear, using all available input components. That is, color space conversion tends to map a three-component vector of three color components to another vector of another three components in another color space. The encoder and decoder, according to embodiments of the present application, can operate independently of the input color space, and thus the luma component can be kept unchanged to form the prediction source. However, according to alternative embodiments, the definition of "luma component" or "first component" can vary from block to block, i.e., from block to block or shape to shape, such as a prediction or transform block (or rectangular shape), and the component serving as the "first component" can be adaptively selected. The adaptation can be indicated to the encoder by signaling in the data stream. For example, FIG. 5 shows the case where component 106 forms the "first component" and component 108 is the inter-predicted component, i.e., the second component, while block 440, as far as it is concerned, can be different for another block of image 10, e.g., component signal 106 is inter-predicted from component 108. Block-wise adaptation of the "prediction source" can be indicated in the data stream, as just outlined. Additionally or alternatively, for each block 440, a syntax element signaling whether prediction is applied or not can be transmitted in the data stream. That is, in the case where component 106 is always used as the "first component," such a syntax element is present only for the dependent components, i.e., components 108 and 110 in FIGS. 1 and 2. In that case, the first component, e.g., the luma component, provides the prediction source, i.e., its residual signal in the case of the embodiment outlined above.

[0063] Additionally or alternatively, if prediction is enabled for a particular block 440, the prediction mode may be signaled in the data stream for that block 440. Note that prediction may be skipped for the case of a zero-valued reconstructed first component signal, i.e., a zero-valued residual luma signal, in the case where the prediction residual is used as the basis for inter-component prediction. In that case, the above-mentioned syntax element signaling whether inter-component prediction is applied or not may be omitted, i.e., not present, in the data stream for the respective block 440.

[0064] In the case of a combined backward driving and forward signaling approach, parameters derived from the respective reconstructed data already coded can serve as starting parameters. In such a case, deltas compared to the selected prediction mode can be transmitted in the data stream. This can be achieved by calculating optimal parameters for a fixed or adapted prediction mode, and the calculated parameters are then transmitted in the bitstream. Another possibility is to transmit some deltas compared to the starting parameters derived by using a backward driving selection approach, or by always using parameters calculated and selected by the backward driving approach only.

[0065] TIFF2026035646000004.tif58170

[0066] Note that a single element in y is replaced after prediction by z. In other words, the reconstruction module 400 receives the corrected signal 404 and replaces it with z 406. Also, note that when using such a configuration, inter-component prediction is simplified to an addition operation: the fully reconstructed residual sample value of the first component (α=1) or its half value (α=0.5) is added to the corrected sample value 404. The halving can be generated by a simple right-shift operation. In that case, for example, the configuration can be implemented by implementing a multiplication between x and α in the predictor within the reconstruction module 400 and further implementing an addition in the adder shown in FIG. 4. In other words, in the given example, on the encoder side, the operation mode is as follows:

[0067] TIFF2026035646000005.tif75170

[0068] TIFF2026035646000006.tif70170

[0069] The parameters for the first sample or group of samples can be initialized by some default values ​​or calculated from adjacent blocks or shapes. Another possibility is to transmit optimal parameters for the first sample or group of samples. To use as many previous samples as possible, a predetermined scan pattern can be used to map the two-dimensional residual block (or rectangular shape) into a one-dimensional vector. For example, the sample or group of samples can be scanned vertically, horizontally, or in a direction similar to the scan direction of the transform coefficients. The specific scan can also be derived by a backward driving scheme, or it can be signaled in the bitstream.

[0070] Another extension is the combination of the already shown example and a transform using all three available components as input. Here, in this example, the residual signal is transformed using a transform matrix that removes correlation for a given transform block. This configuration is useful when the correlation between planes is very large or extremely small. An example configuration uses a Y'CoCg transform in the case of an input color space such as R'G'B' or a principal component analysis approach. For the latter case, the transform matrix must be signaled to the decoder in a forward manner or using a predetermined set and rules known to the encoder and decoder to derive the matrix values. Note that this configuration requires the residual signals of all available components or planes.

[0071] Regarding the prediction domain, the following is noted:

[0072] The prediction domain can be, as mentioned above, the spatial domain, i.e., operating on the residual, or the frequency domain, i.e., operating on the residual after applying a transform such as a DCT or DST. Furthermore, the domain can be both, by transmitting information to the decoder. In addition to the domain for prediction, parameters related to the domain can be transmitted or derived by a backward driving scheme.

[0073] Related to the domain parameters is additional subsampling of the chroma components, i.e., the chroma block (or rectangular shape) is scaled down horizontally, vertically, or both. In such cases, the prediction source may be downsampled as well, or a set of prediction modes must be selected that considers different resolutions in the spatial domain, the frequency domain, or both. Another possibility is to upscale the prediction target so that the dimensions of the prediction source and the prediction target match each other. Downscaling further improves compression efficiency, especially for very flat regions in an image or video. For example, the prediction source contains both low and high frequencies, but the chroma block (or rectangular shape) contains only low frequencies. In this example, subsampling of the prediction source removes high frequencies, and further, less complex prediction modes can be used, further resulting in less information that must be transmitted to the decoder. Note that signaling of downscaling use can be done for each transform block, for each prediction block, or even for a group of prediction blocks or for the entire image.

[0074] In addition to the additional downsampling approach, bit depth adjustments can also be transmitted to the decoder. This occurs when the sample precision differs along different components. One possible method is to reduce or increase the number of bits for the source. Another possible configuration could be to increase or decrease the target bit depth and then correct the final result back to the correct bit depth. A further option is the use of a set of predictors suited to different bit depths. Such prediction modes take into account different bit depths depending on the prediction parameters. The signaling level for bit depth correction can be done for each block or shape or for the entire image or sequence, depending on the content change.

[0075] The following forecast sources are noted:

[0076] The prediction source can be the first component or all available components. For each block (or rectangular shape), the prediction source can be signaled by the encoder. Alternatively, the prediction source can be a dictionary containing all possible blocks of prediction or transform blocks from all available components, and the prediction source to be used for prediction is signaled to the decoder by a syntax element in the bitstream.

[0077] Regarding the parameter derivation, the following is noted:

[0078] TIFF2026035646000007.tif50170

[0079] TIFF2026035646000008.tif33153

[0080] TIFF2026035646000009.tif53170

[0081] TIFF2026035646000010.tif61131

[0082] TIFF2026035646000011.tif28129

[0083] Note that in the above minimization problem, y represents the residual signal that is losslessly transmitted in the data stream for each inter-component predicted block 440. In other words, y is the correction signal for the inter-component predicted block 440. It can be determined iteratively while performing the solution of the LSE optimization problem outlined above at each iteration. In this way, the encoder can optimally decide whether to perform inter-component prediction or not when selecting inter-component prediction parameters, such as prediction mode, prediction source, etc.

[0084] Regarding parameter signaling, the following is noted:

[0085] The prediction itself may be switchable, and a header flag specifying the use of residual prediction should be transmitted at the beginning of the bitstream. When prediction is possible, a syntax element specifying its local use is embedded in the bitstream for block 440, which may be a residual block (or rectangular shape), a transform block (or rectangular shape), or a group of transform blocks (or rectangular shapes), as described above. The first bit may indicate whether prediction is enabled, and the following bits may indicate the prediction mode, prediction source, or prediction domain and related parameters. Note that one syntax element can be used to enable or disable prediction for both chroma components. It is also possible to separately signal the use of prediction and the prediction mode and source for each second (e.g., chroma) component. Also, to achieve high adaptation to the residual signal, a sorting criterion may be used to signal the prediction mode. For example, the most used prediction mode is 1, and the sorted list includes mode 1 at index zero. Then, only one bit is needed to signal the most likely mode 1. Furthermore, the use of prediction may be limited in the case of applying inter-component prediction to the residual; correlation may be high if the same prediction mode for generating the residual is used among different color components. Such a restriction is useful for intra-predicted blocks (or rectangular shapes). For example, this inter-component prediction may be applied only if the same intra-prediction mode used for the block (or rectangular shape) 442, which may be luma, is used for the block (or rectangular shape) 440, which may be chroma.

[0086] That is, in the latter case, the block-wise adaptation of the inter-component prediction process involves checking whether block 440 is relevant to a spatial prediction mode as far as prediction by predictor 2222 is concerned, and whether the spatial prediction mode does not deviate more than a predetermined amount from the spatial mode with which the corresponding or co-located block 442 is predicted by predictor 2221. The spatial prediction mode may comprise, for example, a spatial prediction direction in which neighboring blocks 440 and 442 of already reconstructed samples, respectively, are extrapolated to blocks 440 and 442, respectively, to yield respective prediction signals 2202 and 2201, respectively, which are combined with z and x, respectively.

[0087] In this embodiment, the prediction source is always the luma component. In this preferred embodiment, depicted in Figures 1 and 2, when the interconnection between components 108 / 208 and 110 / 210 is discontinued, the luma processing remains unchanged, and the resulting reconstructed residual of the luma plane is used for prediction. As a result, the prediction source is not transmitted in the bitstream.

[0088] In another embodiment, a prediction source is transmitted for a block or shape 440, such as a residual block, a group of residual blocks or shapes, for example, for a size at which intra or inter prediction is applied.

[0089] For example, the prediction source is luma for the first chroma component and luma or the first chroma component for the second chroma component. This preferred embodiment is similar to the configuration that allows all available planes as prediction sources and also corresponds to Figures 1 and 2.

[0090] To illustrate the above-described embodiment, reference is made to FIG. 6. A three-component image 102 / 202 is shown. The first component 106 / 206 is encoded / decoded using hybrid coding without any inter-component prediction. As for the second component 108 / 208, it is divided into blocks 440, one of which is exemplarily shown in FIG. 6, just as was the case in FIG. 5. In units of these blocks 440, inter-component prediction is adapted, for example, with respect to α, as described above. Similarly, as for the third component 110 / 210, the image 102 / 202 is divided into blocks 450, one such block is exemplarily shown in FIG. 6. As described with respect to the two alternatives just outlined, it may be the case that inter-component prediction of block 450 is necessarily performed by using the first component 106 / 206, i.e., the co-located portion 452 of the first component 106 / 206, as a prediction source. This is indicated using solid arrow 454. However, according to the second alternative just described, syntax element 456 in data stream 104 switches between using the first component as the prediction source, i.e., 454, and using the second component as the prediction source, i.e., inter-component prediction of block 450 based on co-located portion 458 of second component 108 / 208 as indicated by arrow 460. In Figure 6, the first component 106 / 206 may, for example, be the luma component, while the other two components 108 / 208 and 110 / 210 may be chroma components.

[0091] 6, it should be noted that the division into blocks 440 for the second component 108 / 208 could conceivably be signaled in the data stream 104 in a manner independent of, or allowing deviation from, the division of the components 110 / 210 into blocks 450. Of course, the division could equally well be the same and even be taken from one signaled form or used for the first component as is the exemplary case in FIG. 10 described later.

[0092] A further embodiment of the embodiment outlined above is that all components of the image 110 / 210 can alternately serve as prediction sources. See FIG. 7. In FIG. 7, all components 106-110 / 206-210 share a common division into associated blocks 470, with one such block 470 exemplarily shown in FIG. 7. The division or subdivision into blocks 470 can be signaled within the data stream 104. As already mentioned above, block 470 may be a residual or transform block. However, here, any component can form the "first component," i.e., the prediction source. Syntax element 472 and data stream 104 indicate for block 470 which of the components 106-110 / 206-210 forms the prediction source for, for example, the other two components. FIG. 7 also shows another block 472, for example, in which the prediction source is selected differently, as indicated by the arrows in FIG. 7. Within block 470, a first component serves as a prediction source for the other two components, in which case in block 472 the second component assumes the role of the first component.

[0093] 6 and 7, it should be noted that the extent to which syntax elements 456 and 472 indicate the prediction source can be selected differently, i.e., for each block 450 and 470 / 72 individually, or for a group of blocks or even for the entire image 102 / 202.

[0094] In a further embodiment, the prediction sources may be all available components or a subset of the available components. In this preferred embodiment, the weighting of the sources may be signaled to the decoder.

[0095] In a preferred embodiment, the prediction region is in the spatial domain. In this embodiment, the residual for the entire residual block or shape or only certain parts of the residual block or shape can be used, depending on the signaling configuration. The latter case is given when predictions are signaled individually for each transform block or shape, and allows for further subdivision of the residual block or shape into smaller transform blocks or shapes.

[0096] In a further embodiment, the prediction domain is in the frequency domain. In this preferred embodiment, the prediction is tied to the transform block or shape size.

[0097] In a further embodiment, the prediction region may be in the spatial or frequency domain, where the prediction region is identified separately by forward signaling or backward driving estimation depending on local statistics.

[0098] TIFF2026035646000012.tif97170

[0099] Also, syntax element 490 can signal the region used for inter-component prediction for an individual block 440, for a group of blocks 440, or for the entire image, or even to a larger extent, e.g., for a group of images.

[0100] In a further embodiment, both prediction domains are included in the prediction process: in this preferred embodiment of the invention, prediction is first done in the spatial domain and then a further prediction is applied in the frequency domain with both predictions using different prediction modes and sources.

[0101] In an embodiment, a chroma block or shape can be subsampled horizontally, vertically, or both by some factor. In this embodiment, the downscale factor can be equal to a power of 2. The use of a downsampler is signaled as a syntax element in the bitstream, and further, the downsampler is fixed.

[0102] In a further embodiment, a chroma block or shape can be subsampled horizontally, vertically, or both by some factor, which can be transmitted in the bitstream, and further, the downsampler is selected from a set of filters, and the exact filter can be addressed by an index transmitted in the bitstream.

[0103] In a further embodiment, the selected upsampling filter is transmitted in the bitstream, in which the chroma blocks may be subsampled first, and therefore, in order to use prediction with matching block or rectangle sizes, upsampling must be performed before prediction.

[0104] In a further embodiment, the selected downsampling filter is transmitted in the bitstream, in which the luma is downsampled to achieve the same block or rectangle size.

[0105] In an embodiment, a syntax element is signaled to imply bit correction when the source and target bit depths are different. In this embodiment, the luma precision can be reduced, or the chroma precision can be increased to have the same bit depth for prediction. In the latter case, the chroma precision is reduced back to the original bit depth.

[0106] In an embodiment, the number of prediction modes is two, and the set of predictors is defined exactly as in the given example.

[0107] In a further embodiment, the number of prediction modes is 1, and furthermore the configuration is the same as described in the previous embodiment.

[0108] In a further embodiment, the number of predictors is freely adjustable, and the set of predictors is precisely defined as in a given example. This embodiment is a more general description of the example with α=1 / m, where m>0 indicates the prediction number or mode. Thus, m=0 indicates that prediction should be skipped.

[0109] In a further embodiment, the prediction mode is fixed, i.e., prediction is always enabled. For this embodiment, it may enable adaptive inter-plane prediction and further set the number of predictors equal to zero.

[0110] In a further embodiment, prediction is always applied, and further, prediction parameters such as α are derived from neighboring blocks or shapes. In this embodiment, the optimal α for a block or shape is calculated after full reconstruction. The calculated α serves as a parameter for the next block or shape in the local neighborhood.

[0111] In a further embodiment, a syntax element is transmitted in the bitstream indicating the use of parameters derived from a local neighborhood.

[0112] In a further embodiment, the parameters derived from the neighborhood are always used. In addition, the delta compared to the optimal parameters calculated in the encoder can be transmitted in the bitstream.

[0113] In a further embodiment, the backward drive selection scheme for parameters is disabled and the optimal parameters are transmitted in the bitstream.

[0114] In a further embodiment, the use of starting α and the presence of delta α in the bitstream are signaled separately.

[0115] In an embodiment, the signaling of prediction modes, prediction sources, and prediction parameters is restricted to the same regular prediction mode, and information related to inter-plane prediction is transmitted only when the intra prediction mode for the chroma component is the same as that used for the luma component.

[0116] In a further embodiment, the block is divided into windows of different sizes, and further, the parameters for the current window are derived from a reconstructed previous window within the block, hi a further embodiment, the parameters for the first window are derived from a reconstructed adjacent block.

[0117] In a further embodiment, a syntax element is transmitted in the bitstream indicating the use of parameters derived from the local neighborhood used for the first window.

[0118] In a further embodiment, the window can be scanned vertically, horizontally or vertically.

[0119] In a further embodiment, the parameters of the current window are derived from a previous window, where the previous window is determined according to the scan position of the transform coefficient sub-block.

[0120] In a further embodiment, the window scanning is limited to one scan direction.

[0121] In a further embodiment, the parameters are derived using an integer implementation by using lookup tables and multiplication operations instead of division.

[0122] In an embodiment, a global flag transmitted in the bitstream header indicates the use of adaptive inter-plane prediction. In this embodiment, the flag is embedded at the sequence level.

[0123] In a further embodiment, the global flag is transmitted in the bitstream header with embedding at the picture parameter level.

[0124] In a further embodiment, the number of predictors is transmitted in the bitstream header, where the number zero indicates that prediction is always enabled, and a number not equal to zero indicates that the prediction mode is selected adaptively.

[0125] In an embodiment, the set of prediction modes is derived from the number of prediction modes.

[0126] In a further embodiment, the set of prediction modes is known to the decoder and specifies all model parameters for prediction.

[0127] In a further embodiment, the prediction modes are all linear or all affine.

[0128] In an embodiment, the set of predictors is hybrid, i.e., it includes simple prediction modes that use other planes as prediction sources, as well as more complex prediction modes that use all available planes, and also transforms the input residual signal into another component or plane space.

[0129] In an embodiment, the use of prediction is specified per transform block or shape per chroma component, in which case this information may be skipped when the luma component consists of a zero-valued residual at the same spatial location.

[0130] In an embodiment, the modes are transmitted using a truncated unary decomposition. In this embodiment, a different context model is assigned per bin index, but limited to a specific number, e.g., 3. Furthermore, the same context model is used for both chroma components.

[0131] In a further embodiment, different chroma planes use different sets of context models.

[0132] In a further embodiment, different transform block or shape sizes use different sets of context models.

[0133] In a further embodiment, the mapping of bins to prediction modes is dynamic or adaptive, in which, from the decoder's perspective, a decoded mode equal to zero indicates the most used mode up to the time of decoding.

[0134] In a further embodiment, the prediction mode and, when using a configuration that allows different prediction sources, the prediction source is transmitted for the residual block or shape. In this embodiment, different block or shape sizes can use different context models.

[0135] The embodiments described below relate in particular to examples for methods of encoding prediction parameters for the "cross-component decorrelation" described so far.

[0136] Although not limited thereto, the following description may be considered to refer to an alternative in which a dependent (second) component signal is reconstructed based on a reference (first) component signal x and a residual (correction) signal y via the calculation z = αx + y, using z as a prediction of the dependent component signal. For example, prediction may be applied in the spatial domain. As in the above example, inter-component prediction may be applied to the hybrid-coded residual signal, i.e., the first and second component signals may represent the hybrid-coded residual signal. However, the following embodiments focus on the signaling of α when this parameter is coded on a sub-picture basis, e.g., per residual block into which a multi-component image is subdivided. The following points out that the signalable state of α should preferably be variable, also to take into account the fact that the range of optimal values ​​for α depends on the type of image content, which in turn varies in a range / unit larger than the residual block. Thus, in principle, the details described below regarding the transmission of α may also be transferred to the other embodiments outlined above.

[0137] Cross-component decorrelation (CCD) approaches exploit the remaining dependence between different color components, enabling higher compression efficiency. An affine model can be used for such an approach, and furthermore, the model parameters are transmitted in the bitstream as side information.

[0138] To minimize the side information cost, only a limited set of possible parameters is transmitted. For example, a possible CCD implementation in High Efficiency Video Coding (HEVC) may use a linear prediction model instead of an affine prediction model, and furthermore, only the mode parameter, i.e., the slope or gradient parameter α, may be limited in the range from 0 to 1 and may further be non-uniformly quantized. In particular, the limited set of values ​​for α may be α∈{0, ±0.125, ±0.25, ±0.5, ±1}.

[0139] The choice of such quantization for the linear prediction model parameters may be based on the fact that the distribution of α is symmetrically centered around the value 0 for natural video content stored in the Y'CbCr color space. In Y'CbCr, the color components are decorrelated by using a fixed transformation matrix to convert from R'G'B' before entering the compression stage. Due to the fact that global transformations are often suboptimal, the CCD approach can achieve higher compression efficiency by eliminating any remaining dependencies between the different color components.

[0140] However, such an assumption does not hold for different types of content, especially for natural video content stored in the R'G'B' color space domain. In this case, the gradient parameter α is often centered around the value 1.

[0141] Similar to the case given above, the distribution becomes completely different when the CCD is extended to the first chroma component as a prediction source. Therefore, it may be beneficial to adjust the CCD parameters according to the given content.

[0142] For example, for Y'CbCr, the quantization of α can be set to (0, ±0.125, ±0.25, ±0.5, ±1), while for R'G'B', the quantization of α can be the opposite, i.e., (0, ±1, ±0.5, ±0.25, ±0.125). However, allowing different entropy coding paths introduces additional problems. One problem is that the implementation becomes more expensive in terms of area and speed for both hardware and software. To avoid this drawback, parameter ranges can be specified at the Picture Parameter Set (PPS) level, where the use of a CCD is also indicated.

[0143] That is, the syntax element signaling α is transmitted at a sub-picture level / granularity, e.g., individually for each residual block. For example, it may be called res_scale_value. It may be coded, for example, using (truncated) unary binarization combined with binary arithmetic coding of the bin sequence. The mapping of (non-binarized) values ​​of res_scale_value to α may be performed such that the mapping is varied using pps, i.e., for the complete image, or even over a larger range, e.g., on a per image sequence basis. The variations may vary the number of representable α values, the order of representable α values, and the selection of representable α values, i.e., their actual values. While simply allowing switching the order among representable alpha values ​​or limiting representable alpha values ​​to positive or negative values ​​is one way to provide content adaptability, the embodiments outlined below further enable much increased flexibility in changing the mapping from res_scale_value signaled at sub-picture granularity to alpha values ​​(representable set of alpha values), e.g., changing the size, members and member order of the joint region of the mapping, with only minor increased overhead; furthermore, the advantages offered by this provision in terms of bits saved for transmitting res_scale_value are seen through a typical mix of video content encoded in YCC or RGB, and are found to overcompensate the need to signal mapping changes.

[0144] The range specification for α can be done, for example, as follows: In the case of Y′CbCr, a suitable subset can be, for example, (0,±0.125,±0.25,±0.5), while for R′G′B′ it can be (0,±0.5,±1) or (0,0.5,1) or (0,0.5,1,2). To achieve the aforementioned behavior, the range can be specified in the PPS using two syntax elements representing two values. Considering the above example and the fact that prediction is performed with three-point accuracy, i.e., the predicted sample value is multiplied by α and then shifted right by 3, the range configuration for Y′CbCr can be transmitted as [−3,3].

[0145] However, with such signaling, only the second and third cases for R'G'B' can be achieved using [2,3] and [2,4]. To achieve the first example for R'G'B', the codes must be separated using additional syntax. Furthermore, it is sufficient to send a delta for the second value with the first value serving as the starting point. For this example, the second R'G'B' configuration is [2,1] instead of [2,3].

[0146] In the case of prediction from the first chroma component, the range values ​​can be specified separately for each chroma component. Note that this can be done even without support for prediction from the first chroma component.

[0147] Given the constraints specified in PPS, the analysis and reconstruction of the prediction parameter α is modified as follows: For the constraints, i.e., α∈{0,±0.125,±0.25,±0.5,±1} and for the case without three-point accuracy, the final α F is rearranged as follows, where α P denotes the value parsed from the bitstream: α F =1<<α P This is α F =1<<(α P +α L), where α L indicates the offset for the smallest absolute value when the range is entirely in the positive or negative range.

[0148] Both values ​​are P is used to derive a limit on the number of bins parsed from the bitstream when is binarized using a truncated unary code. Note that the parsing of the sign may depend on a given range. A better way to exploit this aspect is to encode the sign before the absolute value of α. After encoding the sign, the maximum number of bins parsed from the bitstream can be derived. This case is useful when the range is asymmetric, for example, when it is [-1,3].

[0149] Often, a different order is desired, e.g., (0,1,0.5) for some R'G'B' content. Such reversal can be achieved by simply setting the range values ​​according to [3,2]. In this case, the number of bins parsed from the bitstream is still 2 (the absolute difference between the two range values ​​is n=1, and the number of bins is always n+1 in the case of truncated unary codes). Reversal can then be achieved in two ways. The first option is to introduce a fixed offset, which is equal to twice the current value if no reversal is desired, or to the maximum representable value in the case of reversal. A second and more elegant way to do this is to expand the range transmitted in the PPS into memory and access the corresponding value through a lookup operation. This approach results in a unified logic and a single path for both cases. For example, the case (0,0.25,0.5,1) is signaled in the PPS by transmitting [1,3], and the memory is then made with the following entries (0,1,2,3). In the other way around, i.e., in the reversed case with the value [3,1] transmitted in the bitstream, the memory is made with the following entries (0,3,2,1). Using such an approach, the final α Fis α F =1<<(LUT[α P ]), where LUT[α P ] indicates a lookup operation.

[0150] To explain the most recently mentioned embodiment in more detail, reference is made to Figures 9a, 9b, and 9c. Figures 9a, 9b, and 9c, hereafter referred to simply as Figure 9, illustratively show one multi-component image 102 / 202 layered, illustratively with a dependent (second) component (108 / 208) followed by a first (reference) component 106 / 206.

[0151] The image 102 / 202 may be part of a video 500 .

[0152] The data stream 104 in which the image 102 / 202 is encoded includes a high-level syntax element structure 510 that relates to at least the entire image 102 / 202, or even to an image sequence from the video 500 that includes the image 102 / 202. This is illustrated in FIG. 9 using curly brackets 502. Additionally, the data stream 104 includes first weight syntax elements at the sub-picture level. One such first weight syntax element 514 is illustratively shown in FIG. 9 as relating to an exemplary residual block 440 of the image 102 / 202. The first weight syntax element is for individually setting a first weight, i.e., α, when inter-component predicting the dependent component 108 / 208 based on the co-located portion of the reference component 106 / 206.

[0153] The first weight syntax element is coded using truncated unary binarization. An example for such a TU2-valuation is shown at 518 in FIG. 9. As shown, such a TU2-valuation consists of a series of bin strings of increasing length. In the manner described above, a high-level syntax element structure 510 defines how the binarized bin string 518 is mapped to possible values ​​for α. An exemplary set of such possible values ​​for α is shown at 520 in FIG. 9. In practice, the high-level syntax element structure 510 defines a mapping 522 from the set of bin strings 518 to a subset of set 520. With this approach, it is possible to keep the signaling overhead for switching between different inter-component prediction weights α at sub-picture granularity with low bit consumption, as sub-picture signaling can simply distinguish between a smaller number of possible values ​​for α.

[0154] As mentioned above, it is possible that the syntax element structure 510 allows a decoder to derive the first and second interval boundaries 524 therefrom. Corresponding syntax elements in the structure 510 can be coded independently / separately or in relation to each other, i.e., differently. The interval boundary values ​​524 identify the elements from the sequence 520, i.e., using the exponential function described above, which can be implemented with respective bit shifts. In this way, the interval boundaries 524 indicate to the decoder the joint region of the mapping 522.

[0155] As mentioned above, the case of α being zero may be signaled to the decoder separately using the respective zero flag 526. If the zero flag has a first state, the decoder sets α to be zero and skips reading any first weight syntax element 514 for the respective block 440. If the zero flag has any other state, the decoder reads the first weight syntax element 514 from the data stream 104 and uses the mapping 522 to determine the actual value of the weight.

[0156] Furthermore, as outlined above, only the absolute portion 528 of the first weight syntax element 514 may be coded using truncated unary binarization, and the sign portion of the first weight syntax element 514 may be coded first. In this way, it is possible for the encoder and decoder to appropriately set the length, i.e., the number of bin sequences, of the binarization 518 for the absolute portion 528 so as to determine whether the sign portion 530 belongs to those members of the joint region 532 of the mapping 522 from the set 520 where the value α of the block 440 has positive α values, or to the other portion consisting of negative α values. Of course, the sign portion 530 of the first weight syntax element 514 may not be present and may not be readable by the decoder, in cases where the decoder derives it from a high-level syntax element structure 510 where the joint region 530 contains only positive α values ​​or only negative α values.

[0157] As will become clear from the above description, the interval boundaries 524 and the order in which they are coded in the high-level syntax element structure 510 can determine the order in which the absolute portion 528 “traverses” the members of the joint region 532.

[0158] With respect to Figure 10, a further specific embodiment is shown that summarizes certain aspects already described above. According to the embodiment of Figure 10, the encoder and decoder commonly subdivide the picture 102 / 202 with respect to all three components as far as hybrid prediction within the predictor 122 / 222 is concerned. In Figure 10, the components are indicated by "1," "2," and "3," respectively, and are further described below the respective indicated images. The subdivision / division of the picture 102 / 202 into prediction blocks 308 may be signaled in the data stream 104 via prediction-related subdivision information 600.

[0159] For each prediction block 308, prediction parameters may be signaled within the data stream 104. These prediction parameters are used for hybrid encoding / decoding of each component of the image 102 / 202. Prediction parameters 602 may be signaled within the data stream 104 for each component individually, commonly for all components, specifically for a partial component, or component-globally. For example, the prediction parameters 602 may distinguish between spatially and / or temporally predicted blocks 308, among others. Furthermore, for example, this indication may be common within a component, and temporal prediction-related parameters within the prediction parameters 602 may be signaled within the data stream 104, specifically within the component. Using the implementation of FIGS. 1 and 2, for example, the predictor 122 / 222 uses the prediction parameters 602 to derive the prediction signal 120 / 220 for each component.

[0160] Furthermore, the data stream 104 signals the subdivision / division of the image 102 / 202 into residual or transform blocks, here indicated by reference sign 604, according to the embodiment of Fig. 10. Residual / transform related subdivision information in the data stream 104 is indicated using reference sign 606. As mentioned above with respect to Fig. 3, the subdivision / division of the image 102 / 202 into prediction blocks 308 on the one hand and residual / transform blocks 604 on the other hand may be at least partially coupled to each other in that the division into residual / transform blocks 604 at least partially forms a hierarchical multi-tree subdivision of the image 102 / 202 into prediction blocks 308 or an extension of several hierarchical intermediate subdivisions, for example into coding blocks.

[0161] For each residual / transform block, data stream 104 may include residual data 6081, 6082, 6083, for example, in the form of quantized transform coefficients. Dequantizing and then inverse transforming the residual data 6081-6083 reveals the residual signal in the spatial domain for each component, i.e., 6101, 6102, and 6103.

[0162] 10, the data stream 104 further includes an inter-component prediction flag 612, which in this example signals globally for the image 102 / 202 whether inter-component prediction is applied / used. If the inter-component prediction flag 612 signals that inter-component prediction is not used, the residual signals 6101-6103 are not combined with each other but are used separately for correcting the prediction signal 120 / 220 for each component. However, if the inter-component prediction flag 612 signals the use of inter-component prediction, the data stream 104 includes, for each residual / transform block, flags 6142, 6143 for each dependent component 2 and 3, which signal whether inter-component prediction is applied for each component 2 / 3. If signaled to be applied, the data stream 104 includes inter-component prediction parameters 6162 and 6163, respectively, for each residual / transform block for each component i=2 / 3, corresponding to, for example, α in the description outlined above.

[0163] Thus, for example, if the flag 6142 indicates for the second component that inter-component prediction is used, the inter-component prediction parameters 6162 indicate the weights to be added to the residual signal 6102 in order to replace the residual signal 6102 with a new residual signal 6102', which is then used instead of the residual signal 6102 to correct the respective prediction signal 1202 / 2202.

[0164] Similarly, when flag 6143 indicates the use of inter-component prediction for the respective residual / transform block, inter-component prediction parameter 6163 indicates a weight α3 by which the residual signal 6101 is added to the residual signal 6103 to result in a new residual signal 6103' that replaces the residual signal 6103 and is used to correct the prediction signal 1203 / 2203 of the third component.

[0165] Inter-component prediction parameters 616 2 / 3 The first flag 614 is followed conditionally by 1 / 2 Instead of transmitting separately, other signaling may be possible. For example, the weights α 2 / 3 Its range of possible values ​​is symmetrically arranged around zero, so that α can be a signed value. 2 / 3 The absolute value of can be used to distinguish between the case of not using inter-component prediction as far as component 2 / 3 is concerned and the case of using inter-component prediction for each component 2 / 3. In particular, when the absolute value is zero, this corresponds to not using inter-component prediction. And each parameter α 2 / 3 Any sign flag signaling for α can be suppressed in the data stream 104. As a reminder, according to the embodiment of FIG. 10, inter-component prediction is varied on a per residual / transform block basis as far as components 2 / 3 are concerned, and the variation is inconsistent with not using inter-component prediction at all. 2 / 3 =0 and α 2 / 3 By changing α 2 / 3 ≠0).

[0166] TIFF2026035646000013.tif24170

[0167] However, it should be noted that the scope of the flag 612 may be chosen differently. For example, the flag 612 may relate to a smaller unit, such as a slice of the image 102 / 202, or to a larger unit, such as a group of images or a series of images.

[0168] TIFF2026035646000014.tif42128

[0169] That is, the above syntax generates, for example, one for each second and third component, such as a chroma component, for each residual or transform block of a double image in the data stream 104, while the luma component forms the base (first) component.

[0170] As shown previously with respect to Figure 10, the syntax example just outlined corresponds to an alternative to the configuration of Figure 10: log2_res_scale_abs_plus1 signals the absolute value of α, and if the syntax element is zero, this corresponds to inter-component prediction not being used for the respective component c. However, if used, res_scale_sign_flag is signaled and also indicates the sign of α.

[0171] The semantics of the syntax elements presented so far can be provided as follows.

[0172] cross_component_prediction_enabled_flag equal to 1 specifies log2_res_scale_abs_plus1 and res_scale_sign_flag may be present in the transform unit syntax for images that reference a PPS. cross_component_prediction_enabled_flag equal to 0 specifies log2_res_scale_abs_plus1 and res_scale_sign_flag is not present for images that reference a PPS. When not present, the value of cross_component_prediction_enabled_flag is inferred to be equal to 0. It is a bitstream conformance requirement that the value of cross_component_prediction_enabled_flag be equal to 0 when ChromaArrayType is not equal to 3.

[0173] log2_res_scale_abs_plus1[c]-1 specifies the logarithm base 2 of the magnitude of the scaling factor ResScaleVal used in cross-component residual prediction. When not present, log2_res_scale_abs_plus1 is inferred to be equal to 0.

[0174] TIFF2026035646000015.tif111170

[0175] In the above, ResScaleVal corresponds to the above-mentioned α.

[0176] TIFF2026035646000016.tif48170

[0177] A right shift ">>3" corresponds to a division by 8. According to the present example, the α values ​​that can be signaled are {0, ±0.125, ±0.25, ±0.5, ±1}, as already exemplified above.

[0178] log2_res_scale_abs_plus1 may be signaled in the data stream using truncated unary binarization and binary arithmetic coding and binary arithmetic decoding and truncated unary inverse binarization, respectively. The binary arithmetic decoding / encoding may be context adaptive. A context may be selected based on a local neighborhood. For example, a different context may be selected for each bin of the binarization of log2_res_scale_abs_plus1. A different set of contexts may be used for both chroma components. Similarly, res_scale_sign_flag may be signaled in the data stream binary arithmetic coding and binary arithmetic decoding, respectively. The binary arithmetic decoding / encoding may be context adaptive. Furthermore, different contexts may be used for both chroma components. Alternatively, the same context is used for both chroma components.

[0179] As described, the mapping from log2_res_scale_abs_plus1 to the absolute value of α, i.e., ResScaleVal>>3, can be done arithmetically using bit-shifting operations, i.e., by exponential functions.

[0180] The signaling of log2_res_scale_abs_plus1 and res_scale_sign_flag for two chroma components can be skipped for a particular residual / transform block if the luma component in the latter is zero, as in the example for signaling 614 and 616 in Figure 10. This allows the decoder to explicitly read sub-picture level syntax elements 6142, 6162, 6143, 6163 from the data stream for the currently decoded portion of the second component 208 and further decode component signal 256' from the spatially corresponding portion 442. 2,3or skip the explicit reading and, optionally, extract the second component signal 256' from the spatially corresponding portion 442. 2,3 , but instead performs a reconstruction of the second component signal 256 2,3 This means that depending on whether the spatially corresponding part 442 of the reconstructed first component signal 2561 is zero, one can possibly check whether it is zero.

[0181] 10 illustrates a method for reconstructing a first component signal 6101 associated with the first component 206 from the data stream 104, and further illustrating a method for reconstructing the first component signal 6101 and a correction signal 610 derived from the data stream. 2,3 a second component signal 610' associated with the second (third) component 208 / 210 of the multi-component image 202 from the spatially corresponding portion 442 of 2,3 6 illustrates a decoder configured to decode a multi-component image 202 that spatially samples a scene with respect to different components 206, 208, 210 by reconstructing a portion 440 of the first component signal 6101. The first component signal 6101 is a prediction residual of a temporal, spatial or inter-view prediction of the first component 206 of the multi-component image 202, and the decoder can further reconstruct the first component 206 of the multi-component image 202 by performing a temporal, spatial or inter-view prediction of the first component 206 and further correcting the temporal, spatial or inter-view prediction of the first component using the reconstructed first component signal 6101. The decoder can further reconstruct the first component 206 of the multi-component image by correcting the temporal, spatial or inter-view prediction of the first component using the reconstructed first component signal 6101 at sub-picture granularity in response to signaling in the data stream. 2,3The decoder is configured to adaptively set α2. For this purpose, the decoder is configured to read the absolute value of the first weight from the data stream at sub-picture granularity and further read the sign of the first weight in a manner that is conditional on whether it is zero. The decoder is configured to skip reading the absolute value of the first weight from the data stream at sub-picture granularity and reading the sign of the first weight in a manner that is conditional on whether it is zero in parts of the first component signal that are zero. When reconstructing the second component signal, the decoder adds a spatially corresponding part of the reconstructed first component signal weighted by the first weight (α2) to the correction signal. The addition can be performed in the spatial domain in a sample-wise manner. Alternatively, it is performed in the spectral domain. The encoder performs it in a prediction loop.

[0182] Although some aspects are described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, with blocks or apparatus corresponding to method steps or features of method steps. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of any of the most important method steps may be performed by such an apparatus.

[0183] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or in software. Implementation can be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, on which electronically readable control signals are stored that cooperate (or can cooperate) with a programmable computer system so that the respective methods are performed. Thus, the digital storage medium may be computer-readable.

[0184] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.

[0185] Generally, embodiments of the present invention can be implemented as a computer program product having program code that operates to perform one of the methods when the computer program product is run on a computer, which program code may for example be stored on a machine-readable carrier.

[0186] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0187] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0188] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium or computer-readable medium) comprising, recorded on it, the computer program for performing one of the methods described herein. The data carrier, digital storage medium or recording medium is typically tangible and / or non-transitory.

[0189] A further embodiment of the inventive method is, therefore, a data stream or sequence of signals representing the computer program for performing one of the methods described herein, which data stream or sequence of signals may for example be arranged to be transmitted via a data communications connection, for example via the Internet.

[0190] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0191] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0192] Further embodiments according to the invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transferring the computer program to the receiver.

[0193] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0194] The apparatus described herein may be implemented using a hardware apparatus, using a computer, or using a combination of a hardware apparatus and a computer.

[0195] The methods described herein may be performed using a hardware apparatus, a computer, or a combination of a hardware apparatus and a computer.

[0196] The above-described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. It is therefore intended that the present invention be limited only by the scope of the claims and not by the specific details set forth herein by way of description and illustration of the embodiments.

Claims

1. A first component signal (256) associated with a first component (206) of the multi-component image (202) from the data stream (104). 1 ;270 1 ) and reconstruct the reconstructed first component signal (256 1 ;270 1 ) and a correction signal (256) derived from said data stream. 2 ;270 2 ) to a second component signal (256') relating to a second component (208) of the multi-component image (202). 2 ;270' 2 1. A decoder configured to decode a multi-component image (202) that spatially samples a scene with respect to different components (206, 208, 210) by reconstructing (400) a portion (440) of the image.

2. The decoder a method for regularly subdividing the multi-component image (202) into tree blocks (302), individually subdividing the tree blocks using recursive multi-tree subdivision into code blocks (304), individually subdividing each code block using recursive multi-tree subdivision into prediction blocks (308) and into residual blocks (312), and further subdividing the residual blocks into transform blocks (316); selecting a prediction mode according to the data stream at a granularity according to the code block or according to the prediction block; setting prediction parameters according to the data stream at the granularity of the prediction block; A prediction signal (220) is generated using the prediction mode and prediction parameters. 1 , 220 2 , 220 3 ) and The residual signal (256 1 , 256' 2 , 256' 3 ) and further Reconstructing the multi-component image (202) by correcting the prediction signal using the residual signal. a block-based hybrid video decoder configured to: The decoder may also include a signaling (614) in the data stream to switch between performing the reconstruction of the second component signals from the spatially corresponding portions of the reconstructed first component signals and the corrected signal at the granularity of the residual block and / or the transform block, and reconstructing the second component signals from the corrected signal without regard to the spatially corresponding portions of the reconstructed first component signals. 2 ;614 3 2. A decoder according to claim 1, responsive to:

3. 3. The decoder of claim 1, wherein the first component signal is a prediction residual of a temporal, spatial or inter-view prediction of the first component of the multi-component image, and wherein the decoder is further configured to reconstruct the first component of the multi-component image by performing the temporal, spatial or inter-view prediction of the first component of the multi-component image and correcting the temporal, spatial or inter-view prediction of the first component using the reconstructed first component signal.

4. The decoder is configured to: determine whether the second component signal is a prediction residual of a temporal, spatial or inter-view prediction of the second component (208) of the multi-component image (202); and further determine whether the second component signal is a prediction residual of a temporal, spatial or inter-view prediction of the second component (208) of the multi-component image (202). 2 10. A decoder according to any preceding claim, configured to reconstruct the second component (208) of the multi-component image (202) by performing the temporal, spatial or inter-view prediction of the first component (208) of the multi-component image (202) and using the reconstructed second component signal to correct the temporal, spatial or inter-view prediction of the multi-component image (202).

5. The decoder performs an inverse spectral transform (226) on the spectral coefficients associated with the second component (208) derived from the data stream (104) to obtain the correction signal in the spatial domain. 2 ) to obtain the correction signal (256 2 10. A decoder according to any preceding claim, configured to obtain

6. The decoder, in reconstructing the second component signal, determines whether the spatially corresponding portions (442) of the reconstructed first component signal have first weights (α) that influence the reconstruction of the second component signal at sub-picture granularity. 2 10. A decoder according to any preceding claim, configured to adaptively set

7. The decoder adjusts the first weight (α 2 7. A decoder as claimed in claim 6, configured to adaptively set .

8. 8. The decoder of claim 6 or claim 7, wherein the decoder is configured to read, at the sub-picture granularity, the absolute value of a first weight from the data stream and further to read the sign of the first weight in a conditional manner on whether it is zero.

9. 9. The decoder of claim 8, wherein the decoder is configured to skip reading the absolute value of the first weight from the data stream at the sub-picture granularity and reading the sign of the first weight in parts where the first component signal is zero in a manner conditional on whether it is zero.

10. 10. A decoder according to claim 6, wherein the decoder is configured to add the spatially corresponding portion (442) of the reconstructed first component signal, weighted by the first weight (α), to the correction signal when reconstructing the second component signal.

11. The decoder of claim 10 , wherein the decoder is configured to perform the summation in the spatial domain in a sample-wise manner when reconstructing the second component signal.

12. The decoder Deriving a high level syntax element structure (510) from the data stream, the high level syntax element structure having at least an image extent; constructing a mapping (522) from a set of regions of a predetermined binarization possible bin sequence (518) to a joint region (520) of possible values ​​of the first weight, at least in the image range; and reading a first weight syntax element (514) from the data stream at sub-picture granularity using the predetermined binarization, and further deriving the first weight by subjecting the bin sequence of the first weight syntax element to the mapping; 12. A decoder as claimed in any one of claims 6 to 11, configured to set the first weight by:

13. 13. The decoder of claim 12, wherein the decoder is configured to derive upper and lower boundaries (524) of a joint region value interval from a predetermined set of possible non-zero values ​​of the first weight from the high-level syntax element structure, and further configured to additionally read, at sub-picture granularity, a zero flag (526) from the data stream indicating whether the first weight should be zero, performing the multiplication conditionally depending on the reading and zero flag of the first weight syntax element when deriving the first weight.

14. 14. The decoder of claim 12 or 13, wherein, in constructing the mapping, the decoder is configured to derive the sign and absolute value of a lower boundary integer value and the sign and absolute value of an upper boundary integer value from the high-level syntax element structure, apply an integer range exponential function to the absolute values ​​of the lower boundary integer value and the upper boundary integer value, and further to obtain the joint range of possible values ​​of the first weight from a joint range of the integer range exponential function excluding zero.

15. 15. The decoder of claim 12, wherein the decoder is configured to use a truncated unary binarization as the predetermined binarization for an absolute value portion of the first weight syntax element, and further configured to read a sign portion (530) of the first weight syntax element from the data stream before the absolute portion (530) of the first weight syntax element when deriving the first weight, and to set a length of the truncated unary binarization of the absolute portion of the first weight syntax element depending on possible values ​​of the sign portion of the first weight and the joint region.

16. 16. The decoder of claim 12, wherein the decoder is configured to derive first and second interval boundaries (524) from the high-level syntax element structure, and further wherein the decoder is configured to use a truncated unary binarization of a TU bin sequence as the predetermined binarization for an absolute value portion of the first weight syntax element, and further wherein, when configuring the mapping, the decoder is configured to reverse the order of the possible values ​​to which the TU bin sequence is mapped so as to traverse the possible values ​​of the joint region in response to a comparison of the first and second interval boundaries.

17. 10. The decoder of claim 1, wherein the decoder is configured to, when reconstructing the second component signal (208), adaptively set a second weight by which the correction signal influences the reconstruction of the second component signal at sub-picture granularity.

18. 10. A decoder according to any preceding claim, wherein the decoder is configured to adaptively set weights of a weighted sum of the correction signal and the spatially corresponding portions (442) of the reconstructed first component signal with sub-picture granularity when reconstructing the second component signal, and further to use, at least for each picture, the weighted sum as a scalar argument of a constant scalar function to obtain the reconstructed second component signal.

19. 10. A decoder according to any preceding claim, wherein the decoder is configured to adaptively set weights of the correction signal, the spatially corresponding portions (442) of the reconstructed first component signal and a constant weighted sum with sub-picture granularity when reconstructing the second component signal, and further to use the weighted sum as a scalar argument of a constant scalar function to obtain the reconstructed second component signal, at least for each picture.

20. 20. A decoder as claimed in claim 18 or claim 19, wherein the decoder is configured to set the weights in a backward-driving manner based on a local neighbourhood.

21. 20. A decoder as claimed in claim 18 or claim 19, wherein the decoder is configured to correct the weights in a forward-driven manner and to set the weights in a backward-driven manner based on a local neighbourhood.

22. A decoder according to claim 20 or claim 21, wherein the decoder is configured to set the weights in a backward driving manner based on attributes of already decoded parts of the multi-component image.

23. The decoder At the first spatial granularity, a combined backward and forward adaptive method, or Forward adaptive methods, In a backward adaptive way, setting the weights to default values; and and refining the weights in a backward-driven manner based on a local neighborhood at a second spatial granularity that is finer than the first spatial granularity. A decoder according to any one of claims 18 to 22, arranged to:

24. 24. A decoder according to any of claims 16 to 23, wherein the decoder is configured to set the weights to one of m different states in response to m array sub-picture level syntax elements, the decoder being configured to derive m from a higher level syntax element (510).

25. The decoder, when reconstructing the second component signal, performs, at sub-picture granularity: performing an inverse spectral transform on the spectral coefficients for the second component (208) derived from the data stream to obtain the correction signal (x) in the spatial domain, and reconstructing (400) the second component signal (z) using the correction signal (x) in the spatial domain; 10. A decoder according to any preceding claim, configured to adaptively switch between deriving the correction signal (x) in the spectral domain from the data stream, reconstructing (400) the second component signal (z) in the spectral domain using the correction signal (x) obtained in the spectral domain, and subjecting the reconstructed second component signal (z) to an inverse spectral transform in the spectral domain.

26. 26. The decoder of claim 25, wherein the decoder is configured to perform the adaptive switching in a backward-adaptive and / or forward-adaptive manner (490).

27. 10. A decoder according to any preceding claim, wherein the decoder is configured to adaptively switch, with sub-picture granularity, a direction of reconstruction of the second component signal between performing the reconstruction of the second component signal from the spatially corresponding parts of the reconstructed first component signal and reversing the reconstruction to reconstruct the first component signal from spatially corresponding parts of the reconstructed second component signal.

28. 28. A decoder according to any of claims 1 to 27, wherein the decoder is configured to adaptively switch a direction of reconstruction of the second component signal between performing the reconstruction of the second component signal from the spatially corresponding portions of the reconstructed first component signal and reversing the reconstruction to reconstruct the first component signal from the spatially corresponding portions of the reconstructed second component signal in response to a syntax element (472) signaling an order among the first and second component signals.

29. 10. A decoder according to any preceding claim, wherein the decoder is configured to adaptively switch, at sub-picture granularity, the reconstruction of the second component signal between reconstructing it based on only the reconstructed first component signal and reconstructing it based on the reconstructed first component signal and a third reconstructed component signal.

30. 10. A decoder according to any preceding claim, wherein the first and second components are colour components.

31. 10. A decoder according to any preceding claim, wherein the first component is a luma component and the second component is a chroma component.

32. The decoder, in response to the first syntax element (612), and enabling the reconstruction of the second component signal based on the reconstructed first component signal, and extracting sub-picture level syntax elements from the data stream when parsing the data stream (614). 2 , 616 2 , 614 3 , 616 3 ), and further adapting the reconstruction of the second component signal based on the reconstructed first component signal at sub-picture granularity based on the sub-picture level syntax elements; and 10. A decoder according to any preceding claim, further comprising: a decoder configured to disable the reconstruction of the second component signal based on the reconstructed first component signal; and further configured to be responsive to a first syntax element (612) in the data stream to alter the parsing of the data stream to address the data stream not including the sub-picture level syntax element.

33. The decoder is a backward driven decoder. enabling the reconstruction of the second component signal based on the reconstructed first component signal; 10. A decoder as claimed in any preceding claim, configured to locally switch between disabling the reconstruction of the second component signal based on the reconstructed first component signal and locally switching between disabling the reconstruction of the second component signal based on the reconstructed first component signal.

34. 34. The decoder of claim 33, wherein the decoder is configured to perform the local switching in a backward driving manner.

35. the decoder is configured to reconstruct the first component of the multi-component image by performing the temporal, spatial or inter-view prediction of the first component of the multi-component image, such that the first component signal is a prediction residual of a temporal, spatial or inter-view prediction of the first component of the multi-component image, and by correcting the temporal, spatial or inter-view prediction of the first component using the reconstructed first component signal; the decoder is configured to reconstruct the second component of the multi-component image such that the second component signal is a prediction residual of a temporal, spatial or inter-view prediction of the second component of the multi-component image, and by performing the temporal, spatial or inter-view prediction of the second component of the multi-component image, and further by correcting the temporal, spatial or inter-view prediction of the multi-component image using the reconstructed second component signal; moreover The decoder is configured to perform the local switching by locally checking whether the first and second component signals are prediction residuals of spatial prediction and whether intra-prediction modes of the spatial prediction match, or by locally checking whether the first and second component signals are prediction residuals of spatial prediction and whether intra-prediction modes of the spatial prediction do not differ by more than a predetermined amount.

34. A decoder according to claim 33.

36. 34. The decoder of claim 33, wherein the decoder is configured to initially determine the local switch in a backward-driven manner, modifying the decision in a forward-adaptive manner in response to signaling in the data stream.

37. The decoder is responsive to a second syntax element in the data stream and, in response to the second syntax element, reading sub-picture level syntax elements from the data stream when parsing the data stream, and adapting the reconstruction at sub-picture granularity of the second component signal based on the reconstructed first component signal based on the sub-picture level syntax elements; and 10. A decoder according to any of the preceding claims, wherein the reconstruction of the second component signal based on the reconstructed first component signal is performed non-adaptively.

38. 10. A decoder according to any preceding claim, wherein the first and second components (206, 208) are two of three color components, and the decoder is configured to also reconstruct a third component signal for a third color component (210) of the multi-component image (202) from spatially corresponding portions of the reconstructed first or second component signal and a correction signal derived from the data stream for the third component, and wherein the decoder is configured to adaptively perform the reconstruction at sub-picture level of the second and third component signals individually.

39. the first component (206) is a luma, the second component (208) is a first chroma component, and the third component (210) is a second chroma component; and the decoder further comprises a first sub-picture level syntax element (614) for adapting the reconstruction of the second component signal for the first color component of the multi-component image. 2 , 616 2 ) and a second sub-picture level syntax element (614) for adapting the reconstruction of the third component signal for the first color component of the multi-component image from the spatially corresponding portions of the reconstructed first or second component signals using the same context in a context-adaptive manner. 3 , 616 3 10. A decoder according to any preceding claim, configured to entropy decode a

40. 39. The decoder of claim 1, wherein the first component is luma, the second component is a first chroma component, and the third component is a second chroma component; and the decoder is configured to entropy decode first sub-picture level syntax elements for adapting the reconstruction of the second component signal for the first color component of the multi-component image and second sub-picture level syntax elements for adapting the reconstruction of the third component signal for the first color component of the multi-component image from the spatially corresponding portions of the reconstructed first or second component signals using a separate context in a context-adaptive manner.

41. The decoder, when parsing the data stream, reads sub-picture level syntax elements from the data stream and adapts the reconstruction of the second component signal based on the reconstructed first component signal at sub-picture granularity based on the sub-picture level syntax elements; when parsing the data stream, checks whether, for a currently decoded portion of the second component, the spatially corresponding portion (442) of the reconstructed first component signal is zero; and, in response to the check, explicitly reading the sub-picture level syntax elements from the data stream and performing the reconstruction of the second component signal from the spatially corresponding portions of the reconstructed first component signal; or Skip the explicit read A decoder according to any preceding claim, arranged to:

42. 10. A decoder according to any preceding claim, wherein the decoder is configured to entropy decode the sub-picture level syntax elements from the data stream using Golomb-Rice coding.

43. 43. The decoder of claim 42, wherein the decoder is configured to binary arithmetically decode bins of the Golomb-Rice code when entropy decoding the sub-picture level syntax elements from the data stream.

44. 44. The decoder of claim 43, wherein the decoder is configured to binary arithmetic decode bins of the Golomb-Rice code at different bin positions using different contexts when entropy decoding the sub-picture level syntax elements from the data stream.

45. 45. The decoder of claim 43 or claim 44, wherein the decoder is configured to binary arithmetically decode bins of the Golomb-Rice code at non-context bin positions above a predetermined value when entropy decoding the sub-picture level syntax elements from the data stream.

46. 10. A decoder according to any preceding claim, wherein the decoder is configured to, when reconstructing the second component signal, spatially rescale and / or perform bit-depth fine mapping on the spatially corresponding parts of the reconstructed first component signal.

47. the decoder is configured to adapt the spatial rescaling and / or the bit depth fine mapping performance in a backward and / or forward adaptive manner.

47. A decoder according to claim 46.

48. The decoder Selecting spatial filters in a backward and / or forward adaptive manner and adapted to adapt the spatial rescaling by 47. A decoder according to claim 46.

49. The decoder In a backward and / or forward adaptive manner, Select a mapping function adapted to adapt the performance of the bit depth fine mapping by 49. A decoder according to claim 48.

50. 10. A decoder according to any preceding claim, wherein the decoder is configured, when reconstructing the second component signals, to reconstruct the second component signals from spatially low-pass filtered versions of the reconstructed first component signals.

51. 51. The decoder of claim 50, wherein the decoder is configured to perform the reconstruction of the second component signal from the spatially low-pass filtered version of the reconstructed first component signal in a forward-adaptive manner or in a backward-adaptive manner.

52. 51. The decoder of claim 50, wherein the decoder is configured to adapt the reconstruction of the second component signal from the spatially low-pass filtered version of the reconstructed first component signal by setting a low-pass filter used for the low-pass filtering in a forward-adaptive manner or in a backward-adaptive manner.

53. 46. ​​The decoder of claim 45, wherein the decoder is configured to perform the spatial low-pass filtering to result in the low-pass filtered version of the reconstructed first component signal using binning.

54. The reconstructed first component signal (256 1 ;270 1 ) by inter-component prediction based on spatially corresponding portions of the second component signal (256') associated with the second component (208) of the multi-component image (202). 2 ;270' 2 ) and further encoding (400) a portion (440) of the inter-component prediction signal (256) for correcting the inter-component prediction. 2 ;270 2 ) into the data stream 1. An encoder configured to encode a multi-component image (202) that spatially samples a scene with respect to different components (206, 208, 210) by:

55. 1. A method for decoding a multi-component image (202) that spatially samples a scene with respect to different components (206, 208, 210), comprising: a first component signal (256) associated with a first component (206) of said multi-component image (202) from said data stream (104); 1 ;270 1 ), and the reconstructed first component signal (256 1 ;270 1 ) and a correction signal (256) derived from said data stream. 2 ;270 2 ) to generate a second component signal (256') for a second component (208) of the multi-component image (202). 2 ;270' 2 ) and reconstructing (400) a portion (440) of A method comprising:

56. A method for encoding a multi-component image (202) that spatially samples a scene with respect to different components (206, 208, 210), comprising: The reconstructed first component signal (256 1 ;270 1 a second component signal (256') related to the second component (208) of said multi-component image (202) by inter-component prediction based on spatially corresponding portions of said second component signal (256') 2 ;270' 2 ) and a correction signal (256) for correcting said inter-component prediction. 2 ;270 2 ) into the data stream.

57. 57. A computer program having a program code for performing the method according to claim 55 or 56, when the computer program runs on a computer.